Enriching Events with Big Data of Dynamic Types for Event Processing
By using CQL processors and Map-Reduce framework in HBase databases, the event streams are processed across multiple processing nodes, which solves the problem of inefficient processing of real-time data event streams in traditional databases, and realizes efficient and real-time event processing capabilities, which are suitable for a variety of distributed event processing scenarios.
Patent Information
- Application Number
- CN202111110048.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2015-06-30
- Filing Date
- 2015-09-21
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2035-09-21
AI Technical Summary
Traditional database models are difficult to efficiently handle real-time, potentially continuous data event streams, especially in continuous event processing (CEP) that requires a large amount of state maintenance. The prior art is often single-threaded, resulting in inefficiency.
The HBase database repository is used as the data source, and the CQL processor is used to process event streams distributed across multiple processing nodes, and events are distributed and processed by defining different distribution streams (such as partitioning, load balancing, fan-in, broadcast streams), and parallel computing is performed in combination with the Map-Reduce framework.
It realizes efficient and real-time processing of large-scale data event streams, improves the efficiency and scalability of event processing, and is suitable for various distributed event processing scenarios such as social media analysis, intelligent instrument energy consumption monitoring and financial risk analysis.
Smart Images

Figure CN113792096B_ABST
Abstract
Description
[0001] This application is a divisional application of the patent application for invention titled "Enriching Events with Big Data of Dynamic Types for Event Processing" with the application date of September 21, 2015 and the application number 201580048664.8.
[0002] Cross-reference to related applications
[0003] This application claims the priority and benefit of U.S. Patent Application No. 14 / 755,088 (Attorney Docket No. 88325-937627 (153500)), filed on June 30, 2015, which claims the priority benefit of U.S. Provisional Application No. 62 / 054,732 (Attorney Docket No. 88325-908381 (153501US)), filed on September 24, 2014. The entire disclosure of each application is hereby incorporated by reference in its entirety for all purposes. Background art
[0004] Databases are commonly used in applications that require data storage and the ability to query the stored data. Thus, existing databases are most suitable for running queries on a finite set of stored data. However, traditional database models are less suitable for an increasing number of modern applications in which data is received as a stream of data events rather than as a bounded set of data. A data stream (also referred to as an event stream) is characterized by a sequence of real-time, potentially continuous events. Thus, a data or event stream represents an unbounded set of data. Examples of sources that generate data streams include sensors and detectors configured to send sequences of sensor readings (e.g., radio frequency identifier (RFID) sensors, temperature sensors, etc.), financial quote providers, network monitoring and traffic management applications that send network status updates, clickstream analysis tools, and others.
[0005] Continuous event processing (CEP) is a technique for processing data in an event stream. CEP is highly stateful. CEP involves continuously receiving events and finding some pattern among those events. Thus, a large amount of state maintenance is involved in CEP. Because CEP involves the maintenance of so much state, the process of applying a CEP query to data within an event stream is always single-threaded. In computer programming, single-threaded means processing one command at a time.
[0006] CEP query processing typically involves continuously executing a query relative to events specified within an event stream. For example, CEP query processing can be used to continuously observe the average price of stocks over the most recent hour. In such a scenario, CEP query processing can be executed relative to an event stream that contains events where each event indicates the current price of a stock at each point in time. The query can aggregate the stock prices over the past hour and then compute the average of those stock prices. The query can output each computed average. As the one-hour window of prices moves, the query can be continuously executed and the query can output various different average stock prices.
[0007] A continuous event processor is capable of receiving a continuous stream of events and is capable of processing the events by applying a CEP query to each event contained within the continuous event stream. Such a CEP query can be formatted to conform to the syntax of a CEP query language, where the CEP query language is such as the Continuous Query Language (CQL) which is an extension of the Structured Query Language (SQL). While SQL queries are typically applied (in response to a user request) once to data that has been stored in the tables of a relational database, CQL queries are repeatedly applied to events as the events in the incoming event stream are received by the continuous event processor. Thus, there is clearly a need for a system that allows for efficient execution of CEP. Summary of the Invention
[0008] Embodiments described herein relate to databases and continuous event processing. According to some embodiments, the processing of CQL queries can be distributed across different processing nodes. The event processing mechanism can be distributed across multiple separate virtual machines. The described embodiments provide a system that efficiently allows for CEP for real-time continuous data queries.
[0009] According to some embodiments, an HBase database repository is used as a data source for a Continuous Query Language processor (CQL processor). Such use allows for the enrichment of events with data present in the repository, similar to how events can be enriched with data present in RDBMS tables. According to some embodiments, the HBase database repository is used as a data sink similar to a table sink feature.
[0010] The foregoing and other features and embodiments will become more apparent by reference to the following specification, claims, and drawings. Brief Description of the Drawings
[0011] Figure 1 is a diagram showing an example of a table in an HBase data repository according to some embodiments.
[0012] Figure 2 is a block diagram showing an example of a simple event processing network according to some embodiments.
[0013] Figure 3 is a block diagram showing an example of a broadcast event processing network according to some embodiments.
[0014] Figure 4 is a block diagram showing an example of a load - balanced event processing network according to some embodiments.
[0015] Figure 5 is a block diagram showing an example of a subsequent state of a load - balanced event processing network according to some embodiments.
[0016] Figure 6 is a block diagram showing an example of a broadcast event processing network in which a channel has two consumers according to some embodiments.
[0017] Figure 7 is a flowchart showing an example of a technique for generating a single token that can be used to request services from multiple resource servers according to an embodiment of the present invention.
[0018] Figure 8 is a block diagram showing an example of a partitioned event processing network according to some embodiments.
[0019] Figure 9 is a block diagram showing another example of a partitioned event processing network according to some embodiments.
[0020] Figure 10 is a block diagram showing an example of a fan - in event processing network according to some embodiments.
[0021] Figure 11 is a diagram showing an example of a line graph according to some embodiments.
[0022] Figure 12 is a diagram showing an example of a scatter plot according to some embodiments.
[0023] Figure 13 is a diagram showing an example of a scatter plot in which a smoothed curve - fitting has been drawn according to some embodiments.
[0024] Figure 14 is a diagram showing an example of a scatter plot in which points have different sizes according to some embodiments.
[0025] Figure 15 is a diagram showing an example of a radar chart according to some embodiments.
[0026] Figure 16 draws a simplified diagram of a distributed system for implementing one of the embodiments.
[0027] Figure 17 is a simplified block diagram of components of a system environment according to an embodiment of the present disclosure, through which services provided by components of an embodiment system can be provided as cloud services.
[0028] Figure 18 Shows an example of a computer system in which various embodiments of the present invention can be implemented.
[0029] Figure 19 Is a diagram showing an example of the shape of a cluster represented superimposed on a scatter plot according to some embodiments. Detailed Description
[0030] In the following description, for the purpose of explanation, specific details are set forth in order to provide a thorough understanding of embodiments of the present invention. However, it will be clear that the present invention may be practiced without these specific details.
[0031] Processing for an event processing application can be distributed. The Oracle Event Processing product is an example of an event processor (continuous query language processor). According to some embodiments, the processing of CQL queries can be distributed across different processing nodes. For example, each such processing node can be a separate machine or computing device. When distributing the processing of CQL queries across different processing nodes, the semantics for ordering events are defined in some way.
[0032] A preliminary approach to ordering events attempts to maintain a first-in, first-out (FIFO) order among events in an event stream. However, some event streaming systems can involve multiple event publishers and multiple event consumers. Each machine in the system can have its own clock. In such a scenario, the event timestamps generated by any single machine may not be definitive across the entire system.
[0033] Within a system involving multiple event consumers, each consumer may have a separate set of requirements. Each consumer can be an event processor that continuously executes CQL queries. In terms of event ordering, each such CQL query may have separate requirements.
[0034] According to some embodiments, a distribution flow is defined. Each distribution flow is a specific way of distributing events between an event producer and an event consumer. One distribution flow can load-balance events among a set of event consumers. For example, when an event producer produces a first event, the load-balancing distribution flow can cause the first event to be routed to a first event consumer. Subsequently, when the event producer produces a second event, the load-balancing distribution flow can cause the second event to be routed to a second event consumer.
[0035] Other types of distribution flows include partitioned distribution flows, fan-in distribution flows, broadcast flows, etc. Depending on the type of distribution flow being used and also depending on the requirements of the event consumers receiving the events, different event ordering techniques can be used to order the events received by the event consumers.
[0036] Overview of Map-Reduce
[0037] Reference is made herein to Map-Reduce, which is a framework for processing parallelizable problems involving huge data sets using a large number of computing machines (nodes). If all nodes are on the same local area network and use similar hardware, these nodes are collectively referred to as a cluster. Alternatively, if the nodes are shared across geographically and administratively distributed systems and use more heterogeneous hardware, these nodes are collectively referred to as a grid. Computational processing can be performed with respect to unstructured data such as might be found in a file system or structured data such as might be found in a database. Map-Reduce can take advantage of data locality to process data on or near the storage assets in order to reduce the distance the data is transmitted.
[0038] In the "map" step, the master node receives a task as input, divides the input into smaller sub-problems, and distributes the sub-problems to worker nodes. A given worker node can repeat this division and distribution, resulting in a multi-level tree structure. Each worker node processes the sub-problem assigned to it and passes the result of the processing back to its master node.
[0039] In the "reduce" step, the master node collects the results of the processing of all sub-problems. The master node combines these results in some way to form the final output. The final output is the product of the task initially given to the master node to execute.
[0040] Map-Reduce allows for the distributed processing of map and reduce operations. Assuming that each map operation is independent of the others, all map operations can be executed in parallel. Similarly, if all the outputs of map operations sharing the same key are presented to the same reducer at the same time, or if the reduce function is associative, a set of reduce nodes can perform the reduce phase. In addition to reducing the total time required to produce the final result, parallelism also provides some possibility of recovering from partial failures of servers or storage devices during the operation. If a map node or a reduce node fails, the work can be rescheduled if the input data is still available.
[0041] Map-Reduce can be conceptualized as a parallel and distributed computation in five steps. In the first step, the map input is prepared. The Map-Reduce system assigns map processors, allocates a first input key-value on which each map processor can work, and provides each map processor with all the input data associated with that first input key-value.
[0042] In the second step, the map nodes execute the map code provided by the user. The map nodes execute the map code once for each of the first key-values. The execution of the map code generates an output organized by second key-values.
[0043] In the third step, the output from the second step is shuffled to the reduce nodes. The Map-Reduce system assigns reduce processors. The Map-Reduce system allocates a second key-value to each reduce processor on which that processor will work. The Map-Reduce system provides each reduce processor with all the data generated during the second step that is also associated with the reduce processor's allocated second key-value.
[0044] In the fourth step, the reduce nodes execute the reduce code provided by the user. The reduce nodes execute the reduce code once for each of the second key-values generated during the second step.
[0045] In the fifth step, the final output is produced. The Map-Reduce system collects all the output data generated by the fourth step and sorts the data by their second key-values to produce the final output.
[0046] Although the above steps can be imagined to run sequentially, in practice, these steps can be interleaved as long as the final output is not affected by the interleaving.
[0047] Event processing scenarios that benefit from distribution
[0048] Since the amount of data to be analyzed has grown tremendously today, a scalable event handling mechanism is very useful. Scalability in this context can involve not only an increase in the number of processing threads involved in event handling execution, but also an increase in the number of computing machines that can process events in parallel. Techniques for distributing event handling applications across multiple virtual machines, such as Java Virtual Machines (JVMs), are disclosed herein.
[0049] Many different event handling scenarios are well-suited for distributed execution. These scenarios tend to have certain characteristics. First, these scenarios are not strictly latency-bound, but can involve latencies in the microsecond range, for example. Second, these scenarios can be logically partitioned, for example, by consumer or by region. Third, these scenarios can be logically divided into individual components or tasks that can be executed in parallel, such that there are no total ordering constraints.
[0050] An example of an event handling scenario that can be usefully executed in a distributed manner is the word count scenario. In this scenario, the system maps incoming sentences into meaningful words and then reduces these words to a count (for each word). The work performed in this word count scenario can be executed using Map-Reduce batching, but can also be executed using stream processing. This is because, with stream processing, a real-time word stream, such as a real-time word stream from Twitter or another social media feed, can be counted. Using stream processing with respect to a social media feed allows for a faster response than might be achievable with other processing methods.
[0051] If stream processing is used to handle social media feeds (e.g., count the words in these feeds), the stream processing mechanism may be subjected to a very large number of incoming words. To handle large amounts of information, the processing of the information can be distributed. Separate computing machines can subscribe to different social media streams, such as Twitter streams. These machines can process the streams in parallel, count the words therein, and then converge the resulting counts to produce a complete result.
[0052] Another example of an event handling scenario that can be usefully executed in a distributed manner is the matrix multiplication scenario. The page ranking algorithm used by Internet search engines can summarize the importance of a web page as a single number. Such an algorithm can be implemented as a series of cascaded large matrix multiplication operations.
[0053] Since matrix multiplication can be highly parallelized, Map-Reduce can be beneficial for performing operations involving matrix multiplication. Matrix multiplication operations can be conceptualized as a natural join, followed by grouping and aggregation.
[0054] Another example of an event processing scenario that can be usefully performed in a distributed manner is the term frequency - inverted document frequency (TF-IDF) scenario. TF-IDF is an algorithm often employed by search engines to determine the importance of terms. Contrary to expectations, if a term is frequently seen in other documents, then that term is less important, and thus the "inverted document frequency" aspect of the algorithm is less significant.
[0055] As with the word count scenario discussed above, it is valuable to be able to perform TF-IDF processing in real time using stream processing. Unlike in the word count scenario, the calculation of TF-IDF values involves accessing historical documents for the "inverted document frequency" calculation. The involvement of historical documents makes the TF-IDF scenario a good candidate for use with Hadoop and / or some indexing layer (such as HBase). Hadoop and HBase and their use in event stream processing are also discussed in this article.
[0056] Another example of an event processing scenario that can be usefully performed in a distributed manner is the smartmeter energy consumption scenario. Today's homes typically collect their energy consumption by using smart meters located in their houses. Generally, these smart meters output energy consumption sensor data in the form of events periodically (e.g., every minute) throughout the day. This sensor data is captured by downstream management systems in regional processing centers. These centers use the captured data to calculate useful descriptive statistics, such as the average energy consumption of a house or a neighborhood. The statistics can reveal how the average energy consumption relates to the historical data for the region. These running aggregates are well-suited to being partitioned. Thus, distributed partitioned streams can be beneficially applied to this scenario.
[0057] Event processing can be performed with respect to this information in order to identify outliers (such as homes with energy consumption above or below the typical range). Event processing can be performed with respect to this information in order to attempt to predict future consumption. The identified outliers and predicted future consumption can be used by energy providers for differential pricing, promotions, and more efficiently controlling the energy purchase and sales processes with their partners.
[0058] Various other scenarios not specifically enumerated above can be suitable for processing in a distributed manner. For example, distributed event stream processing can be used to perform risk analyses that involve calculating the risk exposure of a financial portfolio in real time as derivative prices change.
[0059] Method for distributing an event stream to create a stream
[0060] This disclosure presents several different techniques for distributing an event stream (e.g., a stream originating from a particular data source) across multiple event processing nodes to facilitate parallel event processing. Each technique creates a different kind of stream. Some of these techniques are summarized below.
[0061] Partitioned streams involve partitioning an event stream across several separate computing resources (such as processing threads or virtual machines (e.g., JVMs)) using one or more attributes of the events in the stream as a partitioning criterion. A clustered version of partitioned streams is disclosed herein. Also disclosed herein is partitioning of the streamed events across threads of an event processing network in a single-node configuration.
[0062] Fan-in streams involve gathering multiple previously distributed event streams back into a single computing resource. For example, fan-in streams can be used when a certain state is to be co-located, such as in scenarios involving global aggregation.
[0063] Load-balance streams involve distributing the events of a stream to a group of consuming listeners in such a way that the total load is shared across the group of consuming listeners in a balanced manner. An event processing node that is currently loaded with less work can be selected before an event processing node that is currently loaded with more work to receive new events for processing. This prevents any one event processing node from becoming overloaded while other event processing nodes remain underutilized.
[0064] Broadcast streams involve broadcasting all the events of a stream to all consuming listeners. In this case, all listeners - such as event processing nodes - will receive a copy of all the events.
[0065] Cluster domain generation
[0066] To support a distributed event processing network corresponding to the various streams discussed above, some embodiments involve the generation of a cluster domain. In some embodiments, a configuration wizard or other tool guides a user through the generation of a domain configured to support distributed streams.
[0067] Resource elasticity
[0068] In a cloud computing environment, computing resources can grow or shrink dynamically as demand increases or decreases. According to some embodiments, a distributed event processing system deployed in a cloud computing environment can be plugged into an existing infrastructure. The system can dynamically grow and shrink the number of computing resources currently executing distributed streams. For example, in the case of increasing load, load-balance streams can automatically spawn new computing resources to further share the load.
[0069] Distributed stream defined as an event processing network
[0070] The event processing network corresponding to the stream can be represented as an acyclic directed graph. The graph can be formally defined as a pair (N, C), where N is a set of nodes (vertices) and C is a two-place relation on N representing the connections (arcs) from source nodes to destination nodes.
[0071] For example, an event processing network can be defined as event processing network1 = ({adapter1, channel1, processor, channel2, adapter2}, {(adapter1, channel1), (channel1, processor), (processor, channel2), (channel2, adapter2)}) (event processing network1 = ({adapter1, channel1, processor, channel2, adapter2}, {(adapter1, channel1), (channel1, processor), (processor, channel2), (channel2, adapter2)})). An event is defined as a relation P on any pair (PN, PV) representing an attribute name and an attribute value.
[0072] For another example, given an event stream representing a stock ticker, the following definitions can be used: e1 = {(price, 10), (volume, 200), (symbol, ‘ORCL’)}; e2 = {(p1, v1), (p2, v2), (p3, v3)}, (e1 = {(price, 10), (volume, 200), (symbol, ‘ORCL’)}; e2 = {(p1, v1), (p2, v2), (p3, v3)}). Since an event processing network node can contain more than one event, the set E can be defined as an ordered sequence of events (which is different from some other cases).
[0073] For another example, the following definition can be used: {processor} = {e1, e2} ({processor} = {e1, e2}). The runtime state S = (N, E) of the event processing network can be represented as a two-place relation from N to E. The relation S is not injective, meaning the same event(s) can exist in more than one node. However, the relation S is surjective because all events in the total set of events in the event processing network are in at least one node.
[0074] For another example, the following definition can be used: state = {(processor, {e1, e2}), (adapter2, {e3})} (State = {(processor, {e1, e2}), (adapter2, {e3})}). This provides a logical model of the event flow. This model can be augmented with physical components. This augmentation can be done by allocating the computing resources R of the nodes that host the event handling network. The new model then becomes the three-place relation S = (N, R, E), where R is the set of all computing resources of the cluster.
[0075] For another example, the following definition can be used: state = {(processor, machine1, {e1, e2}), (adapter2, machine1, {e3})} (State = {(processor, machine1, {e1, e2}), (adapter2, machine1, {e3})}).
[0076] Thus, distributed flows can be defined as functions. These functions can take as input the static structure of the event handling network, the current state at runtime, and the specific nodes of a particular computing resource as the subject. The function can return a new configuration of the runtime state that takes into account the flow of events from the subject to the connections of that subject. Formally, this can be defined as: distribute-flows: n ε N, r ε R, C, S → S (Distribute Flows: n ε N, r ε R, C, S → S).
[0077] A number of functions are defined to support patterns for fan-in flows, load-balanced flows, partitioned flows, and broadcast flows, the latter two having two versions: one for the single virtual machine case and another for the clustered virtual machine case. The distribute flow function is:
[0078] These are interpreted as follows: For all sources n that exist in both C and S (i.e., having connections and having events), then for each destination d in C, generate new tuple states (d, r, e). Return the current state S minus the old tuples (n, r, e) plus the new tuples (d, r, e). The removal and addition in the last step represent the movement of events from the source node to the destination node.
[0079]
[0080] In this case, the new state S' includes tuples of all valid permutations of the destination d and the resource t. That is, all computing resources will receive events for each configured destination.
[0081]
[0082] Since threading is not modeled in this definition, there is no difference between local - partition and local - broadcast. However, in practice this is not the case because threading is colored by partition.
[0083]
[0084] As will be seen from the following discussion, the last three are structurally similar, differing only in their scheduling functions. In fact, fan - in can be regarded as a special case of a partition with a single key.
[0085] Scheduling functions
[0086] According to some embodiments, scheduling functions 1b - sched, p - sched, and fi - sched are defined. The implementation of these functions does not change the structure of the distribution. The functions determine the scheduling of resources. The default implementation of the functions is: lb - sched(R) = {R → r ε R: r = round - robin(R)}. lb - sched uses a conventional round - robin algorithm. In this case, a certain cluster state can be maintained.
[0087] According to some embodiments, min - jobs scheduling is used, where min - jobs selects the resource that has had the smallest number of jobs scheduled to run so far.
[0088] According to some embodiments, the target server is randomly selected. Considering the law of large numbers, this embodiment is similar to round - robin except that no centralized state is required: p - sched(e, pn) = {e ε E, pn ε PN → r ε R: r = hash(prop(e, pn)) mod |R|}; and fi - sched(n, R) = {n ε N, R → r ε R: r = user - configured - server(n, R)}.
[0089] In some embodiments, aggregation is performed to a single server having the resources required to process the set of events. For example, if the event is to be output to an event data network / JAVA messaging service, such a server can maintain information about the event data network / JAVA messaging service servers and destination configurations.
[0090] In some embodiments, the fan - in target is by default selected as the cluster member with the lowest member ID, which will indicate the first server of the member to be configured.
[0091] Example event - handling network
[0092] Figure 2 is a block diagram showing an example of a simple event processing network 200 according to some embodiments. In Figure 2 it, the event processing network 200 can be defined using the following syntax: local-broadcast(channel1,machine1,eventprocessing network1,{(channel1,machine1,{e1})})={(processor,machine1,{e1})}(local-broadcast(channel 1,machine 1,event processing network 1,{(channel 1,machine 1,{e1})})={(processor,machine 1,{e1})}).
[0093] The following discusses some cluster situations where R = {machine1, machine2, machine3} (R = {machine 1, machine 2, machine 3}). Figure 3 is a block diagram showing an example of a broadcast event processing network 300 according to some embodiments. In Figure 3 it, the event processing network 300 can be defined using the following syntax: clustered-broadcast(channel1,machine1,event processing network1,{(channel1,machine1,{e1})})={(processor,machine1,{e1}),(processor,machine2,{e1}),(processor,machine3,{e1})}(clustered-broadcast(channel 1,machine 1,event processing network 1,{(channel 1,machine 1,{e1})})={(processor,machine 1,{e1}),(processor,machine 2,{e1}),(processor,machine 3,{e1})}). In this case, the same event e1 is distributed to all of the computing resources machine1, machine2, and machine3.
[0094] Figure 4 is a block diagram showing an example of a load balancing event processing network 400 according to some embodiments. In Figure 4In
[15] , event processing network 400 can be defined using the following syntax: load-balance(channel1,machine1,eventprocessing network1,{(channel1,machine1,{e1})})={(processor,machine2,{e1})}(load-balance(channel1,machine1,eventprocessing network1,{(channel1,machine1,{e1})})={(processor,machine2,{e1})}). In the case of load balancing, the state of machine 2 is changed when executing a job for {e1}.
[0095] Figure 5 is a block diagram illustrating an example of subsequent states of a load-balanced event processing network 500 according to some embodiments. Figure 5 Shows that after e1 is sent to machine 2 Figure 4 The load balancing event handles the state of the network. Figure 5 In
[15] , event processing network 500 can be defined using the following syntax: load-balance(channel1,machine1,eventprocessing network1,{(channel1,machine1,{e2,e3})})={(processor,machine3,{e2,e3})}(load-balance(channel1,machine1,eventprocessing network1,{(channel1,machine1,{e2,e3})})={(processor,machine3,{e2,e3})}). In this case, both e2 and e3 are sent to the same machine. This is because they both exist in the source node, so it makes sense to keep them together.
[0096] In some event processing networks, a single channel can have two consumers, such as event processing network 2 = ({adapter1, channel1, processor1, processor2, channel2, channel3, adapter2, adapter3}, {(adapter1, channel1), (channel1, processor1), (processor1, channel2), (channel2, adapter2), (channel1, processor2), (processor2, channel3), (channel3, adapter3)}) (event processing network 2 = ({adapter1, channel1, processor1, processor2, channel2, channel3, adapter2, adapter3}, {(adapter1, channel1), (channel1, processor1), (processor1, channel2), (channel2, adapter2), (channel1, processor2), (processor2, channel3), (channel3, adapter3)})). In the case of cluster broadcasting in this scenario, all processors in all machines receive the event.
[0097] Figure 6 is a block diagram showing an example of a broadcast event processing network 600 in which a channel has two consumers, according to some embodiments. In Figure 6In [the above], the event processing network 600 can be defined using the following syntax: clustered - broadcast(channel1, machine1, event processing network2,{(channel1, machine1,{e1})}) ={(processor1, machine1,{e1}),(processor2, machine1,{e1}),(processor1, machine2,{e1}),(processor2, machine1,{e1}),(processor,1 machine3,{e1}),(processor2, machine3,{e1})}(clustered - broadcast(channel 1, machine 1, event processing network2,{(channel 1, machine 1,{e1})}) ={(processor1, machine 1,{e1}),(processor2, machine 1,{e1}),(processor1, machine 2,{e1}),(processor2, machine 1,{e1}),(processor,1 machine 3,{e1}),(processor2, machine 3,{e1})}). Within a machine (e.g., machine 1), the dispatch of an event (e.g., e1) to its consuming listeners (e.g., processor 1, processor 2) can occur synchronously (i.e., in the same thread) or asynchronously (i.e., in different threads) depending on the ordering requirements.
[0098] Figure 7 is a block diagram showing an example of a load - balanced event processing network 700 in which a channel has two consumers according to some embodiments. In this case, there are multiple listeners. In Figure 7 [the above], the event processing network 700 can be defined using the following syntax: load - balance(channel1, machine1, event processing network2,{(channel1, machine1,{e1})}) ={(processor1, machine1,{e1}),(processor2, machine1,{e1})}(load - balance(channel 1, machine 1, event processing network2,{(channel 1, machine 1,{e1})}) ={(processor1, machine 1,{e1}),(processor2, machine 1,{e1})}). In this case, the event is sent to all listeners of a single member. In other words, only the next arriving event will be load - balanced to different servers or machines.
[0099] The partitioning scenario and the fan - in scenario can be considered to reuse the simple event processing network again. Figure 8is a block diagram showing an example of a partitioned event processing network 800 according to some embodiments. In Figure 8 it, the event processing network 800 can be defined using the following syntax: partition(channel1,machine1,event processing network1,{(channel1,machine1,{e1(p1,1)})}):-{(processor,machine1,{e1(p1,1)})} (partition(channel1,machine1,event processing network1,{(channel1,machine1,{e1(p1,1)})}):-{(processor,machine1,{e1(p1,1)})}). Next, event e2 can be considered to be on the same partition as event e1, but event e3 is on a different partition. This can be defined using the following syntax: partition(channel1,machine1,eventprocessing network1,{(channel1,machine1,{e2(p1,1)})}):-{(processor,machine1,{e2(p1,1)})}; partition(channel1,machine1,event processing network1,{(channel1,machine1,{e3(p1,2)})}):-{(processor,machine2,{e3(p1,2)})} (partition(channel1,machine1,event processing network1,{(channel1,machine1,{e2(p1,1)})}):-{(processor,machine1,{e2(p1,1)})}; partition(channel1,machine1,event processing network1,{(channel1,machine1,{e3(p1,2)})}):-{(processor,machine2,{e3(p1,2)})}).
[0100] According to some embodiments, in a partitioned event processing network, events reach all machines instead of only reaching certain machines. Figure 9 is a block diagram showing an example of a partitioned event processing network 900 according to some embodiments. In some embodiments, a processor can have multiple upstream channels fed into the processor. This situation is similar to processing multiple events.
[0101] Figure 10 is a block diagram showing an example of a fan-in event processing network 1000 according to some embodiments. In Figure 10In it, the event processing network 1000 can be defined using the following syntax: fan-in(channel1, machine1, eventprocessing network1, {(channel1, machine1, {e1})}):-{(processor, machine1, {e1})}; fan-in(channel1, machine2, event processing network1, {(channel1, machine2, {e2})}):-{(processor, machine1, {e2})} (fan-in(channel1, machine1, event processing network1, {(channel1, machine1, {e1})}):-{(processor, machine1, {e1})}; fan-in(channel1, machine2, event processing network1, {(channel1, machine2, {e2})}):-{(processor, machine1, {e2})}). In the case of fan-in, events are aggregated together in machine1.
[0102] Sorting and query processing semantics
[0103] In some embodiments, events do not instantaneously appear from one node to all other destination nodes; in such embodiments, total order is not always maintained. The sorting requirements can be safely relaxed in certain scenarios without breaking the semantics of the distribution model and the query processing model.
[0104] There are two dimensions to be considered, namely, the dimension of sorting between machine destinations for a single event, and the dimension of sorting of the event itself when sent to the destination. Each dimension can be considered separately.
[0105] For destination sorting, in load-balanced networks, partitioned networks, and fan-in networks, since there is a single destination (i.e., ), destination sorting is not applicable. In cluster broadcast networks, due to the general broadcast nature, sorting guarantees do not need to be envisioned.
[0106] For event sorting, in load-balanced networks, since it cannot be guaranteed that events will initially be sent to the same resource, there is no benefit in guaranteeing the sorting of events, and thus downstream query processing does not attempt to use application time sorting in some embodiments. Additionally, in some embodiments, downstream query processing does not depend on receiving all events and is thus stateless. The types of queries that fall into this criterion are filtering and stream joins (1-n stream relationship joins).
[0107] In a cluster broadcast network, all servers have a complete set of events and thus a full state of processing. Therefore, in such a network, order is maintained in the context of each server (i.e., the destination).
[0108] In a partitioned network, for a particular destination, ordering is guaranteed within the partition. This allows downstream query processing to similarly utilize the partition ordering constraints. In some embodiments, this ordering is guaranteed even if there are multiple upstream nodes feeding events to be partitioned (as in one of the cases described above in connection with Figure 9 the situation).
[0109] In a fan-in network, a determination is made as to how events are initially forked as follows: If the events are load-balanced, there is no ordering guarantee, and the fan-in function does not impose any order. The following is an example of what happens if the events are partitioned in one embodiment:
[0110] Input: {{t4,b},{t3,a},{t2,b},{t1,a}}
[0111] Partition a: {{t3,a},{t1,a}}
[0112] Partition b: {{t4,b},{t2,b}}
[0113] Schedule 1: {{t4,b},{t3,a},{t1,a},{t2,b}}
[0114] Schedule 2: {{t4,b},{t1,a},{t2,b},{t3,a}}
[0115] Schedule 3: {{t4,b},{t3,a},{t2,b},{t1,a}}
[0116] In the case of upstream partitioning, the fan-in may ultimately order the events in a different order than the original input. To avoid this, the fan-in network sorts the events even though they are received from different sources.
[0117] To handle these different scenarios, different semantics are used in different scenarios. In the case of an unordered scenario, there is no ordering guarantee between events according to the timestamps of the events.
[0118] In the case of a partially partition-ordered scenario, according to some embodiments, it is guaranteed that the events are ordered according to their timestamps in the context of source and destination node pairs and within the partitions (i.e., a <= b). In other words, events from different upstream servers are not guaranteed to be ordered, and events going to different partitions are likewise not guaranteed to be ordered.
[0119] In the case of a fully partition - ordered scenario, ensure that events are ordered across all source and destination node pairs according to their timestamps and are ordered within the context of the partitions (i.e., a <= b). To support this pattern, a single - view timestamp can be imposed across the cluster. The applied timestamp can be used for this case.
[0120] In the case of a partially ordered scenario, in some embodiments, ensure that events are ordered according to their timestamps within the context of source and destination node pairs (i.e., a <= b).
[0121] In the case of a fully ordered scenario, in some embodiments, ensure that events are ordered across all source and destination node pairs according to their timestamps (i.e., a <= b).
[0122] These constraints have been presented from the least - restrictive to the most - restrictive. To support these different constraints, in some embodiments, the following additional configurations are used. The property of the applied timestamp is the event property to be used for the total - order criterion. The timeout property indicates the time to wait for upstream events before proceeding. The out - of - order policy indicates whether events should be discarded, presented as an error, or sent to a dead - letter queue if they arrive out of order.
[0123] In some embodiments, each distribution flow can be used with a different set of ordering constraints. In a load - balanced network, the constraints can be unordered. In a broadcast network, all events of an input flow can be propagated to all nodes of the network in the same order in which they are received on the broadcast channel. Each node can maintain a full state, and thus, each node listening to the broadcast channel has exactly the same state for any timestamp. Therefore, the listeners downstream of each of these nodes can receive output events in total order, and the constraints can be fully ordered in a local broadcast. For a cluster broadcast, event delivery across the network may result in the events being unordered, but by definition, the delivery should be fully ordered so that the network can meet the requirements of ordered delivery. In a partitioned network, each node can maintain a partial state and receive a subset of events. The events of the received sub - flow are in the same order as observed in the input flow. Thus, the input flow is ordered across partitions (one partition per node), and the constraints can be partition - ordered in a partitioned network. In a cluster - partitioned network, event delivery across the network may result in the events being unordered, but again, by definition, the delivery should be fully ordered so that the network can meet the requirements of ordered delivery. In a fan - in network, the constraints can be unordered, partially ordered, or fully ordered.
[0124] In addition, if the destination node is a CQL processor and its queries are known, the distribution sorting constraints can be inferred from these queries. For example, if all queries are configured to be partition-ordered, the distribution flow can also be set to be at least partially partition-ordered.
[0125] Deployment plan
[0126] In some embodiments, computing resources are shared across nodes. To allow for better resource sharing, a set of constraint requirements can be used to annotate nodes, such as'memory>1M (memory > 1M)' or 'thread>3 (threads > 3)', and in turn, a set of capabilities can be used to annotate computing resources, such as'memory = 10M (memory = 10M)' or 'thread-pool = 10 (thread pool = 10)'.
[0127] For example, a requirement can be expressed as requirements:{processor1}={threads>3} (requirement: {processor1} = {threads > 3}). A capability can be expressed as capabilities:{machine1}={max-thread-pool = 10, cpu = 8} (capability: {machine1} = {max-thread-pool = 10, cpu = 8}).
[0128] During the scheduling of resources to nodes, the system attempts to match requirements with capabilities, and by doing so, the system dynamically reduces and increases the current value of capabilities as capabilities are assigned to nodes. For example, the scheduling can be expressed as Schedule-1:{processor1}={threads>3, computing-resource = machine1}; Schedule-1:{machine1}={max-thread-pool = 10, current-thread-pool = 7} (Schedule-1: {processor1} = {threads > 3, computing resource = machine1}; Schedule-1: {machine1} = {max-thread-pool = 10, current-thread-pool = 7}).
[0129] In addition, the total capabilities of the cluster itself can be changed, for example, by adding new computing resources to the cluster to handle an increase in application load. This is known as computing elasticity. For example, at t = 0: {cluster}={machine1,machine2} ({cluster} = {machine1, machine2}), but at t = 1: {cluster}={machine1,machine2,machine3} ({cluster} = {machine1, machine2, machine3}). The system copes with these dynamic resource changes.
[0130] There may be a situation where the operator of the system wants to manually assign nodes to specific computing resources. This can be supported by treating 'computing-resource' as a requirement itself. For example, Requirements:{processor1}={threads>3,computing-resource=machine1} (Requirement: {processor1}={threads>3, computing-resource=machine1}). This specification of deployment requirements is known as a deployment plan and can be included in the application metadata.
[0131] Cluster member configuration and domain configuration
[0132] In Hadoop, the map functions at the start of the Map-Reduce system are replicated to distributed tasks and executed in parallel, with each function reading separate input data, or most commonly each function reading a block of the input data. Stream processing is similar. Upstream nodes (i.e., inbound adapters) each subscribe to different streams or different partitions of a stream. This means that in some embodiments, the distributed event processing network allows inbound adapters to work in parallel. There is no need to keep the inbound adapters present in the secondary nodes paused.
[0133] In some embodiments, each inbound adapter can subscribe to different streams or different partitions of a stream. This can be done by using the Clustered Member facility in event processing, where a member can discover whether it is primary, and members in a cluster are associated with a unique ID and can thus use this ID as a key to the stream or stream partition configuration.
[0134] The secondary members can also choose not to subscribe to any events, in which case the input side of the system is not executed in parallel.
[0135] Cost complexity and batch processing
[0136] The communication cost in a distributed system can easily exceed the cost of processing the data itself. In fact, a common problem in Hadoop is to find the best compromise between having too many reducers and thus increasing the communication cost and having too few reducers and thus having too many elements associated with a key and thus not enough memory per reducer.
[0137] To facilitate understanding of this cost and the mechanisms used to address it, the following is provided. The latency metric is calculated as the ratio of the total latency of an event to the communication latency of the event. This is done for a certain sampling rate of the events and can be turned on and off dynamically at runtime. In some embodiments, there is a guarantee that events sent together using the Batching API are indeed batched together throughout the distribution.
[0138] Behavioral Viewpoint
[0139] In some embodiments, cache coherence can be used for both messaging and partitioning. The semantics of cooperation vary by stream and can be implemented using a combination of specific cache schemes, cache keys, and filtering. The sender (source) inserts (i.e., puts()) the event into the cache, and the receiver (destination) removes (i.e., gets / delete) the event from the cache. The MapListener API with Filtered Events can be used to ensure that the correct receiver gets the correct set of events. However, if the event is to be received as a separate action and then deleted, this results in two separate network operations. Therefore, in some embodiments, events are allowed to expire based on their own merit. In this way, cache coherence batches the deletion of events and does this at the appropriate time.
[0140] In some embodiments, the same cache service is shared by all applications for each stream type. For example, there can be a single cache service for all replicated streams, another cache service for partitioned streams, etc. Since locking can be done on a per-entry basis, the handling of an event by a channel in one application does not affect other applications, and this avoids the proliferation of caches in a single server.
[0141] In the case of a broadcast stream, since all events are to be received by all members, a Replicated Cache Scheme can be used. The cache key is a hash of the member ID, application name, event processing network stage name (e.g., channel name), and event timestamp (whether it is application-based or system-based).
[0142] CacheKey = hash(memberId, applicationName, stageName, eventTimestamp)
[0143] The cache value is a wrapper for the (one or more) original events with the application ID, the event handling network stage ID, and the target ID added, where the target ID is set to -1 to represent all members. The application ID and the stage ID are hash values of the original application name and stage name, which are strings set by the user. The wrapper can include an event timestamp (if not application-property-based) and an event kind (i.e., insert, delete, update, heartbeat).
[0144] CacheValue = {applicationId, stageId, eventTimestamp, eventKind, sourceEvents}
[0145] In some embodiments, all cluster members register a MapListener, where a Filter is set to the application ID and the event handling network stage ID for the broadcast channel in question. This means that when a member acting as a sender puts an event into the broadcast cache, all members acting as receivers can be called back on MapListener.entryInserted(MapEvent).
[0146] If the stream is set to be unordered, the MapListener is asynchronous and can use a consistency thread for downstream processing. If the stream is set to be ordered, a SynchronousMapListener can be registered, and the event can be immediately handed over to a single channel thread for downstream processing. In an embodiment, this is done because the entire map can be synchronized, so the work for each channel is queued and the thread immediately returns, allowing other channels to receive their events. The original member node of the sender can receive events through the MapListener.
[0147] In the case of a load-balanced stream, the target computing resource can be selected by finding the total number of members in the cluster and generating a random number between [0, total]. In other words, instead of using the last-used member to maintain a certain cluster state, randomization can be used to achieve load balancing. The key and value can be similar to the broadcast case. However, a MapListener can be registered, where the Filter is set to the application ID, the stage ID, and the target ID, where the value of the target ID is a randomly selected member ID. In other words, in some embodiments, only the randomly selected target will receive the event. According to some embodiments, the load-balanced stream only supports the unordered case, so only asynchronous listeners are used.
[0148] In the case of a fan-in stream, the target computing resource can be directly specified by the user according to a certain configuration mechanism. This user-defined target ID can be set in the cache value wrapper, but otherwise the semantics are similar to the load balancing case. The fan-in stream supports total ordering. In this case, in addition to using a synchronous map listener, the channel can be configured to use application timestamps, and events in the receiver can be reordered until a user-configurable timeout value. The timeout value can be in the range of, for example, several seconds, and can be based on a trade-off between lower latency and a greater chance of out-of-order events.
[0149] In some embodiments, an optimized hash is used that guarantees the cache keys remain ordered. The receiver can then use a filter, where the filter retrieves all entries for a specific channel in a range from the most recent to a certain time in the past. In this case, a Continuous Query Map can be used. The map can be periodically checked using the same timeout configuration.
[0150] In the case of a fan-in stream, the inherent partitioning support of a Partitioned Cache Scheme can be utilized. The cache data affinity can be set up to be associated with a (partition) key consisting of the application ID, stage ID, and the configured partition event attribute (value) (e.g., the value 'ORCL' for the event attribute'symbol'). This can be done by using the KeyAssociation class. The cache keys can remain the same (e.g., with timestamps). However, in some embodiments, all keys have an association to the partition key just described above, thus ensuring the partitions remain co-located.
[0151] If cache coherence is used to place data in the best location, the target member is not selected, and thus an EntryProcessor is used instead of a MapListener, where the filter is set to the application ID, stage ID, and partition event attribute value. In this case, the source node invokes the EntryProcessor and ensures that the EntryProcessor implementation is executed in the member where the data resides, thus avoiding copying the data to a target member that has been explicitly selected. The cache coherence can be optimized using its internal components that determine the correct number of partitions based on the cluster members and data size to fully utilize it.
[0152] The retrievable task implicitly acquires a lock for the entry it is processing. This, along with the fact that the data is co-located, means that the entry can be deleted at the end of processing without causing another network operation (instead of letting the entry expire). If the stream is configured to be unordered, then the task can be handed over to the channel's multi-threaded executor as soon as possible. If the stream is ordered, then again, the task can be handed over to the channel, but to a single channel thread. If the stream is partition-ordered, then the handover can occur per partition. The partition can be determined from the key association and then used to index the correct thread for execution. In other words, the threads can be partition-colored.
[0153] Regarding fault tolerance, if a member is down when the sender publishes an event and if the receiver is using a MapListener, then the event is not received when the member comes back up. One way to address this is to use a combination of a MapListener and a continuous query map. In this case, the event can be deleted as soon as it has been fully processed, rather than deleted slowly. If a member receives an event and goes down before finishing processing the event, then the event is reprocessed, which means the event is not deleted from the cache until it has been fully processed.
[0154] If partition data is migrated to a different server, then live events can be listened for regarding whether the partition has been migrated. In some embodiments, this situation is presented as an error to let the user know that the status is missing until the new window has passed. For example, in the middle of a 10-minute window, then the status is only valid in the next window.
[0155] Structural Viewpoint
[0156] In some embodiments, all events passing through the distribution stream are serializable. In terms of configuration, the following can be added to the channel component. Stream type: local or cluster partition, local or cluster broadcast, load balancing, fan-in; Sorting: unordered, partially partition-ordered, fully partition-ordered, partially ordered, fully ordered; Partition attribute: String; Application timestamp: long; Total ordering timeout: long.
[0157] Deployment Viewpoint
[0158] The consistent cache configuration can be included as part of the server deployment / configuration. This can include the configuration of different cache schemes for each distribution stream in the different distribution streams.
[0159] Design-time Considerations
[0160] In some embodiments, within an integrated development environment (IDE), a developer can select different styles of channels from a palette: regular channels, broadcast channels, partition channels, or fan-in channels. The different channel styles in the palette are visual cues for the user. They can be implemented as configurations to allow the user to change the runtime distribution flow without having to author an application in the IDE. In the case of partition channels, the IDE can keep track by prompting the user to use an ordered set of event attributes as the partitioning criteria. The ordered set of event attributes can exist as defined in the event type of the channel. Similarly, any other channel-specific configurations can be configurable accordingly.
[0161] Management, Operations, and Security
[0162] The distributed channel style and any configurations associated with it can be presented in the management console as part of the channels and channel configurations in the event processing network graph view. The management / monitoring console can also provide a mechanism for visualizing the complete network of computing resources of the cluster. This is the deployment view of the cluster.
[0163] In addition, one aspect involves being able to understand the runtime interactions or mappings between a source node in one computing resource to a destination node in another computing resource. This constitutes the connectivity view of the cluster. In this case, not only are the runtime connections shown, but also the number of events that have passed through the connections.
[0164] Another useful monitoring tool provides the ability to safeguard specific events. For example, a user can ensure that the event e1 = {(p1,1),(p2,2)} has passed through processor 2 in machine 3. In some embodiments, a mechanism for guarding events is provided.
[0165] Plotting Based on Event Data
[0166] Some embodiments allow the real-time identification of situations in the streaming data, such as threats and opportunities. These situations can be identified through visualization graphs. Described below are various graphs and plots that can be used to visualize the data. Some of these graphs are not monitoring graphs, but rather exploration graphs. Exploration graphs can be configurable to allow try-and-see pin-pointing of different situations.
[0167] According to some embodiments, a plotting mechanism receives time series data as input. Time series data is data that changes over time. Time series data can be represented by the following types of graphs: line graphs, scatter plots, and radar plots. Bar graphs and pie charts can be used to represent categorical data.
[0168] A line chart is a way to visualize time series data. The X-axis can be used to represent the passage of time and thus allows for the natural scrolling of data as time moves forward. The Y-axis can be used to view how the dependent variable - the variable of interest - changes over time. Figure 11 FIG. is a diagram showing an example of a line chart 1100 according to some embodiments.
[0169] The dependent variable can be any attribute in the attributes of the output event. For example, the dependent variable can be the attribute 'price' of a stock event, or the attribute 'Sum(price)>10' generated by the source application summaries and conditions. Line charts are suitable for continuous variables. In some embodiments, for line charts, it is allowed to select event attributes as numerical values. In some embodiments, it is not allowed to select attributes of types such as Interval, DateTime, Boolean, String, XML, etc. in the Y-axis. The first numerical attribute of the output event can be initially selected as the Y-axis. The user is allowed to change this to any other numerical attribute.
[0170] As mentioned above, the X-axis can specify the time series. This axis can use the system timestamp at which the output event arrives, converted to the HH:MM:SS:milliseconds (hour:minute:second:millisecond) format, and slide using a known slide criterion for evaluating queries and updating the Live Stream table output table. This may be in the range from 1 / 10 second to 1 / 2 second. In some embodiments, optionally, the actual timestamp of the output event or the element time as indicated by CQL can be used. In the case of a query with an applied timestamp, the timestamp represents the applied time, which may be significantly different from the system time. For example, for each hour of actual time, the applied time may advance by 1 tick.
[0171] Another aspect of analyzing streaming data is understanding the correlation between its variables. Correlation shows the covariance of variable pairs and ranges from -1 (inverse strong correlation), 0 (no correlation) to 1 (direct strong correlation). For example, the weight of a car is directly correlated with its miles per gallon (MPG). However, there is no correlation between the weight of a car and its color. To support this correlation, some embodiments allow a second line (in a different color) to be drawn on the line chart. This second line represents a second variable, which can be a second event attribute selected by the user.
[0172] Each line, which consists of a set of its x and y pairs, is called a data series. Plotting two lines allows the user to visualize whether the variable has a direct or an indirect linearity. Additionally, the correlation coefficient of the variable can be calculated and presented in the graph. The correlation coefficient can be calculated using the CQL function correlate(). The correlation coefficient can be presented using a color gradient, where green is for direct correlation, red is for indirect correlation, and gray means no correlation.
[0173] In some embodiments, a mechanism is provided to the user to select additional variables (i.e., event attributes) to be plotted as lines (series) in a line graph, where the additional variables go up to some convenient maximum value (e.g., between 5 and 10). When the user selects an attribute of an output event, he can select a calculated variable, such as the count result of a categorical attribute. However, since correlation is done in a pairwise manner and correlation can be taxing, some embodiments allow only two variables to be correlated at a time. If the graph has more than two variables, the user can notify which two variables should be used to calculate the correlation coefficient.
[0174] In some embodiments, optionally, the confidence of the result can be calculated along with the calculation of the correlation coefficient. This lets the user know how likely it is that a random sample will produce the same result.
[0175] Correlation is associated with variance and covariance. In some embodiments, optionally, the visible features of the graph show whether the distribution represented within the graph is a normal distribution.
[0176] In some embodiments, the top N correlated pairwise variables can be automatically determined.
[0177] When the time dimension is not important, the data can be better understood through the representation in a scatter plot. Figure 12 is a diagram showing an example of a scatter plot 1200 according to some embodiments. In this case, the user can assign different event attributes to both the X-axis and the Y-axis. However, in some embodiments, both are constrained to be numerical variables. By default, in some embodiments, the first two attributes that are numerical can be selected from the output event (type). The X-axis represents the explanatory variable, and the Y-axis represents the response variable. Thus, an attribute indicating the response or a calculated field (such as totalX or sumOfY or outcomeZ) can be a good candidate for automatic assignment to the Y-axis.
[0178] In the case of a time series, new values enter on the right side of the graph and old values exit on the left side. However, this behavior does not translate to a scatter plot because new points can appear anywhere. Thus, in some embodiments, new points are given some visual cue. For example, new points can initially be plotted in red or blue, unlike other existing points, and gradually phase out as new points arrive.
[0179] Before the oldest points in the graph start to be automatically removed, various numbers of points can be maintained in the graph at one time. There are several ways to perform this removal. One technique keeps as many points as possible as long as these points do not make the visualization worse (or cluttered). Another technique keeps the minimum set of events required to understand the data.
[0180] In some embodiments, lines that limit or shape the points in some form can be drawn in the scatter plot. One technique draws a line above all the values in the Y-axis to represent the maximum value, and another line below all the values in the Y-axis to represent the minimum value. Another technique draws a polygon that encompasses all the points. This gives the user a restricted view of where the values are.
[0181] Another technique draws a smoothed curve fitter (i.e., locally weighted scatterplot smoothing (lowess)). Figure 13 is a diagram showing an example of a scatter plot 1300 in which a smoothed curve fitter has been drawn according to some embodiments. This technique can be used for predictive online processing. As in the case of a line graph, the correlation coefficient of the two variables under discussion can be provided. The smoothed line and regression fit can also indicate linearity.
[0182] Scatter plots are well-suited to support the visualization of a third dimension represented as the size of the points. Figure 14 is a diagram showing an example of a scatter plot 1400 in which the points have different sizes according to some embodiments. In this case, the user can assign a third event attribute to the'size' dimension. The mechanism can scale the sizes in a way that avoids cluttering the graph.
[0183] A radar chart is similar to a line graph except that the X-axis is drawn as a circle representing a time period (such as a 24-hour period), making the radar chart look like a radar. Figure 15 is a diagram showing an example of a radar chart 1500 according to some embodiments. Radar charts are useful for discovering whether a particular situation is cyclic and are thus a useful tool for dealing with time series. For example, such a chart can be used to determine whether the price of a stock is typically high at the start or end of a business day, or to determine whether the number of airline tickets sold on Friday is higher than on any other day of the week.
[0184] The response variables of a radar chart can also be numerical. Finding the correct scale for the X-axis is a consideration for a radar chart. If the window range is defined, then it can be used as the default loop for the radar chart. Otherwise, the loop can vary from milliseconds to hours. The lines attached to the chart can represent different response variables. Such as line type Figure 1 Same as above.
[0185] Numerical variables are not the only variables that can be visualized in a chart. Categorical variables (whether they are nominal (e.g., male, female) or ordinal (e.g., high, medium, low)) can also be visualized in a chart. Categorical variables are usually analyzed as frequencies, such as for example the number of frauds (the factor is "fraud", "no fraud"), the ratio of gold customers to regular customers, the top 5 movies watched last week, etc. Such variables are usually visualized in bar charts and pie charts. In some cases (such as the top n), there may be high CPU / memory consumption involved in calculating the frequencies. Therefore, in some embodiments, such calculations are not performed in the background. In some embodiments, the following operations are performed:
[0186] 1. Select a bar chart (or pie chart)
[0187] 2. Assign an event attribute of string type to the X-axis
[0188] 3. Assign the count result (or any other numerical attribute) to the Y-axis
[0189] The count result can be updated according to the user's definition (e.g., using the window range), so there is no need to explicitly clear the values from the bar chart. A pie chart can also be selected. It can be left to the user to ensure that the total across categories is 100%.
[0190] Line charts are suitable for scenarios involving finding general trends in time series and numerical variables. Scatter plots are suitable for scenarios involving finding correlations and numerical variables. Radar charts are suitable for scenarios involving finding loops and numerical variables. Bar charts are suitable for scenarios involving counting frequencies and categorical variables.
[0191] Since certain types of charts can be better suited for specific types of data (e.g., categorical data or numerical data) and analysis, they can also be resized and updated differently. A line chart focusing on a time series contains events for the last t time and can move (update) as the system (CPU) progresses over time. In other words, even if no events arrive, the chart will still update from right to left and move any previously plotted events. By default, a line chart can be configured to contain events for the most recent 300 seconds. However, the user can change the parameter t (e.g., 300) and change the time granularity from milliseconds to minutes. If the time window is defined, the size of the time window can be used as the default size (scale) for the X-axis of the line chart.
[0192] A radar chart is similar to a line chart, and a configuration of periodic intervals is also added in the radar chart. For example, the configuration can indicate to display 300 seconds in a cycle (period) of every 60 seconds. A scatter plot and a bar chart are not time-oriented, and thus in some embodiments, they are updated only when an event arrives, rather than necessarily being updated over time. Power analysis can be used to determine the size of the X-axis as follows: Considering that a scatter plot is used to spot correlations, and considering that generally if there is 80% covariance between variables, the correlation is considered strong, and then using a 95% confidence level (i.e., there is a 5% chance that random points represent important patterns), it gives:
[0193] pwT.r.test(r =.20, power = 0.95, sig.level =.05, alternative = 'greater') = 265.8005
[0194] That is, the scatter plot should include at least 265 points. If this is relaxed to a 75% correlation with a 10% error margin, then:
[0195] pwr.r.test(r =.25, power = 0.90, sig.level =.10, alternative = 'greater') = 103.1175
[0196] Both options are possible. In some embodiments, the user can customize the size to any arbitrary number.
[0197] A bar chart functions similarly to a scatter plot, and in some embodiments, the bar chart is updated only when a new event arrives. The number of categories can be determined using a process. A default value in the range of 10 to 20 can be assumed. The user can further customize according to needs. In some embodiments, the top n categories can be selected. In some embodiments, this can be encoded into the query.
[0198] The table output can also be sized similarly to a scatter plot. One problem that occurs when plotting points with multiple variables is the issue of scale. Specifically, this is very obvious when using a line chart with multiple series. This can be done in two steps:
[0199] 1. Center the data so that it is closer to the average; this brings outliers.
[0200] 2. Normalize the ratio by dividing the data by its standard deviation.
[0201] The formula is:
[0202] x' = (x - mean) / standard - deviation (x'=(x - average value) / standard deviation).
[0203] Due to the streaming nature of the data, the average value and the standard deviation are continuously updated in some embodiments.
[0204] Cluster unsupervised learning
[0205] The cluster groups together events where variables (features) are closer. This grouping allows users to identify events that come together due to some unknown relationship. For example, in the stock market, it is common for clusters of derivatives to move up or down together. If there is a positive revenue report from IBM, then it is likely that there will subsequently be positive results from, for example, Oracle, Microsoft, and SAP.
[0206] One attraction of the cluster is that it is unsupervised; that is, the user does not need to identify a response variable or provide training data. This framework adapts well to streaming data. In some embodiments, the following algorithm is used:
[0207] If each event i contains j variables (e.g., price, volume) designated as xij, and if the goal is to cluster these events into k clusters such that each k - cluster is defined by its centroid ckj (the centroid ckj is a vector of the average values of the j variables for all events that are part of that cluster), then for each event inserted into the processing window, the k - cluster of that event can be determined as follows:
[0208] 1. If it is the first event, assign it to the smallest k - cluster (e.g., cluster 0).
[0209] 2. For each (new) input event i, calculate its (squared) distance to the centroids of the k - clusters as follows:
[0210] 2a. distik = SUMj((xij - ckj)^2)
[0211] 2b. Assign event i to the cluster k with the smallest distik
[0212] 3. Recalculate the centroid for the selected cluster k:
[0213] 3a. For all j, ckj'=(ckj + xij) / |ck| + 1
[0214] For streaming data, there is additional complexity in handling events that leave the processing window:
[0215] 1. For each (old) event i that is removed from cluster k, recalculate the centroid of its cluster:
[0216] 1a. For all j, ckj’ = (ckj - xij) / |ck| - 1
[0217] As the centroid changes for each event (new and old), there is a potential to relocate existing events (points) to the new clusters where the distance has become the minimum distance. Thus, in some embodiments, the process recalculates the distances for all assigned points until no reallocation occurs. However, this can be a burdensome step. If the processing window is small and the pace is fast, removing and adding new points will have the same effect and slowly aggregate to the best local optimum that can be reached. Since this aggregation is not guaranteed, an option to enable / disable recalculation when an event arrives can be provided.
[0218] Another issue is the problem of scaling; since the distance is calculated as the Euclidean distance between features, in some embodiments the features are scaled equally. Otherwise, during distance calculation, a single feature may overwhelm other features.
[0219] The clustering algorithm can be represented as a CQL aggregation function in the following form:
[0220] cluster(max - clusters: int, scale: Boolean): int
[0221] cluster(max - clusters: int, scale: Boolean, key - property: String): List<String, Integer>
[0222] The parameter max - clusters defines the total number of clusters (i.e., k), and it does not change once the query starts. The latter signature enables recalculating the cluster assignment and returns a list of event keys to the cluster assignment.
[0223] In terms of visualization, the clustered data can be overlaid on top of any of the previously defined supported graphs. In other words, if the user selects a scatter plot, then colors and / or shapes associating points with one of the k clusters can be included as part of each point. Figure 19 is a diagram showing an example of the shape 1900 representing clusters overlaid on a scatter plot according to some embodiments. The shape includes the points belonging to the clusters represented by the shape.
[0224] HBASE
[0225] As discussed above, the calculation of TF-IDF values involves accessing historical documents for the "Inverse Document Frequency" calculation, making the TF-IDF streaming scenario a good candidate for use with an indexing layer such as HBase. HBase is suitable for 'Big Data' storage, where the functionality of a Relational Database Management System (RDBMS) is not required. HBase is a type of 'NOSQL' database. HBase does not support Structured Query Language (SQL) as the primary method of accessing data. Instead, HBase provides a JAVA Application Programming Interface (API) to retrieve data.
[0226] Each row in the HBase data repository has a key. All columns in the HBase data repository belong to a specific column family. Each column family includes one or more qualifiers. Therefore, to retrieve the data from the HBase data repository, a combination of the row key, column family, and column qualifier is used. In the HBase data repository, each table has a row key, similar to how each table in a relational database has a row key.
[0227] Figure 1 is a diagram showing an example of Table 100 in the HBase data repository according to some embodiments. Table 100 includes Column 102 - Column 110. Column 102 stores the row key. Column 104 stores the name. Column 106 stores the gender. Column 108 stores the grades for the database course. Column 110 stores the grades for the algorithm course. The column qualifiers of Table 100 represent the names of Column 102 - Column 110: name, gender, database, and algorithm.
[0228] Table 100 relates to two column families 112 and 114. Column family 112 includes Column 104 and Column 106. In this example, column family 112 is named "Basic Information". Column family 114 includes Column 108 and Column 110. In this example, column family 114 is named "Course Grades".
[0229] The concept of HBase column qualifiers is similar to the concept of minor keys in NoSqlDB. For example, in NoSqlDB, the major key of a record could be a person's name. The minor keys could be different pieces of information to be stored for that person. For example, given the primary key " / Bob / Smith", the corresponding minor keys could include "Date of Birth", "City", and "State". For another example, given the primary key " / John / Snow", the corresponding minor keys could similarly include "Date of Birth", "City", and "State".
[0230] The information contained in the HBase data repository is retrieved without using any query language. The goal behind HBase is to store large amounts of data efficiently without performing any complex data retrieval operations. As mentioned above, in HBase, data is retrieved using the JAVA API. The following code snippet gives an idea of how data can be stored and retrieved in HBase:
[0231] HBaseConfiguration config = new HBaseConfiguration();
[0232] batchUpdate.put("myColumnFamily:columnQualifierl", "columnQualifierlvalue!".getBytes());
[0233] Cell cell = table.get("myRow", "myColumnFamily:columnQualifierl");
[0234] String valueStr = new String(cell.getValue());
[0235] HBase can be used to store metadata information for various applications. For example, a company can store customer information associated with various sales in the HBase database. In this case, the HBase database can use an HBase cartridge that enables writing CQL queries that use HBase as an external data source.
[0236] According to some embodiments, a CQL processor is used to enrich source events with context data contained in the HBase data repository. The HBase data repository is referred to in an abstract form.
[0237] According to some embodiments, an event processing network component is created to represent the HBase data repository. The HBase data repository event processing network component is similar to an event processing network
[0238] <store>Event handling network component: id of the event handling network component, repository location (location in the form of domain:port of the HBase database server), event type (schema of the repository as seen by the CQL processor), table name (name of the HBase table).
[0239] According to some embodiments, the event handling network component has an associated <column-mappings>Component to specify the mapping from CQL event attributes to HBase column families / qualifiers. This component is declared in the HBase cartridge configuration file similar to the JAVA Database Connectivity (JDBC) cartridge configuration. This component has the following attributes: name (for which the mapping is to be declared <store>id of the event handling network component), rowkey (row key of the HBase table), cql-attribute (CQL column name used in the CQL query), hbase-family (HBase column family), and hbase-qualifer (HBase column qualifier). According to some embodiments, in the case where the CQL column is a java.util.map, the user only specifies the 'hbase-family'. According to some embodiments, in the case where the CQL column is a primitive data type, the user specifies both the 'hbase-family' and the 'hbase-qualifier'.
[0240] According to some embodiments, <hbase:store>The component is linked to the CQL processor using the 'table-source' element, as shown in the following example:
[0241] <hbase:store id="User" tablename="User" event-ty="UserEvent" store-
[0242] location="localhost:5000" row-key="username">
[0243] < / hbase:store>
[0244] <wlevs:processor id="P1">
[0245] <wlevs:table-source ref="User" / >
[0246] < / wlevs:processor>
[0247] According to some embodiments, for this <hbase:store>The column mapping of the component is specified in the HBase configuration file of the event processor (e.g., OEP) as shown in the following example:
[0248] <hbase:column-mappings>
[0249] <store>User< / store>
[0250] <mapping cql-attribute="ddress" hbase-family="address" / >
[0251] <mapping cql-attribute="firstname" hbase-family="data" hbase-qualifier="firstname" / >
[0252] <mapping cql-attribute="lastname" hbase-family="data" hbase-qualifier="lastname" / >
[0253] <mapping cql-attribute="email" hbase-family="data" hbase-qualifier="email" / >
[0254] <mapping cql-attibute="role" hbase-family="data" hbase-qualifier="role" / >
[0255] < / hbase:column-mappings>
[0256] According to some embodiments, the UserEvent class has the following fields:
[0257] String userName;
[0258] java.util.Map address;
[0259] String first name;
[0260] String lastname;
[0261] String email;
[0262] String role;
[0263] In the above example, the CQL column "address" is a mapping because it will hold all column qualifiers from the 'address' column family. The CQL columns "firstname", "lastname", "email", and "role" hold primitive data types. These are specific column qualifiers from the 'data' column family. The 'userName' field from the event type is the row key and thus it does not have any mapping to an HBase column family or qualifier.
[0264] According to some embodiments, the HBase schema can be dynamic in nature and additional column families and / or column qualifiers can be added at any point in time after the HBase table is created. Thus, an event processor (e.g., OEP) allows a user to retrieve event fields as a mapping that includes all dynamically added column qualifiers. In this case, the user declares a java.util.Map as one of the event fields in the JAVA event type. Thus, the above 'UserEvent' event type has a java.util.Map field named "address". If the cartridge does not support dynamically added column families, then the event type can be modified if the event processing application needs to use the newly added column family.
[0265] According to some embodiments, the HBase database is executed as a cluster. In this scenario, the host name of the master node is provided in the above configuration.
[0266] According to some embodiments, during the configuration of the HBase source, the name of the event type present in the event type repository is received from the user. When "column-mappings" are received from the user, the user interface can provide the column (field) names in that specific event type as cql-column. Thus, incorrect inputs can be eliminated at the user interface level. Details of the available HBase column families and the column qualifiers therein can be provided to the user to select from. Parsing is verified to be performed in the cartridge.
[0267] Customer Sales Example
[0268] The following example identifies large sales and the associated customers. Sales data is obtained in the incoming stream and customer information is obtained from the HBase database.
[0269] <?xml version="1.0" encoding="UTF-8"?>
[0270] <beans xmlns=″http: / / www.springframework.org / schema / beans″
[0271] xmlns:xsi="http: / / www.w3.org / 2001 / XMLSchema-instance″
[0272] xmlns:osgi="http: / / www.springframework.org / schema / osgi″
[0273] xmlns:wlevs=″http: / / www.bea.com / ns / wlevs / spring″
[0274] xmlns:hbase=″http: / / www.oracle.com / ns / ocep / hbase″
[0275] xmlns:hadoop="http: / / www.oracle.com / ns / ocep / hadoop″
[0276] xsi:schemaLocation=″
[0277] http: / / www.springframework.org / schema / beans
[0278] http: / / www.springframework.org / schema / beans / spring-beans.xsd
[0279] http: / / www.springframework.org / schema / osgi
[0280] http: / / www.springframework.org / schema / osgi / spring-osgi.xsd
[0281] http: / / www.bea.com / ns / wlevs / spring
[0282] http: / / www.bea.com / ns / wlevs / spring / spring-wlevs-v11_1_1_6.xsd″>
[0283] <wlevs:event-type-repository>
[0284] <wlevs:event-type type-name="UserEvent">
[0285] <wlevs:class>
[0286] com.bea.wlevs.example.UserEvent
[0287] < / wlevs:class
[0288] < / wlevs:event-type>
[0289] <wlevs:event-type type-name="SalesEvent">
[0290] <wlevs:class>com.bea.wlevs.example.SalesEvent< / wlevs:class>
[0291] < / wlevs:event-type>
[0292] < / wlevs:event-type-repository>
[0293] <!--Assemble event processing network(event processing network)-->
[0294] <wlevs:adapter id="A1" class="com.bea.wlevs.example.SalesAdapter">
[0295] <wlevs:listener ref="S1" / >
[0296] < / wlevs:adapter>
[0297] <wlevs:channel id="S1" event-tyPe="SalesEvent">
[0298] <wlevs:listener ref="P1" / >
[0299] < / wlevs:channel>
[0300] <hbase:store id="Uset" event-type="UsetEvent″
[0301] store-locations="localhost:5000" table-name="User">
[0302] < / hbase:store / >
[0303] <wlevs:processor id="P1">
[0304] <wlevs:table-source ref="User" / >
[0305] < / wlevs:processor>
[0306] <wlevs:channel id="S2" advertise="true" event-type="SalesEvent">
[0307] <wlevs:listener rcf="bean" / >
[0308] <wlevs:source ref="P1" / >
[0309] < / wlevs:channel>
[0310] <!--Create business object-->
[0311] <bean id="bean" class="com.bea.wlevs.example.OutputBean" / >
[0312]
[0313] Specify the following column mappings in the HBase cartridge configuration file:
[0314] <hbase:col umnmappings>
[0315] <name>User< / name>
[0316] <rowkey>userName
[0317] <mapping cql-column=”firstname”hbase-family=”data”
[0318] hbase-qualifier=”firstname” / >
[0319] <mapping cql-column=″lastname”hbase-family=”data”
[0320] hbase-qualifier=”lastname” / >
[0321] <mapping cql-column=”email”hbase-family=”data”
[0322] hbase-qualifier=”email” / >
[0323] <mapping cql-column=”role”hbase-family=”data”
[0324] hbase-qualifier=”role” / >
[0325] <mapping cql-column=”address”hbase-family=”address” / >
[0326] < / hbase:columnmappings>
[0327] The "User" HBase table in the above example has the following schema:
[0328] Row key: username
[0329] Column families: data, address
[0330] Column qualifiers in the 'data' column family: firstname, lastname, email, role
[0331] Column qualifiers in the 'address' column family: country, state, city, street
[0332] The processor runs the following CQL query that joins the input stream with this table:
[0333] select user.firstname,user.lastname,user.email,user.role,user.address.get("city"),price
[0334] from S1[now],User as user
[0335] where S1.username=user.username and price>10000
[0336] Here, the "address" column family is declared as a "java.util.Map" field in the "com.bea.wlevs.example.UserEvent" class. Therefore, use "user.address.get(' <column-qualifer-name>’)” to retrieve the value of a specific column qualifier from this column family.
[0337] HBASE Foundation for the OPENTSBD Monitoring System
[0338] Some embodiments may involve mapping of mappings: row keys, column families, column qualifiers, and multiple versions. A row can contain multiple column families. However, each family is treated together (e.g., compressed / uncompressed). A column family can contain multiple column qualifiers. Column qualifiers can be added or removed dynamically. Each cell has multiple versions, and by default the most recent version is retrieved. The policy associated with the column family determines how many versions to retain and when they are purged.
[0339] Some embodiments are schema - less and type - less. The API can include get, put, delete, scan, and increment. Filtered scans are used to perform queries on non - keys, and filtered scans support a rich filtering language.
[0340] Schema for the OPENTSBD Monitoring System
[0341] Some embodiments may involve a UID table as follows:
[0342] ROW COLUMN+CELL
[0343] \x00\x00\x01 column=name:metrics,vlue=mysq1.bytes_sent
[0344] \x00\x00\x02 column=name:metrics,value=mysql.bytes_received
[0345] mysq1.bytes_received column=id:metrics,value=\x00\x00\x02
[0346] mysql.bytes_sent column=id:metrics,value=\x00\x00\x01
[0347] Some embodiments will involve a metrics table as follows:
[0348] Row (key):
[0349]
[0350] Column + Cell:
[0351]
[0352] Query of the OPENTSBD Monitoring System
[0353] According to some embodiments, from the perspective of CQL, external relationships are mapped to HBase table sources. When configuring the HBase table source, detailed information about which feature of the external relationship is mapped to which column family or which columnfamily.column qualifier in the HBase table can be received from the user.
[0354] According to some embodiments, the UID can be found from the metric name. The CQL query can be as follows:
[0355] select metrics from S_TABLE where UID_TABLE.rowkey = S.uid
[0356] Here, "metrics" is a feature of the external relationship named UID_TABLE, which is mapped to the "metrics" column qualifier in the "name" column family of the HBase table named UID_TABLE. In addition, "rowkey" is another feature of the external relationship that is mapped to the row key of the HBase table.
[0357] According to some embodiments, all metric names starting with cpu can be found using a query such as:
[0358] select metrics from S[now],UID_TABLE where rowkey like"^cpu”
[0359] Here, a string representing a regular expression to be matched is specified. This string can be a feature of stream S, but it doesn't have to be. The HBase API can be used to support such regular expression-based queries. In HBase, rows can be scanned by specifying an inclusive start and an exclusive end.
[0360] According to some embodiments, a similar technique is used, such as:
[0361] SELECT name:metrics FROM UID_TABLE,S WHERE rowkey>=10AND rowkey<
[0362] Some embodiments utilize
[0363] http: / / hbase.apache.org / apidcs / org / apache / hadoop / hbase / filret / RegexStringComparator.html
[0364] According to some embodiments, predicate support capabilities are specified for external relations. If a given predicate falls within the list of supported predicate capabilities, it is executed on the external relation. Otherwise, all data is brought into memory, and then the CQL engine applies the predicate.
[0365] According to some embodiments, measures of service latency metrics filtered by a host can be found. In this case, the CQL is:
[0366] select measures from S[now]as p,metricstable where rowkey=key-encoding(p.serviceLatency-UID,p.ELEMENT_TIME)and host=‘myhost’
[0367] In the above example, "rowkey", "host", and "measures" are columns of the external relation mapped to the HBase table source, and key-encoding is a user-defined function.
[0368] According to some embodiments, host regions with service latency higher than 1000 (milliseconds) can be found. Column qualifiers can be dynamically added to existing rows. For example, a new column qualifier "region" containing the region where the host is deployed can be added. If the metadata for validating the feature "host" is not available, the following method can be used. In hbase:column-mappings, the user can specify:
[0369] mapping cql-attribute=”cI”hbase-family=”cfl” / >
[0370] Here, "cf1" is the column family name, and c1 is of type java.util.Map. The user can access the qualifier in "cf1" as c1.get("qualifier-name"). Thus, the CQL query can be as follows:
[0371] select info.get("region”)from S[now]as p,metrics_table where rowkey=key-encoding(p.serviceLatencyUID)and measures>p.threshold
[0372] Here, "info" is the name of a property of an external relationship, where the external relationship maps to the column family to which the "region" qualifier is dynamically added.
[0373] Measures can have multiple versions. According to some embodiments, if an application timestamp is older than the most recent version, the previous version is obtained. According to some embodiments, the most recent version is used. There may be transaction-oriented use cases such as:
[0374] SELECT product:price FROM PRODUCT_TABLE,TRANSACTION_STREAM[now]AS S
[0375] WHERE row-key=S.productId
[0376] In other words, although the price may have changed, the price as seen at the time the transaction was issued will still honor its application timestamp.
[0377] Hardware Overview
[0378] Figure 16 A simplified diagram of a distributed system 1600 for implementing one of the embodiments is depicted. In the illustrated embodiment, the distributed system 1600 includes one or more client computing devices 1602, 1604, 1606, and 1608, which are configured to execute and operate client applications such as web browsers, proprietary clients (e.g., Oracle Forms), etc. via the network(s) 1610. The server 1612 can be communicatively coupled to the remote client computing devices 1602, 1604, 1606, and 1608 via the network 1610.
[0379] In various embodiments, the server 1612 may be adapted to run one or more services or software applications provided by one or more of the components of the system. In some embodiments, these services may be provided to users of the client computing devices 1602, 1604, 1606, and / or 1608 as web-based or cloud-based services or under a Software as a Service (SaaS) model. Users operating the client computing devices 1602, 1604, 1606, and / or 1608 may then interact with the server 1612 using one or more client applications to utilize the services provided by these components.
[0380] In the configuration depicted in the figure, the software components 1618, 1620, and 1622 of the system 1600 are shown as being implemented on the server 1612. In other embodiments, one or more of the components of the system 1600 and / or the services provided by these components may also be implemented by one or more of the client computing devices 1602, 1604, 1606, and / or 1608. Users operating the client computing devices may then use one or more client applications to use the services provided by these components. These components may be implemented in hardware, firmware, software, or a combination thereof. It should be understood that a variety of different system configurations are possible, which may differ from the distributed system 1600. Thus, the embodiment shown in the figure is an example of a distributed system for implementing an embodiment system and is not intended to be limiting.
[0381] The client computing devices 1602, 1604, 1606, and / or 1608 may be portable handheld devices (e.g., cellular phones, computing tablets, personal digital assistants (PDAs)) or wearable devices (e.g., Google head-mounted displays), which run software such as Microsoft Windows and / or various mobile operating systems such as iOS, Windows Phone, Android, BlackBerry 17, Palm OS, etc., and are enabled for the Internet, email, short message service (SMS), or other communication protocols. The client computing devices may be general-purpose personal computers, such as personal computers and / or laptop computers that include various versions of Microsoft Apple and / or Linux operating systems. The client computing devices may be running various commercially available A workstation computer of any operating system in a UNIX-like or other operating system, including but not limited to various GNU / Linux operating systems such as, for example, Google Chrome OS. Alternatively or additionally, the client computing devices 1602, 1604, 1606, and 1608 can be any other electronic device capable of communicating via the network(s) 1610, such as a thin client computer, an Internet-enabled gaming system (e.g., a Microsoft Xbox gaming console with or without a gesture input device), and / or a personal messaging device.
[0382] Although the exemplary distributed system 1600 is shown as having four client computing devices, any number of client computing devices can be supported. Other devices such as devices with sensors can interact with the server 1612.
[0383] The network(s) 1610 in the distributed system 1600 can be any type of network familiar to those skilled in the art that can support data communication using any of a variety of commercially available protocols, including but not limited to TCP / IP (Transmission Control Protocol / Internet Protocol), SNA (Systems Network Architecture), IPX (Internetwork Packet Exchange), AppleTalk, etc. By way of example only, the network(s) 1610 can be a local area network (LAN), such as a LAN based on Ethernet, Token Ring, etc. The network(s) 1610 can be a wide area network and the Internet. It can include virtual networks, including but not limited to virtual private networks (VPNs), intranets, extranets, public switched telephone networks (PSTNs), infrared networks, wireless networks (e.g., networks operating under any protocol in the Institute of Electrical and Electronics Engineers (IEEE) 802.11 protocol suite, and / or any other wireless protocol), and / or any combination of these networks and / or other networks.
[0384] The server 1612 can include one or more general-purpose computers, dedicated server computers (including, for example, PC (personal computer) servers, servers, midrange servers, mainframe computers, rack-mounted servers, etc.), server farms, server clusters, or any other suitable arrangement and / or combination. In various embodiments, the server 1612 can be adapted to run one or more services or software applications described in the foregoing disclosure. For example, the server 1612 can correspond to the server for performing the processing according to the embodiments of the present disclosure described above.
[0385] Server 1612 may run an operating system, which includes any of the operating systems discussed above, as well as any commercially available server operating systems. Server 1612 may also run any of a variety of additional server applications and / or middleware applications, including an HTTP (Hypertext Transfer Protocol) server, an FTP (File Transfer Protocol) server, a CGI (Common Gateway Interface) server, a database server, and the like. Exemplary database servers include, but are not limited to, those commercially available database servers from Oracle, Microsoft, Sybase, IBM (International Business Machines Corporation), etc.
[0386] In some implementations, server 1612 may include one or more applications that analyze and integrate data feeds and / or event updates received from users of client computing devices 1602, 1604, 1606, and 1608. As an example, the data feeds and / or event updates may include, but are not limited to a (Twitter) feed, a (Facebook) update, or real-time updates and continuous data streams received from one or more third-party information sources, which may include real-time events related to sensor data applications, financial tickers, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, automotive traffic monitoring, etc. Server 1612 may also include one or more applications that display data feeds and / or real-time events via one or more display devices in client computing devices 1602, 1604, 1606, and 1608.
[0387] Distributed system 1600 may also include one or more databases 1614 and 1616. Databases 1614 and 1616 may reside in various locations. As an example, one or more of databases 1614 and 1616 may reside on a non-transitory storage medium local to server 1612 (and / or reside within server 1612). Alternatively, databases 1614 and 1616 may be remote from server 1612 and communicate with server 1612 via a network-based connection or a dedicated connection. In a set of embodiments, databases 1614 and 1616 may reside in a storage area network (SAN). Similarly, any necessary files for performing the functions belonging to server 1612 may be stored locally on server 1612 and / or remotely as needed. In a set of embodiments, databases 1614 and 1616 may include a relational database adapted to store, update, and retrieve data in response to commands in SQL format, such as the database provided by provided.
[0388] Figure 17 is a simplified block diagram of one or more components of a system environment 1700 according to an embodiment of the present disclosure. Through the system environment 1700, services provided by one or more components of an embodiment system can be provided as cloud services. In the illustrated embodiment, the system environment 1700 includes one or more client devices 1704, 1706, and 1708 that can be used by a user to interact with a cloud infrastructure system 1702 that provides cloud services. A client computing device can be configured to operate client applications such as a web browser, a proprietary client application (e.g., Forms), or some other application that can be used by a user of the client computing device to interact with the cloud infrastructure system 1702 to use the services provided by the cloud infrastructure system 1702.
[0389] It should be understood that the cloud infrastructure system 1702 depicted in this figure may have other components in addition to those depicted. Further, the embodiment shown in this figure is only one example of a cloud infrastructure system that may incorporate embodiments of the present invention. In some other embodiments, the cloud infrastructure system 1702 may have more or fewer components than those shown in this figure, may combine two or more components, or may have a different component configuration or arrangement.
[0390] The client devices 1704, 1706, and 1708 may be devices similar to those described above for 1602, 1604, 1606, and 1608.
[0391] Although the exemplary system environment 1700 is shown as having three client computing devices, any number of client computing devices may be supported. Other devices such as devices having sensors may interact with the cloud infrastructure system 1702.
[0392] (One or more) networks 1710 may facilitate communication and data exchange between the client devices 1704, 1706, and 1708 and the cloud infrastructure system 1702. Each network may be any type of network familiar to those skilled in the art that can support data communication using any of a variety of commercially available protocols, including those described above for (one or more) networks 1610.
[0393] The cloud infrastructure system 1702 may include one or more computers and / or servers, which may include those computers and / or servers described above for server 1612.
[0394] In some embodiments, the services provided by a cloud infrastructure system may include many services that can be used by users of the cloud infrastructure system on demand, such as online data storage and backup solutions, web-based email services, hosted office suites and document collaboration services, database processing, managed technical support services, and so on. The services provided by the cloud infrastructure system can be dynamically scaled to meet the needs of its users. The specific instantiation of the services provided by the cloud infrastructure system is referred to herein as a "service instance". Generally, any service available to a user from a cloud service provider system via a communication network such as the Internet is referred to as a "cloud service". Typically, in a public cloud environment, the servers and systems that make up the cloud service provider's system are different from the customer's own internal servers and systems. For example, the cloud service provider's system can host an application, and the user can order and use the application on demand via a communication network such as the Internet.
[0395] In some examples, the services in a computer network cloud infrastructure can include protected computer network access to storage space, a hosted database, a hosted web server, software applications, or other services provided by a cloud vendor to a user or provided in other ways known in the art. For example, the service can include password-protected access to remote storage space on the cloud via the Internet. As another example, the service can include a web service-based hosted relational database and a scripting language middleware engine for private use by networked developers. As another example, the service can include access to an email software application hosted on the cloud vendor's website.
[0396] In certain embodiments, the cloud infrastructure system 1702 can include a set of application, middleware, and database service offerings that are delivered to customers in a self-service, subscription-based, elastically scalable, reliable, highly available, and secure manner. An example of such a cloud infrastructure system is provided by this assignee public cloud.
[0397] In some embodiments, the cloud infrastructure system 1702 may be adapted to automatically provision, manage, and track a customer's subscription to services provided by the cloud infrastructure system 1702. The cloud infrastructure system 1702 may provide cloud services via different deployment models. For example, the services may be provided under a public cloud model, in which the cloud infrastructure system 1702 is owned by an organization that sells cloud services (e.g., owned by Oracle) and makes the services available to the general public and enterprises in different industries. As another example, the services may be provided under a private cloud model, in which the cloud infrastructure system 1702 operates only for a single organization and may provide services to one or more entities within that organization. Cloud services may also be provided under a community cloud model, in which the cloud infrastructure system 1702 and the services provided by the cloud infrastructure system 1702 are shared by several organizations in a related community. Cloud services may also be provided under a hybrid cloud model, which is a combination of two or more different models.
[0398] In some embodiments, the services provided by the cloud infrastructure system 1702 may include one or more services provided under the software as a service (SaaS) category, platform as a service (PaaS) category, infrastructure as a service (IaaS) category, or other service categories including hybrid services. A customer may order one or more services provided by the cloud infrastructure system 1702 via a subscription order. The cloud infrastructure system 1702 then performs processing to provide the services in the customer's subscription order.
[0399] In some embodiments, the services provided by the cloud infrastructure system 1702 may include, but are not limited to, application services, platform services, and infrastructure services. In some examples, the application services may be provided by the cloud infrastructure system via a SaaS platform. The SaaS platform may be configured to provide cloud services falling under the SaaS category. For example, the SaaS platform may provide the ability to build and deliver a set of on-demand applications on an integrated development and deployment platform. The SaaS platform may manage and control the underlying software and infrastructure for providing the SaaS services. By leveraging the services provided by the SaaS platform, a customer may utilize applications executed on the cloud infrastructure system. The customer may obtain the application services without the customer purchasing separate licenses and support. A variety of different SaaS services may be provided. Examples include, but are not limited to, services that provide solutions for sales performance management, enterprise integration, and business agility for large organizations.
[0400] In some embodiments, the platform services can be provided by a cloud infrastructure system via a PaaS platform. The PaaS platform can be configured to provide cloud services that fall under the PaaS category. Examples of platform services can include, but are not limited to, services that enable an organization (such as Oracle) to integrate existing applications on a shared common architecture and to build new applications that utilize the shared services provided by the platform. The PaaS platform can manage and control the underlying software and infrastructure used to provide the PaaS services. Customers can obtain the PaaS services provided by the cloud infrastructure system without the customers having to purchase separate licenses and support. Examples of platform services include, but are not limited to, Oracle Java Cloud Service (JCS), Oracle Data Cloud Service (DBCS), and other services.
[0401] By leveraging the services provided by the PaaS platform, customers can adopt programming languages and tools supported by the cloud infrastructure system and also control the deployed services. In some embodiments, the platform services provided by the cloud infrastructure system can include database cloud services, middleware cloud services (e.g., Oracle Fusion Middleware services), and Java cloud services. In one embodiment, the database cloud services can support a shared service deployment model that enables organizations to pool database resources and provide database as a service to customers in the form of a database cloud. In the cloud infrastructure system, the middleware cloud services can provide a platform for customers to develop and deploy various business applications, and the Java cloud services can provide a platform for customers to deploy Java applications.
[0402] A variety of different infrastructure services can be provided by the IaaS platform in the cloud infrastructure system. The infrastructure services facilitate the management and control of the underlying computing resources (such as storage devices, networks, and other basic computing resources) so that customers can utilize the services provided by the SaaS platform and the PaaS platform.
[0403] In certain embodiments, the cloud infrastructure system 1702 can also include infrastructure resources 1730 for providing resources used to provide various services to the customers of the cloud infrastructure system. In one embodiment, the infrastructure resources 1730 can include a pre-integrated and optimized combination of hardware (such as servers, storage devices, and networking resources) that execute the services provided by the PaaS platform and the SaaS platform.
[0404] In some embodiments, the resources in the cloud infrastructure system 1702 can be shared by multiple users and dynamically reallocated on demand. Additionally, resources can be allocated to users in different time zones. For example, the cloud infrastructure system 1702 can enable a first set of users in a first time zone to utilize the resources of the cloud infrastructure system for a specified number of hours and then enable the reallocation of the same resources to another set of users located in a different time zone, thereby maximizing resource utilization.
[0405] In certain embodiments, a number of internal shared services 1732 can be provided that are shared by different components or modules of the cloud infrastructure system 1702 and by the services provided by the cloud infrastructure system 1702. These internal shared services can include, but are not limited to, security and identity services, integration services, enterprise repository services, enterprise manager services, virus scanning and whitelisting services, high-availability backup and recovery services, services for enabling cloud support, email services, notification services, file transfer services, and the like.
[0406] In certain embodiments, the cloud infrastructure system 1702 can provide comprehensive management of cloud services (e.g., SaaS, PaaS, and IaaS services) in the cloud infrastructure system. In one embodiment, the cloud management functionality can include the ability to provision, manage, and track the subscriptions of customers received by the cloud infrastructure system 1702, among other things.
[0407] In one embodiment, as depicted in the figure, the cloud management functionality can be provided by one or more modules such as an order management module 1720, an order orchestration module 1722, an order provisioning module 1724, an order management and monitoring module 1726, and an identity management module 1728. These modules can include one or more computers and / or servers or be provided using one or more computers and / or servers, which can be general-purpose computers, dedicated server computers, server groups, server clusters, or any other suitable arrangement and / or combination.
[0408] In exemplary operation 1734, a customer using a client device, such as client devices 1704, 1706, or 1708, can interact with the cloud infrastructure system 1702 by requesting one or more services provided by the cloud infrastructure system 1702 and placing an order for a subscription to one or more services provided by the cloud infrastructure system 1702. In some embodiments, the customer can access the cloud user interface (UI), cloud UI 1712, cloud UI 1714, and / or cloud UI 1716 and place a subscription order via these UIs. The order information received by the cloud infrastructure system 1702 in response to the customer placing an order can include information identifying the customer and the one or more services provided by the cloud infrastructure system 1702 that the customer intends to subscribe to.
[0409] After the customer places an order, the order information is received via cloud UIs 1712, 1714, and / or 1716.
[0410] At operation 1736, the order is stored in the order database 1718. The order database 1718 can be one of several databases operated by the cloud infrastructure system 1718 and in conjunction with other system elements.
[0411] At operation 1738, the order information is forwarded to the order management module 1720. In some cases, the order management module 1720 can be configured to perform billing and accounting functions related to the order, such as verifying the order and booking the order when it passes verification.
[0412] At operation 1740, information about the order is transmitted to the order orchestration module 1722. The order orchestration module 1722 can utilize the order information to orchestrate the provisioning of services and resources for the order placed by the customer. In some cases, the order orchestration module 1722 can orchestrate the provisioning of resources to utilize the services of the order provisioning module 1724 to support the subscribed services.
[0413] In some embodiments, the order orchestration module 1722 enables the management of the business processes associated with each order and applies business logic to determine whether the order should proceed with provisioning. At operation 1742, when a new subscribed order is received, the order orchestration module 1722 sends a request to the order provisioning module 1724 to allocate resources and configure those resources required to fulfill the subscription order. The order provisioning module 1724 enables the allocation of resources for the services ordered by the customer. The order provisioning module 1724 provides an abstraction layer between the cloud services provided by the cloud infrastructure system 1702 and the physical implementation layer that supplies the resources used to provide the requested services. The order orchestration module 1722 can thus be isolated from implementation details, such as whether the services and resources are actually provisioned in real time or pre-provisioned and only allocated / designated upon request.
[0414] At operation 1744, once the services and resources are provided, a notification of the provided services can be sent to the customers on client devices 1704, 1706, and / or 1708 via the order provisioning module 1724 of the cloud infrastructure system 1702.
[0415] At operation 1746, the subscription orders of the customers can be managed and tracked by the order management and monitoring module 1726. In some cases, the order management and monitoring module 1726 can be configured to collect usage statistics of the services in the subscription orders, such as the amount of storage used, the amount of data transferred, the number of users, and the amounts of system uptime and system downtime.
[0416] In certain embodiments, the cloud infrastructure system 1702 can include an identity management module 1728. The identity management module 1728 can be configured to provide identity services, such as access management and authorization services in the cloud infrastructure system 1702. In some embodiments, the identity management module 1728 can control information about customers who wish to utilize the services provided by the cloud infrastructure system 1702. Such information can include information for authenticating the identities of such customers and information describing what actions these customers are authorized to perform relative to various system resources (e.g., files, directories, applications, communication ports, memory segments, etc.). The identity management module 1728 can also include management of the descriptive information about each customer and how the descriptive information can be accessed and modified and by whom.
[0417] Figure 18 An example computer system 1800 in which various embodiments of the present invention can be implemented is shown. The system 1800 can be used to implement any one of the above computer systems. As shown, the computer system 1800 includes a processing unit 1804 communicating with a number of peripheral subsystems via a bus subsystem 1802. These peripheral subsystems can include a processing acceleration unit 1806, an I / O subsystem 1808, a storage subsystem 1818, and a communication subsystem 1824. The storage subsystem 1818 includes a tangible computer-readable storage medium 1822 and a system memory 1810.
[0418] The bus subsystem 1802 provides a mechanism for the various components and subsystems of the computer system 1800 to communicate with each other as intended. Although the bus subsystem 1802 is schematically shown as a single bus, alternative embodiments of the bus subsystem may utilize multiple buses. The bus subsystem 1802 can be any of several types of bus architectures, including a memory bus or memory controller, a peripheral bus, and a local bus that utilize any architecture in various bus architectures. For example, these architectures can include Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Enhanced ISA (EISA) bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus, where the PCI bus can be implemented as a Mezzanine bus manufactured according to the IEEE P1386.1 standard.
[0419] The processing unit 1804, which can be implemented as one or more integrated circuits (e.g., conventional microprocessors or microcontrollers), controls the operation of the computer system 1800. One or more processors can be included in the processing unit 1804. These processors can include single-core or multi-core processors. In certain embodiments, the processing unit 1804 can be implemented as one or more independent processing units 1832 and / or 1834, where each processing unit includes a single or multi-core processor. In other embodiments, the processing unit 1804 can also be implemented as a quad-core processing unit formed by integrating two dual-core processors into a single chip.
[0420] In various embodiments, the processing unit 1804 can execute various programs in response to program code and can maintain multiple concurrently executing programs or processes. At any given time, some or all of the program code to be executed can reside in the processing unit(s) 1804 and / or the storage subsystem 1818. Through appropriate programming, the processing unit(s) 1804 can provide the various functions described above. The computer system 1800 can additionally include a processing acceleration unit 1806, which can include a digital signal processor (DSP), a dedicated processor, etc.
[0421] The I / O subsystem 1808 can include user interface input devices and user interface output devices. The user interface input devices can include a keyboard, a pointing device such as a mouse or trackball, a touchpad or touchscreen integrated into a display, a scroll wheel, a trackpoint, a dial, a button, a switch, a keypad, an audio input device with a voice command recognition system, a microphone, and other types of input devices. The user interface input devices can include, for example, motion sensing and / or gesture recognition devices such as Microsoft motion sensors, Microsoft Motion sensors enable a user to control and interact with an input device such as a Microsoft 360 game controller through a natural user interface that utilizes gestures and spoken commands. The user interface input device may also include an eye gesture recognition device, such as one that detects a user's eye activity (e.g., "blinking" when taking a picture and / or making a menu selection) and transforms the eye gesture into an input to the input device (e.g., a Google ) Google blink detector. Additionally, the user interface input device may include a speech recognition sensing device that enables the user to interact with a speech recognition system (e.g., navigator) via voice commands.
[0422] The user interface input device may also include, but is not limited to, a three-dimensional (3D) mouse, joystick or trackpoint, gamepad and graphics tablet, and audio / video devices such as speakers, digital cameras, digital video cameras, portable media players, webcams, image scanners, fingerprint scanners, barcode readers, 3D scanners, 3D printers, laser rangefinders, and eye gaze tracking devices. Additionally, the user interface input device may include, for example, medical imaging input devices such as computed tomography, magnetic resonance imaging, positron emission tomography, medical ultrasound devices. The user interface input device may also include, for example, audio input devices such as MIDI keyboards, digital musical instruments, etc.
[0423] The user interface output device may include a display subsystem, indicator lights, or non-visual displays such as audio output devices. The display subsystem may be a cathode ray tube (CRT), a flat panel device such as one utilizing a liquid crystal display (LCD) or plasma display, a projection device, a touch screen, etc. Generally, the use of the term "output device" is intended to include all possible types of devices and mechanisms for outputting information from the computer system 1800 to a user or other computer. For example, the user interface output device may include, but is not limited to, various display devices that visually convey text, graphics, and audio / video information, such as monitors, printers, speakers, headphones, automotive navigation systems, plotters, voice output devices, and modems.
[0424] The computer system 1800 may include a storage subsystem 1818 that includes software elements shown as currently residing within system memory 1810. The system memory 1810 may store program instructions that are loadable onto and executable by the processing unit 1804, as well as data generated during the execution of these programs.
[0425] Depending on the configuration and type of the computer system 1800, the system memory 1810 can be volatile (such as random access memory (RAM)) and / or non-volatile (such as read-only memory (ROM), flash memory, etc.). RAM typically contains data and / or program modules that can be immediately accessed by the processing unit 1804 and / or are currently being operated on and executed by the processing unit 1804. In some implementations, the system memory 1810 can include various different types of memory, such as static random access memory (SRAM) or dynamic random access memory (DRAM). In some implementations, the basic input / output system (BIOS), which contains basic routines that help transfer information between components within the computer system 1800 during startup, can typically be stored in the ROM. By way of example, and not limitation, the system memory 1810 also shows an application 1812, program data 1814, and an operating system 1816, where the application 1812 can include client applications, web browsers, middle-tier applications, relational database management systems (RDBMS), etc. By way of example, the operating system 1816 can include various versions of Microsoft Apple and / or Linux operating systems, various commercially available or UNIX-like operating systems (including but not limited to various GNU / Linux operating systems, Google OS, etc.) and / or mobile operating systems such as iOS, Phone, OS, 18OS and OS operating systems.
[0426] The storage subsystem 1818 can also provide a tangible computer-readable storage medium for storing the basic programming and data constructs that provide the functionality of some embodiments. Software (programs, code modules, instructions) that provides the above functionality when executed by a processor can be stored in the storage subsystem 1818. These software modules or instructions can be executed by the processing unit 1804. The storage subsystem 1818 can also provide a repository for storing data used in accordance with the present invention.
[0427] The storage subsystem 1818 can also include a computer-readable storage medium reader 1820, and the computer-readable storage medium reader 1820 can be further connected to a computer-readable storage medium 1822. The computer-readable storage medium 1822, together with and optionally in combination with the system memory 1810, can comprehensively represent remote, local, fixed, and / or removable storage devices plus storage media for temporarily and / or more persistently containing, storing, transmitting, and retrieving computer-readable information.
[0428] The computer-readable storage medium 1822 that contains code or portions of code may also include any suitable medium known or used in the art, including storage media and communication media, such as but not limited to volatile and non-volatile, removable and non-removable media implemented in any method or technology for the storage and / or transmission of information. This may include tangible computer-readable storage media such as RAM, ROM, electrically erasable programmable ROM (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical storage devices, magnetic tape cartridges, tapes, disk storage devices or other magnetic storage devices, or other tangible computer-readable media. This may also include non-tangible computer-readable media such as data signals, data transmissions, or any other medium that can be used to transmit the desired information and can be accessed by the computer system 1800.
[0429] As an example, the computer-readable storage medium 1822 may include a hard disk drive that reads from or writes to a non-removable non-volatile magnetic medium, a disk drive that reads from or writes to a removable non-volatile disk, and an optical disk drive that reads from or writes to a removable non-volatile optical disk (such as a CD ROM, DVD, and disk or other optical medium). The computer-readable storage medium 1822 may include but is not limited to drives, flash memory cards, universal serial bus (USB) flash drives, secure digital (SD) cards, DVD disks, digital audio tapes, etc. The computer-readable storage medium 1822 may also include solid-state drives (SSDs) based on non-volatile memory (such as flash memory-based SSDs, enterprise flash drives, solid-state ROMs, etc.), SSDs based on volatile memory (such as solid-state RAM, dynamic RAM, static RAM, DRAM-based SSDs, magnetoresistive RAM (MRAM) SSDs), and hybrid SSDs that use a combination of DRAM-based SSDs and flash memory-based SSDs. Disk drives and their associated computer-readable media may provide non-volatile storage for computer-readable instructions, data structures, program modules, and other data for the computer system 1800.
[0430] The communication subsystem 1824 provides an interface to other computer systems and networks. The communication subsystem 1824 functions as an interface for receiving data from other systems and transmitting data from the computer system 1800 to other systems. For example, the communication subsystem 1824 may enable the computer system 1800 to connect to one or more devices via the Internet. In some embodiments, the communication subsystem 1824 may include radio frequency (RF) transceiver components for accessing wireless voice and / or data networks (e.g., using cellular telephone technologies such as advanced data network technologies like 3G, 4G, or EDGE (Enhanced Data Rates for Global Evolution), Wi-Fi (IEEE 1602.11 standard family), or other mobile communication technologies, or any combination thereof), a Global Positioning System (GPS) receiver component, and / or other components. In some embodiments, as an addition or alternative to the wireless interface, the communication subsystem 1824 may provide a wired network connection (e.g., Ethernet).
[0431] In some embodiments, the communication subsystem 1824 may also receive input communications in the form of structured and / or unstructured data feeds 1826, event streams 1828, event updates 1830, etc. on behalf of one or more users who may use the computer system 1800.
[0432] As an example, the communication subsystem 1824 may be configured to receive data feeds 1826 from users of social networks and / or other communication services in real time, such as feeds, updates, web feeds such as Rich Site Summary (RSS) feeds, and / or real-time updates from one or more third-party information sources.
[0433] In addition, the communication subsystem 1824 may also be configured to receive data in the form of a continuous data stream, which may include an event stream 1828 and / or an event update 1830 of real-time events. The data in the form of a continuous data stream may be continuous or unbounded in nature without a definite end. Examples of applications that generate continuous data may include, for example, sensor data applications, financial tickers, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, automotive traffic monitoring, etc. The communication subsystem 1824 may also be configured to output structured and / or unstructured data feeds 1826, event streams 1828, event updates 1830, etc. to one or more databases, where the one or more databases may communicate with one or more streaming data source computers coupled to the computer system 1800.
[0434] The computer system 1800 may be one of various types, including a handheld portable device (e.g., cellular phone, computing tablets, PDAs), wearable devices (e.g., Google head-mounted displays), PCs, workstations, mainframes, kiosks, server racks, or any other data processing system.
[0435] The computer system 1800 and its components of this embodiment can be configured to perform any appropriate operations based on the principles of the present invention as described in the previous embodiments. For example, the communication subsystem included in such a computer system can receive an event stream. The processor or processing unit included in the computer system can create a table component representing data repository events of a data repository, and use the table component as an external relationship source to implement an engine of an event processing network to select specific data corresponding to a specific event of the event stream from the data repository, and add the specific data to the specific event.
[0436] Preferably, the engine of the event processing network is configured to identify specific data at least partially based on a mapping. Preferably, the mapping can be configured to be updated after the creation of the table component. The mapping can be included as a column family in the table component. The data repository can include a cluster of data storage locations. At least one of the data repositories can include a non-relational database or the engine includes a continuous query language engine.
[0437] It will be apparent to those skilled in the art that for the specific operation processes of the above units, reference can be made to the corresponding steps / components in the related method / system embodiments sharing the same concept, and such reference is also regarded as the disclosure of the related units. Therefore, for the sake of brevity of description, some of the specific operation processes will not be repeated or described in detail.
[0438] Due to the ever-changing nature of computers and networks, the description of the computer system 1800 depicted in the figures is intended only as a specific example. Many other configurations with more or fewer components than the systems depicted in the figures are possible. For example, custom hardware can also be used and / or specific elements can be implemented in hardware (including FPGAs, ASICs, etc.), firmware, software (including applets), or a combination thereof. Additionally, connections to other computing devices such as network input / output devices can be employed. Based on the disclosure and teachings provided herein, those of ordinary skill in the art will understand other ways and / or methods of implementing the various embodiments.
[0439] In the foregoing specification, aspects of the invention have been described with reference to specific embodiments of the invention, but those skilled in the art will recognize that the invention is not limited thereto. The various features and aspects of the above invention may be used alone or in combination. In addition, embodiments may be utilized in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of this specification. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive. < / rowkey> < / hbase:store> < / hbase:store> < / store> < / store> a component and is used as an external relationship source in a CQL processor. The HBase data repository event handling network component is typed using event types. The HBase database is started by its own mechanism and is accessible. The HBase database does not need to be directly managed by an event processor such as an Oracle Event Processor (OEP). According to some embodiments, the HBase data repository event handling network component is provided as a cartridge. The HBase cartridge provides
Claims
1. A method for processing an event stream, comprising: Receiving a plurality of events as the event stream; Creating an event processing network to represent a data repository, the data repository having tables and storing context data, the tables including column families, the column families including one or more columns, and the one or more columns being specified by column qualifiers representing the names of the columns; In the event processing network, creating a table component representing a data repository event of the data repository, the table component having a mapping component for specifying a mapping from event characteristics to column qualifiers, the mapping component including a row key of the table, the name of the column used in a query, a column family, and a column qualifier; And For an event in the event stream, using the table component as an external relational source to operate a processor of the event processing network to perform operations including: Selecting context data corresponding to the event from the data repository based on the mapping; And Adding the selected context data to the event.
2. The method according to claim 1, wherein The context data is retrieved from the data repository by using an application programming interface method call.
3. The method according to claim 2, wherein The application programming interface method call includes a Java method call.
4. The method according to any one of claims 1 to 3, wherein, The data repository includes a non-relational database.
5. The method according to claim 4, wherein, The non-relational database includes an HBase database.
6. The method according to any one of claims 1-3 and 5, wherein, The processor includes a continuous query language processor.
7. The method according to any one of claims 1-3 and 5, further comprising: Holding the event in a cache until the cache is completely processed.
8. A computer-readable storage medium including computer-executable instructions for processing an event stream, the computer-executable instructions when executed by one or more processors configure one or more computer systems to at least perform: Instructions for causing the one or more processors to create a table component representing a data repository event of a data repository, The data repository having tables and storing context data, the tables including column families, the column families including one or more columns, and the one or more columns being specified by column qualifiers representing the names of the columns; and The table component has a mapping component for specifying a mapping from event characteristics to column qualifiers, the mapping component including the row key of the table, the name of the column used in the query, the column family, and the column qualifier; And Instructions for causing the one or more processors to use the table component as an external relational source to implement an engine of an event processing network to perform: Instructions for causing the one or more processors to select context data corresponding to an event in the event stream from the data repository based on the mapping; And Instructions for causing the one or more processors to add the selected context data to the event.
9. The computer-readable storage medium according to claim 8, wherein, The one or more computer systems are further configured to execute instructions for causing the one or more processors to create a mapping component; And Wherein the mapping component is configured to specify a mapping between characteristics of the event and information associated with the context data.
10. The computer-readable storage medium according to claim 9, wherein, The information associated with the context data includes at least one of a qualifier or a family associated with the context data.
11. The computer-readable storage medium according to claim 9 or 10, wherein, The one or more computer systems are also configured to execute instructions that cause the one or more processors to update a mode of the information associated with the context data after creation of the table component.
12. The computer-readable storage medium of claim 11, wherein when implementing the engine of the event processing network, the mode is updated dynamically, and the mode defines at least one of a qualifier type or a family type associated with the context data.
13. The computer-readable storage medium according to any one of claims 8-10 and 12, wherein, The engine includes a continuous query language engine.
14. A system for processing an event stream, comprising: a memory that stores a plurality of instructions; and a processor configured to access the memory, the processor further configured to execute the plurality of instructions to at least: create a table component that represents a data repository event of a data repository, the data repository having tables and storing context data; and implement an engine of an event processing network using the table component as an external relational source to: select context data corresponding to an event in the event stream from the data repository; and add the selected context data to the event, wherein: the engine of the event processing network is configured to identify the context data at least in part based on a mapping, the mapping is configured to be updated after creation of the table component, the mapping is included in the table component as a column family, the mapping holding all column qualifiers from the column family, the column qualifiers representing names of columns included in the table component, the data repository includes a cluster of data storage locations, and there is at least one of the following configurations: the data repository includes a non-relational database, or the engine includes a continuous query language engine.
15. The system according to claim 14, wherein The context data is retrieved from the data repository using an application programming interface method call.
16. The system according to claim 15, wherein, The application programming interface method call includes a Java method call.
17. The system according to claim 16, wherein The non-relational database includes an HBase database.
18. The system according to any one of claims 14 to 17, wherein, The processor includes a continuous query language processor.
19. The system of any one of claims 14 to 17, further comprising: holding the event in a cache until the cache is fully processed.
20. The system according to claim 14, wherein The information associated with the context data includes at least one of a qualifier or a family associated with the context data.