System and method for calculating risk metrics on a network of processing nodes
A network of processing nodes using probability models addresses the limitations of CEP systems by processing both concrete and probabilistic events, ensuring robust and timely threat detection in diverse data environments.
Patent Information
- Application Number
- JP2022569138
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-05-13
- Filing Date
- 2021-05-12
- Publication Date
- 2025-07-10
- Estimated Expiration
- 2041-05-12
AI Technical Summary
Current complex event processing (CEP) systems assume uniformity in input data sources and concreteness of inputs, failing to handle missing or uncertain data effectively.
Implementing a network of processing nodes using probability models (Bayesian networks, Markov chains) to process both concrete and probabilistic events, generating outputs that account for missing or uncertain data through probabilistic calculations.
Enhances security measures by timely and accurately identifying potential threats by handling diverse and uncertain data inputs, improving the reliability of risk metric calculations.
Smart Images

Figure 0007705887000001 
Figure 0007705887000002 
Figure 0007705887000003
Abstract
Description
Related Applications
[0001] (Cross - References to Related Applications) This application claims the benefit of U.S. Provisional Application No. 63 / 024,244, filed on May 13, 2020, the disclosure of which is incorporated herein by reference in its entirety.
Technical Field
[0002] Embodiments of the present disclosure generally relate to event processing. More particularly, embodiments of the present disclosure relate to systems and methods for calculating risk metrics on a network of processing nodes.
Background Art
[0003] The security of governments, companies, organizations, and individuals is becoming increasingly important. This is because such security is being increasingly compromised by some individuals and groups. Therefore, it is important to have security measures that can process useful information in a timely and effective manner when detecting and preventing potential threats, as well as when responding to threats in the development stage.
[0004] With the availability of large amounts of data from several sources such as transaction systems, social networks, web activities, and history logs, it has become necessary to use data technologies to mine and correlate useful and important information. Stream processing techniques and event - based systems incorporating complex event processing (CEP) are widely accepted as solutions for handling big data in several application areas. CEP refers to the detection of events with complex relationships, often including temporal or geographical components. CEP is a way of tracking and analyzing (i.e., processing) a stream of information (data) about things (events) that occur and then drawing conclusions. Generally, the goal of CEP is to identify meaningful events (e.g., opportunities or threats) in real - time situations and respond to them as quickly as possible.
Summary of the Invention
Problems to be Solved by the Invention
[0005] Unfortunately, current CEP systems have drawbacks such as the assumption that input data (events) is obtained from similar data sources, or the assumption that the input data is concrete (no missing inputs or events).
Means for Solving the Problems
[0006] Embodiments of the present disclosure are shown by way of example and not limitation in each of the accompanying drawings, in which like reference numerals indicate like elements.
Brief Description of the Drawings
[0007]
Figure 1A
Figure 1B
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6A
Figure 6B
Figure 6C
Modes for Carrying Out the Invention
[0008] Various embodiments and aspects of the present disclosure are described with reference to the details discussed below, and the accompanying drawings illustrate the various embodiments. The following description and drawings are examples of the present invention and should not be construed as limiting the present invention. Numerous specific details are set forth in order to provide a thorough understanding of the various embodiments of the present disclosure. However, in some instances, well-known details or conventional details are not described in order to provide a concise discussion of the embodiments of the present disclosure.
[0009] References to "one embodiment" in this specification mean that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present disclosure. The appearances of the phrase "in one embodiment" in various places in this specification are not necessarily all referring to the same embodiment.
[0010] Embodiments of the present disclosure provide systems and methods for calculating risk metrics on a network of processing nodes using probability models (e.g., Bayesian networks, Markov chains, etc.) to improve security measures of governments, companies, organizations, individuals, etc. and to address deficiencies in existing event-based systems. Processing nodes include receiving input events from one or more sources and generating one or more events from combinations of input events. Input events can be marked as being of the nature of creation, update, or deletion in the sense that the input event is a new piece of data, a revision to one piece of past data, or a deletion of one piece of past data. Output events can also be marked as creation, update, or deletion in the sense that the output event can be a new piece of derived data, or a revision to one piece of derived data, or an indication that one piece of past derived data is no longer true.
[0011] In one embodiment, the system can process two types of events, such as concrete events and probabilistic events. Concrete events can be received from outside the system or generated by the system when all required inputs are available. Probabilistic events are generated by the system to propagate risk / reliability when not all inputs are available and concrete. Thus, probabilistic events serve to propagate ongoing evaluations that are not yet complete, to indicate areas of risk before they are fully mature, or to compensate for missing information that has not been collected or has been made unclear. When a probabilistic event is propagated through the calculation of a node, the result is a probabilistic event that also compensates for all concrete and probabilistic node inputs.
[0012] One type of computational or processing node is an event - pattern - matching node, and another type is an explicit computational node from input events. The computational node can be an SQL (Structured Query Language) node that performs a streaming query on an input event stream or the use of an ML (Machine Learning) model to synthesize output events from input events.
[0013] According to one aspect, a method for calculating a risk metric on a network of processing nodes is provided. The method includes receiving a plurality of events at a plurality of processing nodes. The method further includes, at a first processing node, processing with respect to a first event and a known instance of a second event to determine whether the first event matches the known instance of the second event. The method further includes, in response to determining that the first event does not match the known instance of the second event, ending the process without generating an output, and generating a first output event having a resulting probability calculated based on a confidence value of the first event and a first probability value of a first missing event. Or, in response to determining that the first event matches the known instance of the second event, generating a first output event having a resulting probability calculated based on the confidence value of the first event. The method further includes, at a second processing node, receiving the first output event as a first input event. The method further includes matching the first input event with a first pattern. The method further includes, when the first input event matches the first pattern, generating a new event having a confidence value calculated based on the confidence value of the first input event. Or, when the first input event does not match the first pattern, generating a new event having a confidence value calculated based on the confidence value of the first input event and a second probability value of a second missing event.
[0014] In one embodiment, the method further includes, at a third processing node, matching a third event with a second pattern. The method further includes, when the third event matches the second pattern, generating a second output event having a resulting probability calculated based on the confidence value of the third event, or when the third event does not match the second pattern, generating a second output event having a resulting probability calculated based on the confidence value of the third event and a third probability value of a third missing event.
[0015] In one embodiment, the method further includes receiving, at a second processing node, a second output event as a second input event. The method further includes matching the second input event to a second pattern. When the first input event matches the first pattern and the second input event matches the second pattern, a confidence value of a new event is calculated based on the confidence value of the first event and the confidence value of the second input event. Or, when the first input event does not match the first pattern or the second input event does not match the second pattern, a confidence value of a new event is calculated based on the confidence value of the first event, the confidence value of the second input event, and a first probability value of a first missing event.
[0016] In one embodiment, processing with respect to a first event and a known instance of a second event includes performing a calculation, a Structured Query Language (SQL) query operation, or a machine learning (ML) model to associate the first event with the known instance of the second event.
[0017] In one embodiment, matching a first input event to a first pattern includes determining whether the first input event matches a plurality of elements of the first pattern, and in response to determining that the first input event matches an element of the first pattern, creating a partial solution for each element that matches the first input event and populating the matching elements for the event type of the first input event.
[0018] In one embodiment, matching a first input event to a first pattern further includes combining matching elements for event types other than the event type of the first input event and filtering the partial solution according to constraints within the first pattern.
[0019] In one embodiment, the method further includes determining that the first input event does not match the first pattern if a partial solution fails a constraint, or determining that the first input event matches the first pattern if each partial solution passes all constraints.
[0020] In one embodiment, the probability and confidence value of a new event are calculated by performing probabilistic calculations on a probability table.
[0021] In one embodiment, the method further includes, at a first processing node, receiving a second event, determining whether the second event matches the first event, and in response to determining that the second event matches the first event, revising, by generating a specific output event, a first output event having a probability calculated based on the confidence value of the first event and a first probability value of the first missing event. The method further includes, at a second processing node, revising the confidence value of the new event based on the confidence value of the specific output event.
[0022] FIGS. 1A and 1B are block diagrams showing an event processing system according to one embodiment. Referring to FIG. 1A, event processing system 100 includes one or more user devices 101-102 communicatively coupled to server 150 via network 103, without limitation. User devices 101-102 can be any type of device such as a host or server, personal computer (e.g., desktop, laptop, and tablet), "thin" client, personal digital assistant (PDA), web-enabled appliance, mobile phone (e.g., smartphone), wearable device (e.g., smartwatch), etc. Network 303 can be any type of wired or wireless network such as a local area network (LAN), wide area network (WAN) such as the Internet, fiber network, storage network, or a combination thereof.
[0023] User devices 101-102 may be provided with an electronic display along with input / output functionality. Alternatively, separate electronic displays and input / output devices, such as a keyboard, that communicate directly with server 150 may be utilized. Any of a variety of electronic displays and input / output devices may be utilized with system 100. In one embodiment, a user may access a web page hosted by server 150 through network 103 using user devices 101-102. Server 150 may supply the web page, and the user may access it by using a conventional web browser or viewer, such as Safari, Internet Explorer, etc., that may be installed on user devices 101-102. Issued events may be presented to the user through the web page. In another embodiment, server 150 may provide a computer application that may be downloaded to user devices 101-102. For example, a user may access a web page hosted by server 150 to download a computer application. The computer application may be installed on user devices 101-102, and user devices 101-102 provide an interface for the user to view issued events.
[0024] Continuing to refer to FIG. 1A, server 150 is communicatively coupled or connected to external system 171 and data storage unit 172, which may be via a network similar to network 103. Server 150 can be any type of server or cluster of servers, such as a web or cloud server, application server, backend server, or a combination thereof. As will be described in more detail below herein, server 150 may rely on one or more external sources of input events (e.g., text messages, social media posts, stock market feeds, traffic information, weather information, or any other type of data) and generate a set of one or more outputs (and further events). In some embodiments, server 150 may process various types of events, such as specific events or probabilistic events. Specific events may be received from an external system (e.g., external system 171) or generated by server 150 when all required inputs are available. In one embodiment, probabilistic events may be generated by server 150 to propagate risk / reliability when not all inputs are available and specific. Thus, probabilistic events may serve to propagate ongoing evaluations that are not yet complete, indicate areas of risk before they are fully mature, or compensate for missing or unclear information that has not been collected. When a probabilistic event is propagated through the calculation of nodes, the result is a probabilistic event that also compensates for all specific and probabilistic node inputs.
[0025] External system 171 can be any computer system having computing capabilities and network connection capabilities for interfacing with server 150. In one embodiment, external system 171 can include a plurality of computer systems. That is, external system 171 can be a cluster of machines that share computing and source data storage workloads. In one embodiment, data storage unit 172 can be any memory storage medium, computer memory, database, or database server suitable for storing electronic data. External system 171 can be a part of a government, company, or organization that respectively implements government functions, company functions, or organizational functions.
[0026] Data storage unit 172 can be a separate computer independent of server 150. Data storage unit 172 can also be a relational data storage unit. In one embodiment, data storage unit 172 can be on external system 171, or alternatively, can be configured to be separately located in one or more locations.
[0027] Referring to FIG. 1B, server 150 includes, without limitation, an input data receiving module 151, a computing node 152, an SQL node 153, an event pattern matching node 154, a machine learning (ML) node 155, and an event pattern matching node 156. Some or all of modules / nodes 151 - 156 can be implemented in software, hardware, or a combination thereof. For example, these modules / nodes can be installed on a persistent storage device 182 (e.g., hard disk, solid state drive), loaded into memory 181, and executed by one or more processors (not shown) of server 150. Note that some or all of these modules / nodes can be communicatively coupled to or integrated with some or all of the modules / nodes of server 150. Some of modules / nodes 151 - 156 can be integrated with each other as integrated modules / nodes.
[0028] In one embodiment, the input data receiving module 151 may receive source data from an external system 171 or a data storage unit 172. For example, when data is directly supplied to the server 150 when it becomes available (i.e., direct feed), the source data may be pushed. Alternatively, when the server 150 (i.e., module 151) periodically requests source data from an external system 171 through a database query such as Solr, Structured Query Language (SQL), etc., the source data may be pulled. The source data may include input events (e.g., text messages, social media posts, stock market feeds, traffic information, weather information, or other types of data) and may take any form of structured and unstructured data (e.g., non-uniform content). Any data such as Extensible Markup Language (XML), Comma-Separated Value (CSV), Java Script Notation (JSON), and Resource Description Framework (RDF) data may be used as structured data. In one embodiment, the source data may be source data that can be mapped (or filtered using vocabulary data) with a vocabulary from an external source, as described in U.S. Patent No. 9,858,260, entitled "System and method for analyzing items using lexicon analysis and filtering process," the disclosure of which is incorporated herein by reference. Upon receiving the source data, the module 151 may store the source data as input source data 161 on the persistent storage device 182.
[0029] Computing node The computing node 152 can receive source data or events (input source data 161) from one or more sources (e.g., external system 171) and compute results. The computing operations can include the joining of streams (data or event streams) and logical operations (e.g., functions, operators, window operations, etc.). Each time an event is received, the result of the node can be updated. Input events may not generate new outputs if the events are propagated through the computation but do not produce different outputs. In such computing nodes, each operation is linked to other operations by explicit input / output associations, so that each source feeds an operation, and then the output of the operation is input to another operation until the final output operation is reached. As with other nodes, probabilistic outputs can be obtained if the output is not generated from a set of data. In either case, a probabilistic model (e.g., a Bayesian network or another graphical or statistical model) or an ML algorithm / model, and the input event reliability are used to compute the reliability results.
[0030] SQL Node The SQL node 153 can be similar to the computing node 152, but uses SQL query operations to join multiple streams into a single output stream. In the case of the SQL node 153, a result is either generated as a whole or not. In the case where a result is generated, probabilistic calculations (e.g., Bayesian calculations, or another graphical or statistical calculation) or an ML algorithm / model can be used to compute the output reliability using the input event reliability. In the case where no output is obtained, probabilistic calculations can also be used to compute the output reliability, but the probabilistic calculations produce probabilistic events as outputs. In some embodiments of SQL processing, it is not possible to detect the absence of an output, and in that case, the output event can only be enhanced with a risk reliability value based on its input events.
[0031] ML Model Node The ML model node 155 may be functionally equivalent to the SQL node 153, but the ML model / algorithm generates, or does not generate, an output. In one embodiment, probabilistic calculations (e.g., Bayesian calculations, or another graphical or statistical calculation) or an ML algorithm / model may be used, similar to assigning a confidence level to the output or generating a probabilistic output event. In some embodiments of ML processing, it may not be possible to detect the absence of an output, and in that case, the output event may only be enhanced with a risk confidence value based on its input event.
[0032] Event Pattern Matching Node Each of the event pattern matching nodes 154, 156 may include a set of elements that collect events and a set of constraints that are satisfied between events. A constraint refers to the specification of a restriction on an event. When a pattern is matched, a new event may be issued. In the case of a match to the pattern, a new event is issued and has a confidence level based on the confidence level of the input event. For example, in the case where an event is not generated due to a partial match of an event to a pattern, a placeholder event (sometimes called fi) is generated with a probabilistic confidence value based on the available input and any probability values for missing input. In either case, a probabilistic model (e.g., a Bayesian network or another graphical or statistical model) or an ML algorithm / model may be used to calculate the confidence level of the output event.
[0033] Risk Model Propagation Embodiments of the present disclosure are used in situations where a set of nodes is used to detect / predict the interpretation of events, and the calculated confidence is the expected level of risk that the interpretation or prediction is correct, assuming a known event. In other words, if a node consumes (matches) two of three events, for example, the question may be asked, "What is the risk that the third event occurred or occurred undetected?" Thus, the pattern correctly interprets / interpreted the data. For example, if the pattern is looking for money laundering, the question may be asked, "Given the data at hand, did it occur or is it about to occur?"
[0034] Nodes form a hierarchy, and low-level nodes generate probabilistic events or concrete events, respectively, when they partially match or are complete. Such events are used when calculating upstream or higher-level abstraction nodes. Server 150 ensures that a risk value is generated from each node and is used when calculating the risk level for any pattern that is an input thereto.
[0035] FIG. 2 is a block diagram showing an exemplary network of processing nodes according to one embodiment. In FIG. 2, risks are calculated across node network 200. For example, at 201, there are two events 210-211 each having event types A and B and input to calculation node 152 (also referred to as calculation node H). Event types A and B can be the same event type or different event types. At 202, there are three events 212-214 each having event types C, D, and E and input into event pattern matching node 154 (also referred to as pattern node J). Event types C, D, and E can be the same event type or different event types. At 203, there are two events 215-216 each having event types F and G and input into ML model node 155 (also referred to as ML model node K). Event types F and G can be the same event type or different event types. The outputs from nodes 152-155 are supplied to event pattern matching node 156 (also referred to as pattern node L), and event pattern matching node 156 outputs an event 217 having event type M (at 207).
[0036] In one embodiment, risk analysis is made available on all nodes, and the nodes can have various assigned probability tables (as described in more detail below). The following is, for example, a possible sequence of events and resulting actions by the system.
[0037] Continuing to refer to FIG. 2, in one embodiment, event 210 arrives and is transferred to computing node H. Generally, an event is an instance of an event type and has a structure defined by that type. Computing node H may attempt to associate event 210 with a known instance of event 211. If there is no known instance of event 211, computing node H ends without output, and a probability (e.g., a probability calculated using Bayesian probability, another graphical / statistical model, or an ML algorithm / model) is calculated assuming the confidence level of event 210 and the probability table associated with computing node H. As a result, a probabilistic event is input to pattern node L, and pattern node L is used to construct a probabilistic event 217 of type M.
[0038] In one embodiment, event 215 arrives and is transferred to ML model node K. If ML model node K generates an output, the output is transferred to pattern node L. For example, if event 215 and event 210 are associated with the same result event, a previous or prior instance of probabilistic event 217 can be revised with a new probability. Otherwise, a new instance of event 217 is generated based on the confidence level from ML model node K. For example, the confidence level from ML model node K can be calculated based on the confidence level of event 215 and the probability table associated with ML model node K.
[0039] If event 211 arrives and event 211 matches the previous event 210, computing node H revises the previously generated probabilistic event by generating a specific output event for pattern node L. Assuming that the pattern within pattern node L requires all inputs to pass, the previous probabilistic output can be revised based on the new confidence level from computing node H.
[0040] Pattern node J includes a set of elements that collect events 212 - 214 and a set of constraints that are satisfied during events 212 - 214. When a pattern is matched, a new event can be issued or output to pattern node L. In the case of a pattern match, a new event is issued and has a confidence based on the confidence of input events 212 - 214. For example, in the case where an event is not generated due to a partial match of an event to a pattern, a placeholder event is generated with a probabilistic confidence value based on the available inputs and any probability values for missing inputs. In either case, a probability model (e.g., a Bayesian network or another graphical or statistical model) or an ML algorithm / model can be used to calculate the confidence of the output event.
[0041] This continues until pattern node L receives sufficient input to create a match, at which point pattern node L revises the previously generated probabilistic output by generating a concrete output with the confidence of its input and the confidence calculated from the probability table for pattern node L. That is, the matching of the number of input events is non - limiting and such matching can be repeated for as many input events as required, including revisions of previous input events or instances of previous input events. Node network 200 of FIG. 2 shows four processing nodes (e.g., nodes 152, 154, 155, and 156), but note that the number of processing nodes is non - limiting. In other words, any number and any combination of processing nodes (e.g., pattern matching, SQL, ML models, calculations) can exist in one embodiment.
[0042] FIG. 3 is a flowchart showing an exemplary method of calculating a risk metric on a network of processing nodes according to one embodiment. Process 300 can be implemented by processing logic that can include software, hardware, or a combination thereof. For example, process 300 can be implemented by server 150, e.g., any combination of nodes 152 - 156 of FIG. 1B.
[0043] Referring to FIG. 3, at 301, a plurality of events are received at a plurality of processing nodes (e.g., computing nodes, event - pattern - matching nodes, SQL nodes, ML model nodes, etc.). At 302, at the first processing node (e.g., a computing node), the first event and a known instance of the second event are processed, and it is determined whether the first event matches the known instance of the second event. At 303, in response to determining that the first event does not match the known instance of the second event, the process ends without generating an output, and a first output event is generated. The first output event has a resulting probability calculated based on the confidence value of the first event and the first probability value of the first missing event. Or, in response to determining that the first event matches the known instance of the second event, a first output event is generated. The first output event has a resulting probability calculated based on the confidence value of the first event. At 304, at the second processing node (e.g., an event - pattern - matching node), the first output event is received as the first input event. At 305, the first input event is matched with the first pattern. At 306, if the first input event matches the first pattern, a new event is generated. The new event has a confidence value calculated based on the confidence value of the first input event. Or, if the first input event does not match the first pattern, a new event is generated. The new event has a confidence value calculated based on the confidence value of the first input event and the second probability value of the second missing event.
[0044] Streaming Considerations One embodiment is in the context of a streaming computing system, in which events are continuously processed and the risk values and the events issued are continuously updated. This adds additional issues such as the frequency of the events issued and the use of streaming semantics to ensure accurate full and partial risk propagation and calculation. Independent time windows may be used to control the amount between probabilistic events and concrete events and thus control the overhead of risk calculation and propagation time.
[0045] FIG. 4 is a flowchart illustrating an exemplary method implemented by an event pattern matching node according to one embodiment. Process 400 may be implemented by processing logic that may include software, hardware, or a combination thereof. For example, process 400 may be implemented by server 150, such as nodes 154, 156 of FIG. 1B. Process 400 may be repeated for each event pattern matching node for which risk assessment is desired and configured.
[0046] Referring to FIG. 4, at 401, input events are received for processing against the pattern. At 402, a partial solution is created with the supplied or received events, and the appropriate matching elements for the event type of the received events are filled in. If an event matches multiple elements, a partial solution that fills that element is created for each element that the event matches.
[0047] In 403, the process performs joins for other event types and filters partial solutions specified by the constraints in the pattern. If the constraint limits the values of two elements and the current solution has not yet joined with such elements, the join is performed. If such elements are not joined, the solution is filtered against the constraint condition. If, during the process of performing the join or filter operation, the solution is found to fail the constraint, it is transferred (at 404) to be treated as a probabilistic solution. The probabilistic solution proceeds through the remaining constraints and, when possible, collects specific event data when it matches the solution data. As a result, a solution with all the necessary elements populated by events, or a probabilistic solution with partially filled elements, is obtained. If the solution passes all the constraints, the solution must further match all multiplicity constraints (e.g., the minimum number of events for each element, the maximum number of events for each element). If the solution passes all the constraints and matches all the multiplicity constraints, the solution (result event) is given to 405.
[0048] If the solution cannot match the multiplicity constraint or includes any probabilistic events for the input, the solution is treated as not matching the pattern and thus as a probabilistic solution (proceed to 406). In a streaming system, a key extraction operation that extracts the key used in the join precedes the join from the event. If such an operation is required and a value is missing from the event or the event is missing from the solution, a probabilistic solution is created and transferred to the next join / filter operation. This enables each join to have a chance of being complete even if some data in the solution is missing.
[0049] In some embodiments, considerations are given to how the joins are performed and the order in which the joins are performed. In a streaming environment, joins generally involve maintaining state information on both sides of the join. In cases where one side is a more limited data set, performance can be obtained by distributing the data set to all nodes so that the join can be performed without additional state on the opposite side. Another optimization is the ordering of the constraint processing such that subsequent joins with shared values are performed. For example, if all constraints involve equality against other events on fields of an event, the constraints can be processed sequentially as joins without the need to shuffle data between threads / cores.
[0050] When update and delete events are received, the update and delete events are processed such that the solution is updated accordingly and may generate an output event if no output event has been generated previously, or may not generate an output event if it has been done so previously. In cases where previous results have been generated and are no longer generated, old events are deleted (the sent delete event), and probabilistic events are generated.
[0051] At 405, a fully populated solution uses a confidence value (e.g., a confidence value from input events or calculated by a prior pattern - risk calculation) to calculate a confidence value for the solution being processed. This is performed using a probabilistic calculation (e.g., Bayesian calculation) that treats the input confidence as a value for an independent variable and a composite probability table for calculating the resulting probability (used as the risk confidence for output events from the solution).
[0052] For example, referring to FIG. 6A, in a pattern having three inputs A, B, C and a solution, each with events a, b, c respectively, and in a probability table 600 (e.g., a Bayesian probability table), the calculation of the output reliability takes the probability table 600, substitutes the reliability from a, b, c in all cells for the existing events and the reciprocal of the reliability for all cells where the event is absent, then takes the product of each input reliability including the row probability, and sums all rows to obtain the output reliability. Thus, when a = 0.75, b = 1.0, and c = 0.9, table 620 (illustrated in FIG. 6B) is used to calculate the product of each row and the sum of the results to obtain the output reliability. The output reliability is then the sum of the last column of table 620 and is equal to 0.475725.
[0053] In 406, when a partially populated / probabilistic solution is available (e.g., available from input events or calculated by a prior pattern - risk calculation), the reliability value is used to calculate the reliability value attributed to a probabilistic output event. The calculation treats the input reliabilities as the values of independent variables and is performed using a probabilistic calculation (e.g., Bayesian calculation) that uses a formula (e.g., Bayes' formula) to eliminate missing variables and a composite probability table for calculating the probabilities (used as the risk reliability for the probabilistic output events generated from the solution).
[0054] For example, in a pattern with three inputs A, B, C and a solution involving only events a and b, and in probability table 600 (shown in FIG. 6A), the calculation of the output reliability takes the probability table 600, substitutes the reliability from a and b in each row for the existing events, and the reciprocal of the reliability in the cells where A or B is missing, takes the product of each row including the row probability, and sums them up to obtain the output reliability. For example, when a has a reliability of 0.2 and b has a reliability of 0.4, table 640 (shown in FIG. 6C) is obtained from table 600, table 640 has substitutions for a, b, and their reciprocals, is then multiplied across the rows, and then summed downwards. Note that the values of tables 600, 620, 640 are given only as examples, are non-limiting, and that tables 600, 620, 640 can contain any values.
[0055] In 407, a fully populated solution is used to generate or issue an output or result event according to the pattern specification, and the calculated reliability is attached to the issued result event.
[0056] In 408, a partially populated solution is used to generate or issue a probabilistic output event according to the pattern specification, and the calculated reliability is attached to the issued output event.
[0057] In 407 and 408, when the pattern specifies an output window (i.e., time frame) for the pattern, this can be applied separately to the complete event and the probabilistic event. In either case, the output window specification can limit the rate at which events are issued per solution. This allows the amount of output events to be limited separately from the rate of input event arrivals. By doing so separately, it is possible to generate actual output events with a lower latency than probabilistic events and limit the overhead of calculating risk values in high-throughput or low-latency application examples.
[0058] FIG. 5 is a flowchart showing an exemplary method implemented by a non-pattern node according to an embodiment. Process 500 may be implemented by processing logic that may include software, hardware, or a combination thereof. For example, process 400 may be implemented by server 150, such as nodes 152, 153, 155 of FIG. 1B.
[0059] Referring to FIG. 5, at 501, an event is received. At 502, a calculation (e.g., an instructed calculation, SQL, or ML) is performed and an output is generated. Unlike process 400 and pattern-based processing of FIG. 4, the results or lack of results in the cases of calculation, SQL, and ML are all or zero, and thus apply only in those two cases and there is no related partial processing.
[0060] Accordingly, at 503, if the calculation generates an output (or result), the output proceeds to 504 and then to 506. Otherwise, the calculation does not generate an output (or result) and the process proceeds to 505 and then to 507. The aspects at 504-507 are the same as or identical to the aspects at 405-408 of FIG. 4 (described above). Accordingly, for the sake of brevity, the aspects at 504-507 are not described herein.
[0061] Note that some or all of the components illustrated and described above may be implemented as software, hardware, or a combination thereof. For example, such components may be installed and stored as software within a persistent memory device and the software may be loaded and executed in memory by a processor (not shown) to perform the processes or operations described throughout this application. Alternatively, such components may be programmed or implemented as executable code incorporated within dedicated hardware such as an integrated circuit (e.g., an application specific IC or ASIC), a digital signal processor (DSP), a field programmable gate array (FPGA), etc., and the executable code may be accessed from an application via a corresponding driver and / or operating system. Further, such components may be implemented as specific hardware logic within a processor or processor core as part of an instruction set accessible by software components via one or more specific instructions.
[0062] Some portions of the foregoing detailed description have been presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. An algorithm is here, and generally, conceived to be a self-consistent sequence of operations leading to a desired result. The operations are those requiring physical manipulations of physical quantities.
[0063] However, it should be noted that all of these terms and similar terms should be associated with appropriate physical quantities and are merely convenient symbols applied to these quantities. Unless otherwise specified, as is apparent from the above discussion, throughout the description, the discussion of using terms as described in the following claims refers to the operation and process of a computer system or a similar electronic computing device that manipulates data represented as physical (electronic) quantities in the registers and memories of the computer system and converts it into other data similarly represented as physical quantities in the computer system memory or registers or other such information storage, transmission device, or display device.
[0064] Embodiments of the present disclosure also relate to an apparatus for performing the operations herein. Such a computer program is stored in a non-transitory computer-readable medium. A machine-readable medium includes any mechanism for storing information in a form readable by a machine (e.g., a computer). For example, a machine-readable (e.g., computer-readable) medium includes a machine (e.g., a computer) readable storage medium (e.g., read-only memory ("ROM"), random access memory ("RAM"), magnetic disk storage media, optical storage media, flash memory devices).
[0065] The processes or methods shown in the preceding figures can be implemented by processing logic including hardware (e.g., circuits, dedicated logic, etc.), software (e.g., implemented on a non-transitory computer-readable medium), or a combination thereof. Although the processes or methods are described above by a number of sequential operations, it should be understood that some of the described operations can be performed in a different order. Further, some operations can be performed in parallel rather than sequentially.
[0066] Embodiments of the present disclosure are not described with reference to any particular programming language. It will be understood that various programming languages may be used to implement the teachings of the embodiments of the present disclosure described herein.
[0067] In the foregoing specification, embodiments of the present disclosure have been described with reference to its particular exemplary embodiments. It will be apparent that various changes may be made thereto without departing from the broader spirit and scope of the present disclosure as set forth in the following claims. Accordingly, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense.
Claims
1. A computer-implemented method for processing a data stream including events, comprising: receiving, by a server, a data stream including a plurality of events at a plurality of processing nodes; at a first processing node, processing, by the server, with respect to a first event and a known instance of a second event to determine whether the first event matches the known instance of the second event; in response to determining that the first event does not match the known instance of the second event, ending, by the server, the processing without generating an output, and generating, by the server, a probabilistic event having a probability calculated based on a confidence value of the first event and a first probability value of a first missing event, or in response to determining that the first event matches the known instance of the second event, generating, by the server, a first output event having a probability calculated based on the confidence value of the first event; at a second processing node, receiving, by the server, the first output event or the probabilistic event as a first input event; matching, by the server, the first input event with a first pattern; when the first input event matches the first pattern, generating, by the server, a new event having a confidence value calculated based on the confidence value of the first input event, or when the first input event does not match or partially matches the first pattern, generating, by the server, a first placeholder event having a confidence value calculated based on the confidence value of the first input event and a second probability value of a second missing event; A computer-implemented method comprising the above steps.
2. at a third processing node, matching, by the server, a third event with a second pattern; when the third event matches the second pattern, generating, by the server, a second output event having a probability calculated based on the confidence value of the third event, or When the third event does not match or partially matches the second pattern, generating, by the server, a second placeholder event having a resulting probability calculated based on the confidence value of the third event and the third probability value of the third missing event The method according to claim 1, further comprising **Claim 3** at the second processing node receiving, by the server, the second output event as a second input event matching, by the server, the second input event with a second pattern wherein when the first input event matches the first pattern and the second input event matches the second pattern, the confidence value of the new event is calculated based on the confidence value of the first event and the confidence value of the second input event, or when the first input event does not match or partially matches the first pattern, or the second input event does not match or partially matches the second pattern, the confidence value of the new event is calculated based on the confidence value of the first event, the confidence value of the second input event, and the first probability value of the first missing event, the matching The method according to claim 2, further comprising **Claim 4** Processing with respect to the first event and the known instances of the second event includes the server performing a calculation, a structured query language (SQL) query operation, or a machine learning (ML) model to associate the first event with the known instances of the second event. The method according to claim 1 **Claim 5** Matching the first input event with the first pattern includes determining, by the server, whether the first input event matches a plurality of elements of the first pattern in response to determining that the first input event matches the elements of the first pattern, creating, by the server, a partial solution for each element that matches the first input event and populating the matching elements for the event type of the first input event The method according to claim 1, comprising **Claim 6** Matching the first input event with the first pattern includes combining, by the server, the matching elements for event types other than the event type of the first input event; filtering, by the server, the partial solution according to the constraints within the first pattern; The method according to claim 5, further comprising. **Claim 7** when the partial solution fails to meet the constraints, determining by the server that the first input event does not match the first pattern, or when each partial solution meets all the constraints, determining by the server that the first input event matches the first pattern; The method according to claim 6, further comprising. **Claim 8** The method according to claim 1, wherein the obtained probability and the confidence value of the new event are calculated by performing probabilistic calculations regarding a probability table. **Claim 9** at the first processing node, receiving, by the server, the second event; determining, by the server, whether the second event matches the first event; in response to determining that the second event matches the first event, revising, by the server, the first output event having the obtained probability, which is calculated based on the confidence value of the first event and the first probability value of the first missing event, by generating a specific output event; at the second processing node, revising, by the server, the confidence value of the new event based on the confidence value of the specific output event; The method according to claim 1, further comprising. **Claim 10** The method according to claim 1, wherein each of the first output event and the new event is a specific event. **Claim 11** the first processing node is a computing node, a Structured Query Language (SQL), or a machine learning (ML) model node, the second and third processing nodes are event pattern matching nodes; **Claim 12** a processor; a memory coupled to the processor to store instructions that, when executed by the processor, cause the processor to perform operations, the operations including receiving a data stream including a plurality of events at a plurality of processing nodes; at the first processing node, Process with respect to a first event and a known instance of a second event to determine whether the first event matches the known instance of the second event, In response to determining that the first event does not match the known instance of the second event, end the process without generating an output, and generate a probabilistic event having a resulting probability calculated based on the confidence value of the first event and a first probability value of a first missing event, or In response to determining that the first event matches the known instance of the second event, generate a first output event having a resulting probability calculated based on the confidence value of the first event, At a second processing node, Receive the first output event or the probabilistic event as a first input event, Match the first input event with a first pattern, If the first input event matches the first pattern, generate a new event having a confidence value calculated based on the confidence value of the first input event, or If the first input event does not match or partially matches the first pattern, generate a first placeholder event having a confidence value calculated based on the confidence value of the first input event and a second probability value of a second missing event including a memory A server system comprising.
13. The operation is, At a third processing node, Match a third event with a second pattern, If the third event matches the second pattern, generate a second output event having a resulting probability calculated based on the confidence value of the third event, or If the third event does not match or partially matches the second pattern, generate a second placeholder event having a resulting probability calculated based on the confidence value of the third event and a third probability value of a third missing event The server system according to claim 12, further comprising.
14. The operation is, At the second processing node, Receive the second output event as a second input event, matching the second input event with the second pattern, when the first input event matches the first pattern and the second input event matches the second pattern, calculating the confidence value of the new event based on the confidence value of the first event and the confidence value of the second input event, or when the first input event does not match the first pattern, or partially matches, or the second input event does not match the second pattern, or partially matches, calculating the confidence value of the new event based on the confidence value of the first event, the confidence value of the second input event, and the first probability value of the first missing event, and matching The server system according to claim 13, further comprising.
15. Processing with respect to the first event and the known instance of the second event includes performing a calculation, a Structured Query Language (SQL) query operation, or a machine learning (ML) model to combine the first event with the known instance of the second event. The server system according to claim 12.
16. Matching the first input event with the first pattern includes determining whether the first input event matches a plurality of elements of the first pattern, and in response to determining that the first input event matches the element of the first pattern, creating a partial solution for each element that matches the first input event and filling the matching element for the event type of the first input event. The server system according to claim 12, comprising.
17. Matching the first input event with the first pattern includes combining the matching elements for event types other than the event type of the first input event, and filtering the partial solution according to the constraints within the first pattern. The server system according to claim 16, further comprising.
18. The operation includes when the partial solution fails to meet the constraints, determining that the first input event does not match the first pattern, or If all the partial decompositions pass all the constraints, determine that the first input event matches the first pattern The server system according to claim 17, further comprising
19. The server system according to claim 12, wherein the obtained probability and the confidence value of the new event are calculated by performing probabilistic calculations regarding a probability table.
20. The operation is at the first processing node, receiving the second event; determining whether the second event matches the first event; in response to determining that the second event matches the first event, revising, by generating a specific output event, the first output event having the obtained probability calculated based on the confidence value of the first event and the first probability value of the first missing event; at the second processing node, revising the confidence value of the new event based on the confidence value of the specific output event The server system according to claim 12, further comprising
21. The server system according to claim 12, wherein each of the first output event and the new event is a specific event.
22. The first processing node is a computing node, a Structured Query Language (SQL), or a machine learning (ML) model node, The server system according to claim 13, wherein the second and third processing nodes are event pattern matching nodes.
Citation Information
Patent Citations
System and method for detecting pattern of events
JP2020017254A
Method and system for complex event processing with latency constraints
US20170195204A1