Analysis device and analysis method
By employing Concurrent Context-free Grammar and an extended left-corner branching method, the limitations of conventional process mining methods are overcome, allowing for the analysis of business processes with parallel structures and expanding the application area of process mining.
Patent Information
- Application Number
- PCT/JP2023/043901
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-07
- Publication Date
- 2025-06-12
AI Technical Summary
Conventional process mining methods, such as those based on Probabilistic Generative Process Models (PGPM), struggle to represent and analyze business processes with parallel structures and indefinite numbers of parallel executions, limiting the application area of process mining.
The introduction of a Concurrent Context-free Grammar (CCFG) and an extended left-corner branching method allows for the representation and analysis of process models with parallel structures, enabling compliance checking with traces that include Parallel Multiple Instance (PMI) sub-processes.
This approach expands the application area of process mining by enabling the analysis of business processes with complex parallel structures, improving compliance checking and data analysis capabilities.
Smart Images

Figure JP2023043901_12062025_PF_FP_ABST
Abstract
Description
Analysis device and analysis method
[0001] The present disclosure relates to analysis of business logs for business improvement.
[0002] In recent years, opportunities to put process mining into practice have increased, and conformance checking has become increasingly important in society. Process mining is a general term for a data analysis method for event logs using process models. Event logs (business logs), which are data sets that accumulate the execution history of tasks in a certain business, are analyzed using process models. A "process model" is a graph that represents the execution order of activities in the target business. In recent years, with the acceleration of digital transformation, it has become easier to accumulate event logs for many business processes. As a result, much attention is being paid to process mining techniques that analyze accumulated data.
[0003] Process mining involves a task called conformance checking, which considers whether a process model and a trace match correctly. A "trace" is a log that corresponds to one business event in a business process (event) log. "Conformance checking" is an effort to verify the compatibility of a trace with a process model to confirm whether the business process was executed correctly according to the specified process.
[0004] Conventionally, conformance checking methods based on process models and tracing based on generative grammar have been proposed (see Non-Patent Document 1). "Generative grammar" is a language model used to clarify the structure of sentences in the field of language processing. The generative grammar model obtained by conversion using this conventional method is called the Probabilistic Generative Process Model (PGPM).
[0005] A. Watanabe, Y. Takahashi, H. Ikeuchi and K. Matsuda, "Grammar-Based Process Model Representation for Probabilistic Conformance Checking," 2022 4th International Conference on Process Mining (ICPM), Bolzano, Italy, 2022, pp. 88-95, doi: 10.1109 / ICPM57379.2022.9980588.
[0006] However, conventional methods based on PGPM cannot express the relationships between activities when a process has a parallel structure. Furthermore, they cannot express subprocesses with a parallel structure (PMI: Parallel Multiple Instance), which have an indefinite number of executions and are executed in parallel with each other, without a predetermined execution count. Thus, the application area of process mining, a technique for analyzing business logs, has traditionally been limited.
[0007] The present invention has been made in consideration of the above-mentioned problems, and aims to expand the application area of process mining, which is a technique for analyzing business logs, more than ever before.
[0008] In order to solve the above problems, the present disclosure provides an analysis device that analyzes business logs, the analysis device having: a generative grammar conversion unit that converts a process tree representing an input process model into a parallel context-free grammar, which is a data format in which probability can be calculated, by using parallel generation rules that add the concept of a parallel structure to the generation rules; and a trace analysis unit that outputs a correspondence between the parallel context-free grammar and the business log as a set of syntax trees, based on a trace representing the parallel context-free grammar and the business log converted by the generative grammar conversion unit.
[0009] As described above, the present invention has the effect of expanding the scope of application of process mining, which is a technique for analyzing business logs, compared to conventional techniques.
[0010] 12. A diagram showing an example of a process model expressed in BPMN. A diagram showing conformance checking of a process model and traces. An example of a process tree equivalent to the process tree shown in FIG. 1. A table for explaining each operator in the process tree. An example of a trace tree showing the analysis results of conformance checking using PGPM, an existing method. An example of a process model having Parallel Multiple Instance, which is an indefinite number of parallel executions. A functional configuration diagram according to this embodiment. A diagram showing an example of a process tree. A diagram showing a sequence of specific generation rules included in the trace tree. A table showing the conversion of process tree nodes into parallel generation rules according to operators. A diagram showing an example of a syntax tree. A diagram showing a graph representing the syntax tree shown in FIG. 11. A diagram showing the process of rearranging tasks by serializing the graph representing the syntax tree shown in FIG. 12. A diagram showing an example of an algorithm for the Left Corner Branching method. A diagram showing an example of forward processing. A diagram showing an example of backward processing. A diagram showing parallel generation rules of CCFC (Concurrent Context-free Grammar), a new generation grammar that can represent parallel structures. A diagram showing the process of acquiring a syntax tree using the Extended Left Corner Branching method of this embodiment. FIG. 17 is a diagram showing a process of acquiring a syntax tree using the extended left corner branching method of this embodiment. FIG. 18 is a diagram showing a process of acquiring a syntax tree using the extended left corner branching method of this embodiment. FIG. 19 is a diagram showing a process of acquiring a syntax tree using the extended left corner branching method of this embodiment. FIG. 20 is a diagram showing a process of acquiring a syntax tree using the extended left corner branching method of this embodiment. FIG. 21 is a diagram showing a process of acquiring a syntax tree using the extended left corner branching method of this embodiment. FIG. 22 is a diagram showing a process of acquiring a syntax tree using the extended left corner branching method of this embodiment. FIG. 23 is a diagram showing an example of a process model expressed in BPMN, corresponding to FIGS. 17 and 18A to 18H. FIG. 24 is an electrical hardware configuration diagram of an analysis device according to an embodiment.
[0011] [Outline of the embodiment] Hereinafter, an outline of the embodiment of the present invention will be described with reference to the drawings, in comparison with conventional techniques.
[0012] First, let us consider conformance checking for a process model written in a notation called Business Process Modeling Notation (BPMN), as shown in Figure 1. This is an example of a process model expressed in BPMN.
[0013] In BPMN, a process model can be represented as a graph consisting of nodes consisting of tasks (squares in Figure 1), branches (diamonds with no letters in Figure 1), parallel branches (diamonds with a + sign in Figure 1), start symbols (black circles in Figure 1), and end symbols (white circles in Figure 1), as well as directed edges connecting each node. A process model represents the order relationships between tasks. For example, an edge from the start point to task a indicates that task a will be executed immediately after the start, and an edge from task e to task f indicates that task f will be executed after task e is executed. Multiple edges coming out of a branch indicate that one of the tasks beyond that edge will be executed; for example, either task b or c will be executed. A parallel branch indicates that all tasks beyond the edge will be executed in any order. For example, tasks e, f, and g will all be executed, but it is unclear which of tasks e and g will be executed first, or which of tasks f and g will be executed first (in Figure 1, task e will always be executed before task f).
[0014] Furthermore, "the trace conforms to the process model" means that the trace can be output when the process model is executed in the order indicated by the edges. Figure 2 shows conformance checking of the process model and the trace. For example, as shown in Figure 2,<a,b,d,c,e,g,f> If we consider that the tasks were executed in the order of S1 to S7, all activities (task execution history) will fit into the process model without any order breakdown.<a,b,f,e,g> does not conform to the process model because there is no way in the process model to reach task e after task f.
[0015] Conformance checking is the task of investigating the correspondence between process models and traces in this way. By obtaining the correspondence between traces and process models, it is possible to perform tasks such as checking for deviation activities that do not correspond to any tasks in the trace, predicting the completion time of a project based on the actual execution log, and model evaluation, which evaluates how well the process model matches the event log. As these are mathematical principles that form the basis of data analysis, conformance checking can be said to be a core technique in process mining.
[0016] The process model shown in Figure 1 is written in Business Process Modeling Notation (BPMN), the most well-known graph format for expressing business processes. However, in reality, business processes can be expressed in a variety of ways, and a well-known format other than BPMN is the process tree. An example of a process tree is shown in Figure 3. Figure 3 is an example of a process tree equivalent to the process tree shown in Figure 1.
[0017] In a process tree, symbols called operators shown at each node other than leaf nodes represent the relationship between child nodes. For example, if b and c are present under a node with an x representing a choice, this indicates that either b or c will be executed. The meaning of each operator is shown in Figure 4. Figure 4 is a table explaining each operator in a process tree.
[0018] Furthermore, the presence of another node other than a leaf node (a node with no children) under a node indicates a subprocess inclusion relationship. For example, the presence of g and → (with ef as a child node) under a + representing parallelism indicates that the subprocess that executes e and f in order and g are executed in parallel. A process tree can be converted into a process model, and the process tree in Figure 3 represents a business process equivalent to the process model shown in Figure 1. This means that the models in Figure 1 and Figure 3 can both execute only the same trace.
[0019] In recent years, methods for conformance checking of process trees and traces based on generative grammar have also been proposed. Generative grammar is a language model used in the field of language processing to clarify the structure of sentences. Non-Patent Document 1 shows that if a process model can be represented as a tree structure called a process tree, it can be converted into an equivalent generative grammar, and conformance checking can be performed using syntactic analysis techniques in the field of language processing. The generative grammar model obtained by the conversion in Non-Patent Document 1 is called a Probabilistic Generative Process Model (PGPM). In this PGPM-based method, if a trace tree that shows the correspondence between the trace and the model is obtained, it means that the trace is compatible with the target process model. Figure 5 shows a trace<a,b,d,c,e,f,g> Figure 2 shows the trace tree obtained as a result of conformance checking the process tree in Figure 3. Both Figure 2 and Figure 5 show the same process model and the correspondence between traces. The trace tree is a graph that represents the process of generating traces in the process tree, and by using the same algorithm for parsing, it is possible to perform relatively efficient and accurate conformance checking.
[0020] However, PGPM does not represent the relationships between activities under a parallel structure. As shown in Figure 5, all the operations under PARALLEL(+) are output in a single line. In fact, referring to Figure 3, e and f are connected by SEQUENCE(→) under PARALLEL, but PGPM does not represent such a hierarchical structure. The hierarchical structure is useful information that indicates the relationships between data, and losing this information makes data analysis difficult.
[0021] Furthermore, there may be process models that cannot be expressed using PGPM. Of the operators shown in Figure 4, four widely used operators in process trees are SEQUENCE, SELECTION, PARALLEL, and LOOP. However, these operators alone cannot express subprocesses that execute an indefinite number of times in parallel. Business processes can utilize subprocesses that execute other processes from any process, as shown in Figure 6. In particular, subprocesses that do not have a predetermined number of executions and that execute in parallel with each other are called Parallel Multiple Instance (PMI) in BPMN. In BPMN, PMI is depicted as a subprocess with "||" as shown in Figure 6. The following traces can be generated from the business process shown in Figure 6:
[0022] -<a,b,c,d> -<a,b,c,b,c,d> -<a,b,b,c,c,d> -<a,b,b,b,c,c,c,d> -<a,b,c,b,b,c,c,d> When this process is executed, b and c are always executed the same number of times. In addition, for every b, there is a corresponding c, and b is always executed before the corresponding c. A normal process tree does not have operators that correspond to PMI, and PMI cannot be expressed using only other operators. Therefore, existing conformance checking methods based on generative grammars cannot handle business processes with such PMI.
[0023] Business processes with PMI structures often appear in common business processes. For example, if a restaurant receives an order and then executes cooking tasks in parallel, the number of cooking tasks executed in parallel is determined by the customer's order. In this way, realizing conformance checking that can handle PMI will greatly expand the scope of application of process mining.
[0024] In consideration of the above-mentioned conventional problems, this embodiment realizes a method for performing conformance checking when a process model with a parallel structure and a business log (trace) are given. Specifically, we propose Concurrent Context-Free Grammar (CCFG), a new generative grammar that can represent parallel structures, and explain how to represent a process model using CCFG. Furthermore, we also propose an extended left corner branching method that estimates the generation process assuming that the trace is generated based on the CCFG and enables conformance checking.
[0025] [Details of the embodiment] Next, the details of the present embodiment will be described with reference to the drawings.
[0026] 7 is a functional configuration diagram of the present embodiment. An analysis device 30 of the present embodiment receives as input a process tree PT representing a business process and a trace σ representing the history of an actually executed business operation, and outputs whether or not the trace is likely to be observed when the business operation is performed according to the process tree, i.e., whether or not the trace matches the process tree.
[0027] The analysis device 30 of this embodiment has a generative grammar conversion unit 31, a trace analysis unit 33, and a conformance determination unit 35. These units each have a function that is realized by instructions from a CPU 301 shown in Fig. 20, which will be described later, based on a program.
[0028] The generative grammar conversion unit 31 converts a process tree representing an input process model into a new generative grammar that generalizes CFG, called a concurrent context-free grammar (CCFG), which is a data format that enables probability calculation by using parallel generation rules that add the concept of parallel structure to the generation rules. In this case, the generative grammar conversion unit 31 converts the process tree into a set of generation rules with parallel structure using a simple conversion method.
[0029] The trace analysis unit 33 analyzes the generation process of the trace from the CCFG based on the trace representing the CCFG and the business log converted by the generative grammar conversion unit, and outputs the correspondence between the CCFG and the business log as a syntax tree set t. In this case, the trace analysis unit 33 uses an algorithm that extends the left corner branching method, which is an analysis algorithm for a context-free grammar (CFG), which is a set of generation rules for the process tree and is a data format for which probabilities can be calculated, to enable analysis of a CCFG, which is a grammar with a parallel structure.
[0030] The compatibility determination unit 35 determines whether the process model and the business log are compatible based on whether or not the syntax tree set t is output. If the syntax tree set t is not empty, the compatibility determination unit 35 determines that the process model and the business log are "compatible," and if the syntax tree set t is empty, the compatibility determination unit 35 determines that the process model and the business log are "incompatible."
[0031] In this embodiment, a trace, which is sequence data of work for each case, is given as one of the inputs to the analysis device 30. If the activity space that can be performed in a certain task is A, then the trace σ∈A * indicates the order in which activities are performed for each item. For example, σ =<a,b,d,c,e> means that activities a, b, d, c, and e were executed in that order. In analyzing actual business data, an event log is provided that records all activities related to the business, but in this embodiment, the event log is considered to be a collection of traces, and the traces for each case are analyzed in order. Generally, event logs are often assigned a case ID that identifies the case, and most event logs meet this prerequisite. The process tree PT, which is another input in this embodiment, is an ordered tree in which each node has an operator or task, and is specifically defined as follows, and is a set of activities that can take A.
[0032] PT = [(p1, o1, l1),…,(p |PT| ,o |PT| ,l |PT| )] However, p iis the parent node number of the node. i When i = -1, the i-th node is the root. There is only one root. There is no rule, but usually the first node is the root. The operators for each node are shown below.
[0033] The symbols in {} represent, from the left, SEQUENCE, SELECTION, PARALLEL, LOOP, and PMI, respectively, as shown in FIG. 4, and are represented as T if they are leaf nodes.
[0034] Also, l i ∈A∪{τ} represents the mapping of node i to a label indicating an operator or task. If node i is not a leaf node, then l i Let =τ.
[0035] Next, an example of a process tree is shown in FIG. 8, and the same process tree is shown graphically in FIG.
[0036] Before explaining the CCFG proposed in this embodiment, we will explain the Probabilistic Generative Process Model (PGPM), a closely related conformance checking method. The "PGPM" is another process model representation in which a process tree is converted into a context-free grammar, which is a set of generation rules and is a data format in which probabilities can be calculated. A "generation rule" is expressed in the form A⇒α using a symbol A and a symbol string α, and indicates that another symbol string α can be generated from A.
[0037] The components of PGPM are non-terminal symbol M, terminal symbol E, production rule set R, and start symbol $. Non-terminal symbol M corresponds to a node in the process tree that has an operator other than T. Terminal symbol M corresponds to an activity. The production rule set is A⇒α(A∈M,α∈(M∪E) * ), which expresses the transformation of any non-terminal symbol into another string of symbols.
[0038] PGPM checks the compatibility of a process tree and a trace by exploring the process of generating a trace, called a trace tree, which is equivalent to a path, instead of a path.<a,b,d,c,e,g,f> Figure 3 shows a trace tree that demonstrates compatibility with the given sequence of terminal symbols. Figure 9 also shows the sequence of specific generation rules contained in the trace tree. The trace tree applies one of the generation rules $⇒α to the start symbol $. If each element of the resulting symbol string α is a non-terminal symbol, another generation rule can be applied to it. The process of repeatedly applying generation rules until all symbols become terminal symbols can be represented as a tree structure, which is the trace tree corresponding to the relevant terminal symbol string. Figure 3 illustrates that σ1 can be obtained by applying the generation rules in the order shown in Figure 9. This means that σ1 is compatible with the process model expressed in PGPM.
[0039] However, referring to Figure 5, all the processes under PARALLEL(+) are output in a row. In fact, referring to the process tree in Figure 3, e and f are connected by SEQUENCE(→) under PARALLEL, but PGPM does not represent such a hierarchical structure. Also, PGPM cannot represent an indefinite parallel structure such as PMI.
[0040] The reason why the conventional method, PGPM, cannot handle parallel structures is that CFG cannot represent the discontinuity of dependencies that arise due to parallel structures. Figure 5 shows a trace tree that has SEQUENCE(→) under PARALLEL(+) and represents a hierarchical structure. When there is a parallel structure, the T in Figure 5 11 Edge from to f and T 12 Because the order of two subtrees is indeterminate, such as the edge from g to g, correspondences may cross over on the graph. This crossover of relationships between input sequences is called a "discontinuity." CFGs are known to be unable to represent grammars with discontinuities, which prevents PGPM from properly handling parallel structures.
[0041] Therefore, in this embodiment, we propose a new generative grammar called Concurrent CFG (CCFG) that can represent parallel structures, and use CCFG instead of CFG. To represent parallel structures, CCFG uses parallel generation rules in the generation rule set R, which are generalized by adding the concept of parallel mechanisms to the generation rules. The parallel generation rules have the following structure:
[0042] A⇒([α1]…[α K ]) Here, the following relationship holds:
[0043] In a column generation rule, elements enclosed in "[ ]" in a certain region are in a parallel relationship. Here, "two elements are parallel" means that there is no defined order relationship between the two elements. Intuitively, if the left-right arrangement of a symbol string represents an order relationship, then the "[ ]" are lined up in the same column as shown below, and represent that there is no left-right order relationship.
[0044] A generation rule A⇒α in a normal CFG can be considered as a parallel generation rule with one parallel element, such as A⇒([α]). In other words, a parallel generation rule can be said to be a generalization of a CFG with a parallel structure added. For simplicity of notation, in this embodiment, one parallel sequence ([α]) can be expressed as α. For example, A⇒([α])=A⇒α, and <([([ABC])([D][E])([FG])])> =<ABC([D][E])FG> A CCFG is defined as the following set of elements:
[0045] CCFG: G=(M,E,R,$) is called a concurrent context-free grammar (CCFG).
[0046] M: Set of non-terminal symbols E: Set of terminal symbols R: Set of parallel production rules $: Start symbol ($∈M) Parallel production rule: r=A⇒([γ1]…[γ |r| ]) - R:M→((M∪E) * )* - Input A (input), ([γ1]…[γ |r| ]) is called the output.
[0047] - γ i is called the i-th argument of r. In this paper, each term is denoted by [ ].
[0048] -|r| represents the number of terms in r.
[0049] Each node in the process tree can be converted into a CCFG. Suppose the symbol corresponding to the i-th child node of a node v in the process tree can be expressed as follows:
[0050] In this embodiment, the generative grammar conversion unit 31 converts an operator o of an arbitrary node v into a v Parallel generation rules are added according to the process tree, as shown in FIG. 10. Note that for simplicity of explanation, the total number of child nodes is set to three in FIG. 10, but in reality, any number of child nodes can be supported. The CCFG constructed from the parallel generation rule set thus obtained can be combined with the serialization process described below to generate a trace identical to the original process tree.
[0051] In CCFG, traces are generated through a two-stage process. The process of generating traces in CCFG is called a "trace tree." A trace tree consists of two parts: a syntax tree and a correspondence relationship.
[0052] Syntax tree: A tree-structured sequence z=[(r1,j1,d1),…,(r |z| ,j |z| ,d |z| )] is called a syntax tree.
[0053] -r i ∈R: parallel production rule - j i ∈{-1,1,…,|z|}: the index of the parent node of the node j i If it is -1, then the i-th node is the root. There is one and only one root in any syntax tree.
[0054] The index of the parent node that indicates the number of items in which the node is included. i If it is -1, the i-th node is the root.
[0055] An example of a syntax tree is shown in Figure 11. Figure 12 shows the same syntax tree as a graph. In Figure 12, child nodes included in the same term are represented by hyperedges, which have multiple outputs for one input. Note that child nodes included in different terms form different hyperedges. For example, in Figure 11, at node 2, child nodes 3, 4, 13, and 14 are the same term, so four nodes are linked to one hyperedge, but at child node 4, child nodes 5 and 6 are different terms from child node 7, so two hyperedges link two nodes and one node, respectively.
[0056] In Z, the order of tasks executed in parallel is undetermined. Therefore, in the second step, a serialization process is performed to remove the parallel relationships within the task sequence and assign an order relationship to all elements of Z. This serialization allows us to obtain a trace.
[0057] To serialize Z, we define a serialization function f(Z, δY) that orders each element. Here, δY is a random numerical sequence with the same number of elements as Z, which is used to give Z an ordering relationship. The serialization function must be defined so that subsequences in a parallel relationship can be arranged in any order. However, in this case, the order of tasks that already have an ordering relationship must not be changed.
[0058] First, consider how to obtain the variable δY to give an order to the task sequence. i (i=1…|Z|) i The work time variable corresponding to i (δy≧0). δy i As the name suggests, Task Z i is a variable that corresponds to the time it takes to complete the iAlthough any setting is possible for how to obtain this, the easiest method is to obtain it according to a Poisson distribution, exponential distribution, etc. By setting the parameter λ of the probability distribution for each task, it is also possible to express the expected working time for each task in a business process.
[0059] δY=(y1,…,y |Z| ) does not have any constraints on the order of tasks, but the following time conversion algorithm, duration(Z, δY, y0), can be used to obtain a sequence Y in which the work time δY is converted into the work completion time.
[0060] And - duration(Z,δY,y0) - y last ←y0 - For i=1…|Z| - If z i If the element of is not enclosed in ([]) (if it is a single parallel element), - y i ←y last +δy i -y last ←y i - Else - z i The parallel partial task sequence is shown below.
[0061] The corresponding work time variable string is shown below:
[0062] - For k=1,…,K i - Y k ←duration(Z k ,δY k ,y last ) - Y for Y k are concatenated.
[0063] Also, -y k last Y k The final element of
[0064] - Return Y However, y_0 is an arbitrary real value that represents the start time of the task, and no matter how you set the initial value, it does not affect the order of the serialized task sequence that is generated at the end, so you can just set y0=0. duration(Z,δY,y0) is the duration of any task zi For the task that was just completed, the working time for the task is y last As, y last +δy i z i However, all tasks at the beginning of a task sequence in a parallel structure are considered to be completed at the same time y last The completion time of the work is obtained based on this.
[0065] Next, Figure 13 shows the serialization of the syntax tree shown in Figure 12. In serialization, first, δY, which corresponds to the execution time of each task, is determined probabilistically. In Figure 13, the width of each task corresponds to δY. Once δY is determined, f(Z, δY) calculates Y from duration(Z, δY, 0). In Figure 13, Y corresponds to the rightmost position of each task. Then, Z is sorted in ascending order of the corresponding Y value. If there is an order relationship between corresponding tasks in Y, the order is maintained. Therefore, in the task sequence serialized based on Y, only elements with a parallel structure are probabilistically sorted. Since an activity is uniquely determined for each task, a trace is obtained as a result. In Figure 13, the bottommost row corresponds to the trace.
[0066] So far, we have explained the process of generating traces in CCFG.
[0067] The purpose of the analysis device 30 of this embodiment is to check whether a given trace can be generated from a CCFG.
[0068] In language processing, checking whether an input sentence fits the generative grammar is called "syntactic analysis." Many syntactic analysis methods have been proposed for CFGs, but ordinary syntactic analysis methods cannot be applied directly to CCFGs, which have a parallel structure.
[0069] Therefore, the trace analysis unit 33 of this embodiment improves the left corner branch method, which is an existing parsing algorithm, to realize parsing in CCFG, that is, conformance checking.
[0070] In CCFG, a generative grammar that can represent parallel structures, the trace analysis unit 33 performs syntactic analysis of traces using an algorithm that adapts left corner branching to CCFG. Left corner branching is a bidirectional breadth-first algorithm that is an improvement over bottom-up sequential analysis. In left corner branching, the intermediate state of analysis up to the point where a syntax tree is obtained is held in a tree structure called a "parse tree" and stored in a database called a chart. In left corner branching, the trace analysis unit 33 examines the generation rule corresponding to the incomplete edge at the bottom left of the parse tree. However, in the case of CCFG, it must be noted that when there is a parallel structure, there may be multiple leftmost incomplete branches.
[0071] [Terminology and definitions for the left corner branching method] In the following, we will define the concepts related to the left corner branching method. Unless otherwise specified, A, B, C, D∈M, α i ,β i ,γ i ∈(M∪E) * Let (i∈N). Also, let ε represent a string of length 0.
[0072] <About the parse tree> Node: q=A⇒([α1・β1]…[α K ・β K ]) is called a clause. "・" indicates the progress of analysis at a certain point in time, and the string to the left of "・" indicates a string that has already been analyzed, and the string to the right of "・" indicates a string that is currently being predicted and is expected to appear. For any i∈1,…,K, β i If = ε, then we say that the clause q is complete. If not, we say that it is an incomplete clause.
[0073] branch: node q=A⇒([α1・β1 ]…[α K ・β K ]), b=(A,α i ,β i) is called the i-th branch. The concept of a branch is different from that of a syntax tree in normal parsing, and represents a parallel series. Therefore, care must be taken because the number of child nodes in the syntax tree does not match the number of branches. For example, there are two branches for node A⇒([B・CD][E・FG]), (A,B,CD) and (A,E,FG). Let |q| represent the number of branches for q.
[0074] Parsing tree: sequence t=[(q1,i1,d1),(q2,i2,d2),…,(q |t| ,i |t| ,d |t| )] is called a "parse tree." k represents the parent node number. k =-1 indicates that the k-th node is the root. There is only one root in the parse tree. There is no clear definition, but usually the node of the first element is the root. d k indicates the branch number of the parent node in which the node is included. k =-1 indicates that the kth node is the root.
[0075] Nodes and branches are components of a parse tree, and represent the state during parsing. The syntax tree that we ultimately want to obtain has no state (no "・" symbol) and no incomplete branches, so it is a different concept from the parse tree. However, as the analysis progresses, the parse tree may become a state where all branches are complete, and the result corresponds to the syntax tree. In other words, the syntax tree is obtained using the parse tree.
[0076] In the node q = A⇒([α1・B1β1]) included in t, the branch set {(A,α1,B1β1),(B1,α2,B2β2),…,(B n-1 ,α n ,B n β n )} exists, that is, derivation A⇒ from A * ([γ1][α n B n β n ][γ3]), A is a node (B n-1 ,α n ,B n β n) is an ancestor of t. By definition, all ancestors of an incomplete node are incomplete. In t, the sibling node ((i k ,d k The order of nodes ) whose ) are the same is the order of their appearance in t.
[0077] <Left corner incomplete branch (lcib)> Node q=A⇒([α1・B1β1]…[α K ・B K β K ]), Bi ≠ ε and any incomplete clause q' in t is B i If node q does not have an ancestor, then the i-th branch g i is called the "leftmost incomplete branch". Also, a node that has at least one leftmost incomplete branch is called the "left corner incomplete node" of t.
[0078] From the above definition, the conditions for a branch to be leftmost incomplete are that it is an incomplete branch and that it is not an ancestor of any other incomplete node. Based on this, the following function lcib(t) can be used to obtain the set of all leftmost incomplete branches of any parse tree t. lcib(t) - For k=1,...,|t|, check all branches i=1,...,|q| of node q in (q,j,d)=t[k], and if they are incomplete, put k into the incomplete node set and (k,i) into the incomplete branch set. - for k=1,...,|t| - If k is included in the incomplete node set, put (j,d) of (q,j,d)=t[k] into the non-leftmost incomplete branch set (unless j=-1). - The leftmost incomplete branch set is the incomplete branch set minus the non-leftmost incomplete branch set. - (k among the elements (k,i) of the incomplete edge set is set as the leftmost incomplete node set.) In this procedure, in Figure 14, the first line confirms that it is an incomplete edge, and acquires both the incomplete edge and the node. In lines 2 and 3, the branch of the parent node of the incomplete node is considered to be a branch that is not the leftmost, and in line 4, the set excluding these branches is considered to be the leftmost incomplete edge. The function for obtaining the leftmost incomplete edge set using the above procedure is called lcib(t).
[0079] In a CFG, there is only one leftmost incomplete branch in any analysis tree, but in a CCFG, there can be multiple leftmost incomplete branches. The extended left corner branching method, which will be described later, is an improvement over the left corner branching method, which analyzes all leftmost incomplete branches when there are multiple incomplete branches that fall into the so-called "left corner."
[0080] <Definition of the acquisition function for the first terminal symbol> As will be described later, in the left corner branch method, σ :l If the symbol corresponding to the leftmost incomplete branch lcib(t) of the parse tree obtained for t is A, then A is the next terminal symbol σ l To do this, we need a way to obtain all terminal symbols that can be at the beginning of a string that can be derived from any symbol A. Therefore, we define a function first(A) that obtains the set of relevant terminal symbols as follows:
[0081] - For every terminal symbol w, let first(w) = {w}. - For every production rule A⇒([B1γ1]…[B K gamma K ])∈R,
[0082] Here, the first function indicates the terminal symbol at the beginning of a string derived from an arbitrary symbol A. However, in the case of parallel production rules, A⇒ * In the case of ([bc][de][fg]), the first elements of each parallel term, b, d, and f, are all called "located at the beginning."
[0083] Note that the first function is recursively defined by the first function. In fact, all functions can be obtained by searching for functions in order from the leaf nodes of the process tree toward the root node. Also, since the first function is determined by R, it is sufficient to calculate it once as long as the given process tree does not change.
[0084] <Definition of other useful functions> Up to this point, we have defined lcib(t) and first(A), but in addition to these, we will define some simple functions that are useful for calculations in the left corner branch method. - wait(q,d): Returns B in the incomplete branch d=A⇒α・Bβ of q. - step(q,d): Returns B in the incomplete node q=A⇒([α1・B1β1]…[α d ・B d β d ]…[α K ・B K β K ]) of the dth branch (A,α d ,B d β d ) analysis is advanced one step to the next: q'=A⇒([α1・B1β1]…[α d B d ・β d ]…[α K ・B K β K ]).
[0085] [Extended Left Corner Branching Method] Now that the definition of the relevant elements is complete, we will explain the extended left corner branching method, which performs syntactic analysis from traces and CCFGs. As the name suggests, the extended left corner branching method is an extension of the left corner branching algorithm so that it can be applied to CCFGs. Since ordinary CFGs can be considered a special case of CCFGs in which the parallel generation rule is limited to one term, the extended left corner branching method can also be applied to the analysis of ordinary CFGs.
[0086] Here, the sequence of terminal symbols from 1 to l-1 of the trace σ is denoted as σ :l Let the lth terminal symbol be σ l Let's say.
[0087] The "left corner branching method" is a sequential algorithm that analyzes the trace σ from the first terminal symbol. As shown in FIG. 14, the string σ is analyzed in order from l (el)=1. :l The parse tree in which appears at the beginning is called chart[σ :lAs shown in Figure 15, the analysis is performed by "forward processing" which extends the current lcib set until it reaches a given terminal symbol, and then, as shown in Figure 16, when it reaches the relevant terminal symbol, it performs "backward processing" which completes the branches in the reverse direction toward the root node to obtain the next parse tree to be processed.
[0088] In the left corner branching method, if there is a parse tree in chart[σ] where all nodes are completed (or the root is completed), then the trace σ is considered to be generable from the target CCFG, and the result of conformance checking is considered to be conformance. In this case, if a syntax tree is obtained at the same time, all nodes A⇒([γ1・]…[γK・]) of the obtained parse tree are converted to A⇒([γ1]…[γ K ]) and create a tree that converts it.
[0089] [Process for Acquiring a Syntax Tree by the Extended Left Corner Branching Method] Next, the process for acquiring a syntax tree by the Extended Left Corner Branching method will be described with reference to Fig. 17 and Fig. 18A to Fig. 18H. Fig. 17 shows parallel generation rules of CCFC (Concurrent Context-free Grammar), a new generative grammar that can represent parallel structures.
[0090] Here, we consider a CCFG with parallel generation rules shown in FIG. 17 and a trace σ=<a,d,b,c,e,e,f,f> Analysis is performed using the above. In FIG. 10 and the following explanation, parentheses are omitted as appropriate for simplicity when there is only one term in the parallel generation rule. Also, FIG. 19 shows the BPMN representation of the process model corresponding to FIG. 17. Note that in FIG. 19, subprocesses are used to represent PMI, which is an indefinite number of executions with no predetermined number of executions, and it is assumed that subprocesses consisting of e and f are executed in parallel an indefinite number of times.
[0091] In the initial state, the analytic tree t0 = [($ → [· → 1], -1, -1)] corresponding to the derivation from the start symbol $, as shown in FIG. 18A, is inserted into chart[ε]. This corresponds to the second line of the algorithm shown in FIG. 14. In the figure, completed nodes are shown in white, and incomplete nodes are shown in gray. In addition, the node connected to the leftmost incomplete branch of each analytic tree is surrounded by a red dashed line.
[0092] Next, consider the symbol σ1=a in l=1. Considering lcib(t0) for the only parse tree t0 contained in chart[ε], the only (incomplete) branch is the i=1th element (q,j,d)=(($,ε,→1),-1,-1) of the k=1th element. In this case, wait(q,1)=→1 and a∈first(→1)={a} (lines 8 and 9 of the algorithm shown in Figure 14). Therefore, by expanding →1, we can derive a symbol string starting with the first character a, so we can use forward( ,t,1,1,a) is executed (line 10 of the algorithm shown in FIG. 14).
[0093] forward( ,t,1,1,a), since →1 is a terminal symbol, we obtain all parallel production rules that have →1 as input from R according to the 10th line of the forward processing shown in Figure 15. This is because there is only the following relationship.
[0094] Also, since first(T2)={a}, we express the new clause as follows and add (q child ,1,1) is added to create a tree t1 (lines 12, 13, and 14 of forward).
[0095] At this time, t1 is expressed as shown in FIG. child In the first term of child , 1)=T2. →1≠T2 and a∈first(T2), so the condition on line 18 of the forward processing shown in FIG. 15 is met, so recursively, forward( ,t1,2,1,a) is executed.
[0096] I will omit the details, but forward( , t1,2,1,a), (q,j,d) = (T2⇒[·a],2,1) is newly added, and the parse tree t2 corresponding to FIG. 18C is obtained, and forward( ,t2,3,1,a) is executed recursively.
[0097] forward( , t2,3,1,a), q=T2⇒[·a] and wait(q,1)=a, so this corresponds to the third line of the forward process shown in FIG. new =step(q,1)=T2⇒Create [a・] and change the last element q of t to q new t3 corresponds to Figure 18D. As a result, q new is completed because there is no symbol to the right of the symbol ".", so it follows the 7th line of the forward process shown in FIG. 15 and ,t3,3).
[0098] backward( , t3, 3), the object to be investigated is q = T2⇒[a·]. Since q is completed, the parent node q of q shown below is found according to the 5th and 6th lines of the backward processing shown in FIG. parent Step function to advance by one, and q new Create a.
[0099]
[0100] And then q to t3 parent , and set it as t4 (line 7 of the backward processing shown in FIG. 16). This t4 corresponds to FIG. 18E. ,t4,2) is executed recursively.
[0101] Next, backward( ,t4,3), the object q to be investigated is expressed as follows, but this is incomplete.
[0102] Therefore, chart[ ], and then t4 (corresponding to FIG. 18E) is added to the [?] and the process ends.
[0103] This is the series of processes for l = 1. It may look complicated, but in reality, the forward function adds clauses according to the generation rules from the starting symbol $ until it reaches the l = 1 symbol a, and once it reaches a, it simply uses the backward function to complete as many clauses as possible.
[0104] Next, consider the symbol σ2=d at l=2. Consider lcib(t4) for the only parse tree t4 contained in [], and the incomplete branches are the d=1 branches of the first and second nodes. Since the first node is the ancestor of the second node, the leftmost incomplete branch of t4 is (2,1), that is, it can be expressed as follows:
[0105] This is the node enclosed by the dashed line in Figure 18E. In this case, first(+3) = {b, d}, so according to line 10 of the algorithm shown in Figure 14, forward(<a,d> ,t4,2,1,e) is executed.
[0106] Next, forward(<a,d> ,t4,2,1,e) Consider the following relationship.
[0107] Since wait(q,1)=+3 and e∈first(+3)={b,d}, we add a new node to q child =+3⇒([・→4][ ・T7]), and then (q child ,2,1) is added to create a parse tree t5 (lines 12, 13, and 14 of the forward processing shown in Figure 15). At this time, t5$ is represented as shown in Figure 18F. child has two terms, wait(q child ,1)=→4, wait(q child , 2) = T7. Although d∈first(→4) is not true, since d∈first(T7), the analysis is continued only for T7 that satisfies the condition on line 18 of the forward processing shown in FIG. 15, and then recursively forward(<a,d> ,t5,4,2,e) is executed.
[0108] I won't go into detail here, but forward(<a,d> In t5,4,2,e), after node T7⇒([·d]) is generated, step processing is executed and (T7⇒([d·]),4,2) is added to the parse tree. Then, backward processing is executed to execute (+3⇒([·→4][·T7]),2,1) to (+3⇒([·→4][T7]),2,1), and the parse tree t6 in Figure 18G is added to chart[<a,d> ] is added to
[0109] The subsequent analysis process is omitted, but finally, two parse trees with all nodes completed to chart[σ] are obtained. Figure 18H shows the two completed parse trees, with the only difference being that the edges to the two terminal symbols f are shown with two types (solid and dashed lines). These indicate which of the two e's correspond to the two f's, respectively.
[0110] [Hardware Configuration] Next, the electrical hardware configuration of the analysis device 30 will be described with reference to Fig. 20. Fig. 20 is a diagram showing the electrical hardware configuration of the analysis device according to the embodiment. The analysis device 30 is configured with one or more computers. When the analysis device 30 is configured with multiple computers, it may be referred to as an "analysis device" or an "analysis system."
[0111] As shown in FIG. 20 , the analysis device 30 is a computer and includes a CPU (Central Processing Unit) 301, a ROM (Read Only Memory) 302, a RAM (Random Access Memory) 303, an SSD (Solid State Drive) 304, an external device connection I / F (Interface) 305, a network I / F 306, a media I / F 309, and a bus line 310.
[0112] Of these, the CPU 301 controls the overall operation of the analysis device 30. The ROM 302 stores programs such as an IPL (Initial Program Loader) used to drive the CPU 301. The RAM 303 is used as a work area for the CPU 301.
[0113] The SSD 304 reads or writes various data under the control of the CPU 301. Note that instead of the SSD 304, a hard disk drive (HDD) may be used.
[0114] The external device connection I / F 305 is an interface for connecting various external devices, such as a display, a speaker, a keyboard, a mouse, a USB (Universal Serial Bus) memory, and a printer.
[0115] The network I / F 306 is an interface for performing data communication via the communication network 100 .
[0116] The media I / F 309 controls reading and writing (storing) of data from and to a recording medium 309m such as a flash memory, etc. The recording medium 309m includes a DVD (Digital Versatile Disc) and a Blu-ray (registered trademark) Disc.
[0117] The bus line 310 is an address bus, a data bus, etc. for electrically connecting the components such as the CPU 301 shown in FIG.
[0118] [Major Effects of the Present Embodiment] As described above, according to the present embodiment, by using CCFG and the Extended Left Corner Branch Method, it is possible to realize conformance checking between a process model with an indefinite number of parallel structures and a trace, which was not possible with conventional methods. Tasks with an indefinite number of parallel structures appear very frequently in business processes, so by using the present embodiment, it is possible to expand the scope of application of process mining, which is a technique for analyzing business logs.
[0119] Furthermore, although the main target of this embodiment is the analysis of business logs, parallel structures also appear slightly in natural language, which is the language spoken by humans, and therefore application to syntactic analysis in natural language processing is also expected.
[0120] [Supplementary Information] The present invention is not limited to the above-described embodiment, and may have the following configurations or processes (operations): (1) The analysis device 30 can be realized by a computer and a program, but this program can also be recorded on a (non-temporary) recording medium or provided via a communication network. (2) The CPU 301 as a processor may be a single processor or multiple processors.
[0121] 30 Analysis device 31 Generative grammar conversion unit 33 Trace analysis unit 35 Conformance determination unit 301 CPU 302 ROM 303 RAM (an example of a storage unit) 304 SSD (an example of a storage unit)
Claims
1. An analysis device for analyzing business logs, comprising: a generation grammar conversion unit that converts, from a process tree representing an input process model, into a parallel context-free grammar which is a data format in which probabilities can be calculated by using a parallel generation rule obtained by adding a concept of a parallel structure to a generation rule; and a trace analysis unit that outputs, as a syntax tree set, a correspondence relationship between the parallel context-free grammar and the business logs based on the parallel context-free grammar converted by the generation grammar conversion unit and a trace representing the business logs.
2. The analysis device according to claim 1, wherein the trace analysis unit uses an algorithm obtained by extending a left corner branching method, which is an analysis algorithm of a context-free grammar that is a set of generation rules and is a data format in which probabilities can be calculated, so as to be able to analyze the parallel context-free grammar which is a grammar having a parallel structure.
3. The analysis device according to claim 1, further comprising a conformity determination unit that determines whether or not the process model and the business logs conform based on whether or not the syntax tree set is output.
4. An analysis method executed by an analysis device for analyzing business logs, the analysis device performing: a generation grammar conversion process of converting, from a process tree representing an input process model, into a parallel context-free grammar which is a data format in which probabilities can be calculated by using a parallel generation rule obtained by adding a concept of a parallel structure to a generation rule; and a trace analysis process of outputting, as a syntax tree set, a correspondence relationship between the parallel context-free grammar and the business logs based on the parallel context-free grammar converted by the generation grammar conversion process and a trace representing the business logs.
Citation Information
Patent Citations
System and method for checking the conformance of the behavior of a process
US20140047445A1
Process model repairing method based on structure replacement
US20210049147A1
Visual conformance checking of processes
US20210200574A1