Evaluation device and evaluation method
The Bayesian estimation-based method integrates fitness, precision, generalizability, and simplicity into a unified index, addressing the fairness issues in conventional evaluations by offering a comprehensive assessment of process models.
Patent Information
- Application Number
- PCT/JP2024/000081
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-05
- Publication Date
- 2025-07-10
AI Technical Summary
Conventional methods for evaluating process models in process mining fail to provide a fair and simultaneous assessment across multiple viewpoints, leading to subjective weight assignments that compromise the fairness of evaluations.
A Bayesian estimation-based method for evaluating process models that calculates an approximation of generalization loss, integrating fitness, precision, generalizability, and simplicity into a single evaluation index, ensuring a fair and comprehensive assessment.
The method allows for a fair and comprehensive evaluation of process models by eliminating subjective weight assignment, providing a unified index that reflects the model's performance across all four viewpoints.
Smart Images

Figure JP2024000081_10072025_PF_FP_ABST
Abstract
Description
Evaluation device and evaluation method
[0001] The present disclosure relates to an evaluation device and an evaluation method.
[0002] Process mining is a well-known method for understanding the business processes that business operations are carried out in. In process mining, a key step is called "model evaluation," which evaluates whether the process model that models the business process appropriately represents the actual business process.
[0003] Generally, model evaluation is performed from four perspectives: "goodness of fit," "accuracy," "versatility," and "simplicity." A conventional method for evaluating the fit and accuracy of a process model is to use entropy calculated from log-likelihood (Non-Patent Document 1).
[0004] Leemans, SJJ, & Polyvyanyy, A. (2020). Stochastic-aware conformance checking: An entropy-based approach. In Advanced Information Systems Engineering: 32nd International Conference, CAiSE 2020, Grenoble, France, June 8-12, 2020, Proceedings 32 (pp. 217-233). Springer International Publishing.
[0005] However, evaluation indices for the above four perspectives are proposed based on different theories, and there is no conventional method that can simultaneously evaluate all four perspectives. Therefore, when evaluating whether a certain process model satisfies all four perspectives, for example, the evaluation is performed using a weighted sum of the evaluation values of each perspective. However, since the weights for each perspective are determined subjectively by the evaluator, arbitrariness cannot be eliminated and fairness for each perspective cannot be guaranteed.
[0006] The present disclosure has been made in consideration of the above points, and aims to provide a technology that can simultaneously and fairly evaluate a process model from multiple perspectives.
[0007] An evaluation device according to one aspect of the present disclosure is an evaluation device that evaluates a process model that represents a business process, and includes: an input unit that inputs the process model to be evaluated and an event log that is a log of traces that represent one or more operations observed in the business process; and an evaluation value calculation unit that calculates, as an evaluation value of the process model, an approximation of a generalization loss between a true distribution of the traces when the distribution of the traces is predicted from the process model by Bayesian estimation, based on the process model and the event log.
[0008] A technique is provided that allows a process model to be evaluated from multiple perspectives simultaneously and fairly.
[0009] FIG. 1 is a diagram showing an example of an event log. FIG. 2 is a diagram showing an example of a process model (part 1). FIG. 3 is a diagram showing an example of a process model (part 3). FIG. 4 is a diagram showing an example of a process model. FIG. 5 is a diagram showing an example of a hardware configuration of a process model evaluation device according to an embodiment. FIG. 6 is a diagram showing an example of a functional configuration of a process model evaluation device according to an embodiment. FIG. 7 is a flowchart showing an example of a process model evaluation process according to an embodiment.
[0010] Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings.
[0011] <Process Mining> When various large and small tasks are performed in a business, workers carry out the tasks according to some kind of business process. A business process represents the activities (tasks) performed by workers and the order in which they are performed. For example, in a bank's loan screening process, the first step is to "check the documents" from the applicant, and if the documents are inappropriate, they are "pushed back." On the other hand, if the documents are appropriate, they are "screened" according to certain conditions, and the screening results are "approved" by the person in charge's superior, before being "notified" of the results to the applicant. These activities - "check the documents," "pull back," "screening," "approval," and "notifying the results" - and the order in which they are performed make up a business process. Note that each activity can also be expressed as even more detailed activities.
[0012] While business processes are often explicit to workers, such as the loan screening process mentioned above, they are often implicit, not explicitly written down, especially for tasks that are not performed frequently. However, in order to consider streamlining and automating business processes, it is necessary to at least understand what activities exist in the business process and how long each activity takes to execute. Therefore, in order to consider streamlining and automating business processes in various businesses, it is necessary to accurately understand each business process of the target business.
[0013] A method known as process mining is known for understanding such business processes. Process mining is a general term for data analysis methods that use process models, which are data that represent business processes, and event logs, which are execution logs of activities in business. In recent years, with the acceleration of digital transformation, it has become easier to accumulate event logs in many business processes. For this reason, much attention has been focused on process mining, which analyzes these event logs. An event log is a history of activities performed by workers.
[0014] In process mining, a crucial step is called "model evaluation," which evaluates whether a process model accurately represents the actual business process. Process models are created and edited, for example, by analysts who perform data analysis or by workers who actually carry out the business process, or they are acquired by inferring them from event logs using a method called process discovery. At this point, whether the process model accurately represents the actual business process has a significant impact on the accuracy of subsequent data analysis. For this reason, various methods for evaluating the quality of process models are being considered in process mining.
[0015] Process models are generally evaluated from four perspectives: "fitness," "precision," "generalization," and "simplicity." Except for simplicity, these perspectives are usually evaluated using the process model and event logs.
[0016] <Explanation of Adaptability, Accuracy, Versatility, and Simplicity> Hereinafter, using specific examples of a process model and an event log, we will explain what aspects of the process model are evaluated in terms of adaptability, accuracy, versatility, and simplicity. First, an example of an event log is shown in Figure 1. The event log L shown in Figure 1 0 is made up of one or more events, and each event contains a "Case ID" that identifies the business case, a "Timestamp" that indicates the time the event occurred, and an "Activity" that indicates the activity (work) performed in that event. Note that events often contain other additional information (e.g., worker name, etc.), but for simplicity, the following discussion will consider a case where an event contains only a "Case ID," a "Timestamp," and an "Activity."
[0017] In the example shown in FIG. 0Visually, the event log is represented as a collection of traces, but mathematically, the event log is represented as a collection of traces. A trace is a string that lists the activities executed in one business case. For example, a trace abef means that activities a, b, e, and f were executed in that order in a business case with a certain Case ID. The event log L shown in Figure 1 0 If we express L as a collection of traces, 0 = {abef,acg,abdcef}.
[0018] In the following, the event log L shown in FIG. 0 Traces acdbg, abef, and acdcdbef are further accumulated for L 1 The explanation will continue assuming that the following holds: abef, acg, abdcef, acdbg, abef, acdcdbef.
[0019] A process model represents activities and the order in which they are executed, and is expressed, for example, as a graph such as those shown in Figures 2 to 5. As shown in Figures 2 to 5, in a process model, a business operation starts from a start point, each activity is executed according to conditions, and ends when it reaches an end point. At this time, the sequence of activities that are passed from the start point to the end point corresponds to a trace, and if a path corresponding to a certain trace exists, the process model is said to be able to generate that trace.
[0020] The process model G shown in FIG. 0 is a process model that represents the true business process, and this process model G 0 It is assumed that the business process represented by the process model G is executed in the actual business. 0 Under these circumstances, the process model G shown in FIG. 1 and the process model G shown in FIG. 2 and the process model G shown in FIG. 3 and which of these process models is the process model G 0 Event log L 1 Consider the situation you are trying to infer from.
[0021] - Goodness of Fit Goodness of fit is an index to evaluate to what extent a process model can generate traces included in an event log. A process model is a process model G that represents at least a true business process. 0 Event log L generated from 1 However, for example, in the process model G shown in FIG. 1 does not include activity g. Therefore, the process model G shown in FIG. 1 Event Log L 1 Therefore, it is not possible to generate the traces acg and acdbg included in the process model G 1 Event Log L 1 Since only a portion of the traces included in can be generated, its fitness is low.
[0022] Accuracy Accuracy, in contrast to goodness of fit, is an index that evaluates the extent to which a process model does not generate traces that are not included in the event log. If a "process model that can generate all possible traces" is prepared, the goodness of fit will show the best value no matter what traces are in the event log. The more a process model can generate traces that are not in the event log, the worse the accuracy will be. For example, in the case of process model G shown in Figure 4, 2 Event Log L 1 It is possible to generate all traces contained in the event log L 1 Traces that do not exist in the true process model G 0 Therefore, the process model G 2 is evaluated as having high fitness but low accuracy. A process model with sufficiently good fitness and accuracy (in other words, a process model that satisfies both fitness and accuracy) is a process model that generates only traces that exist in the event log, and does not generate traces that exist in the event log L, which is observed observation data. 1 This means that it is a good model that accurately represents the
[0023] Versatility Versatility is an index that evaluates how many traces a process model can generate that are not included in the event log but may actually appear. An event log is a collection of all traces observed during a certain period for a certain business operation. However, it is not guaranteed that all traces that can be generated will be observed during that period. For example, trace abg is generated by process model G. 0 This is a trace that can be generated by the event log L 1 does not include trace abg.
[0024] Process model G shown in FIG. 3 The process model G is designed to enumerate all observed traces. This allows the best possible fit and precision. However, in reality, this process model G 3 Event Log L 1 Since it is impossible to generate any trace not in a, it is also impossible to generate the trace abg. Versatility is an index used to evaluate poorly models that only generate such "observed traces."
[0025] ・Simplicity Simply put, simplicity is an index that evaluates the number of components in a process model. Process models are not only used for data analysis, but also serve as procedure manuals that show workers the details of their work. For this reason, even if a process model has sufficient suitability, precision, and versatility (in other words, it satisfies all of suitability, precision, and versatility) and accurately represents a business process, it is difficult to say that it is excellent if it is so complex that it is difficult for humans to understand. Simplicity is an index that indicates how easy a process model is to understand, and is generally evaluated by the number of elements in the process model (activities, edges that represent order, etc.).
[0026] Process model G shown in FIG. 1 and the process model G shown in FIG. 2 Both of them have a small number of activities and are relatively simple. On the other hand, the process model G shown in Figure 5 3The number of activities is extremely large, and the simplicity is low. Note that, unlike other viewpoints, the simplicity is generally calculated only from the process model, without using the event log.
[0027] As described above, the perspectives of suitability, accuracy, versatility, and simplicity each highly evaluate different process models. Therefore, when evaluating a process model in a model evaluation, it is necessary to combine indicators from various perspectives.
[0028] <Problems when Evaluating Process Models> While the above four perspectives are important in model evaluation, conventional methods can only evaluate some of the perspectives, and there is no method that can simultaneously evaluate all four of the above perspectives. This is because evaluation indicators for each of the above four perspectives are proposed based on different theories.
[0029] In particular, while compatibility, accuracy, and versatility evaluate the "fit between the process model and the event log," simplicity evaluates the "process model" itself, and is generally considered to be completely different. Therefore, when evaluating whether a certain process model satisfies all of the above four perspectives, for example, weights are assigned to the evaluation values of each perspective, and the evaluation is carried out by the weighted sum of the evaluation values. However, because these weights are determined subjectively by the evaluator, arbitrariness cannot be eliminated, and there is a problem in that fairness for each perspective cannot be guaranteed.
[0030] <Proposed method> Below, we propose a method that can simultaneously and fairly evaluate process models from the above four perspectives under a single theoretical system called Bayesian estimation. This proposed method eliminates arbitrariness on the part of the evaluator and makes it possible to naturally and fairly evaluate process models from the four perspectives of suitability, accuracy, versatility, and simplicity.
[0031] In the proposed method, a process model G to be evaluated and an event log L are given as input. Hereinafter, a set of possible activities (tasks) is defined as A, and the event log L is a trace σ∈A, which is a sequence of activities. *That is, the number of traces in the event log is |L|, and the i-th trace is σ i Then, L = (σ 1 , ..., σ |L| ) where A * represents the Cartesian product of any number of A's.
[0032] In this case, the proposed method outputs an evaluation index value W(G) that indicates how superior the given process model G is. 1 , G 2 When each evaluation index value W(G 1 ), W(G 2 ) to compare the process model G 1 , G 2 Here, the smaller the value of W, the better the model. For example, W(G 1 ) < W(G 2 ), then the process model G 1 Process Model G 2 It can be said that this is a better process model than the previous one. A good process model means one that is superior in all four aspects: suitability, accuracy, versatility, and simplicity.
[0033] Therefore, the proposed method makes it possible to, for example, re-edit the process model G according to the value of the evaluation index W(G) or evaluate the performance of the search algorithm of the process model.
[0034] <Conditions to be Satisfied by Process Model G> In the proposed method, a process model G that satisfies the following three conditions, Condition 1 to Condition 3, is given as input.
[0035] Condition 1: The process model G is a probabilistic model that follows a parameter θ that defines the generation probability of a trace σ, and the generation probability Pr(σ|G, θ) of a trace σεL can be calculated.
[0036] Condition 2: The parameter θ of the process model G is some prior parameter α θ The probability distribution Pr(θ|αθ ) and the posterior distribution Pr(θ|L,α θ ) the parameter θ can be sampled probabilistically.
[0037] Condition 3: The process model G itself has some prior parameter α G The probability distribution Pr(G|α G ) can be generated probabilistically.
[0038] There are various data formats for the process model G, but one that satisfies the above three conditions 1 to 3 is the Probabilistic Generative Process Model (PGPM) proposed in Reference 1. PGPM is a process model in which a process tree is converted into generative grammar. In the following, as an example, the process model G will be described as a PGPM.
[0039] It is noted in Reference 1 that PGPM satisfies the above condition 1. Furthermore, when considering a probabilistic model, it is usually constructed using a combination of general probability distributions (exponential distribution families), so if condition 1 is satisfied, condition 2 is also generally satisfied, and this also applies to PGPM. Furthermore, since a random generation method for process trees is also known (Reference 2), PGPM also satisfies condition 3.
[0040] However, the fact that the process model G is PGPM is just one example, and the proposed method is not limited to PGPM as long as it satisfies the above three conditions 1 to 3. For example, GDT-SPN (Reference 3), a process model based on Petri nets, also satisfies the above conditions 1 and 2, and therefore, if it also satisfies the above condition 3, it can be used as the process model G. In other words, if there is a method for generating Petri nets, GDT-SPN also satisfies the above condition 3, and it can be used as the process model G.
[0041] Note that the above condition 3 is a necessary condition for simultaneously evaluating simplicity. Therefore, for example, when simultaneously evaluating the three aspects of suitability, accuracy, and versatility, it is not necessary to satisfy the above condition 3.
[0042] <Calculation formula for evaluation index value W(G)> The evaluation index value W(G) for simultaneously evaluating the four aspects of the suitability, accuracy, versatility, and simplicity of the process model G can be calculated using the following formulas (1) to (5).
[0043] Here, s is a preset hyperparameter and represents the number of sampling times of the parameter θ. j ∝Pr(θ|L, α θ ) is the posterior distribution Pr(θ|L,α θ ) to the parameter θ j Furthermore, α G , α θ is a hyperparameter that is set in advance. In the case of PGPM, the probability of generating a trace follows a multinomial distribution, so that Pr(θ|L,α θ ) is expressed as a Dirichlet distribution, which is the conjugate prior of the multinomial distribution.
[0044] <<Explanation of Evaluation Index Value W(G)>> Hereinafter, an intuitive explanation will be given of how the evaluation index value W(G) simultaneously evaluates four viewpoints.
[0045] First, T in the above formula (2) is an index for evaluating the conformance and accuracy of the process model G. T is a parameter θ j The probability of generating any trace of the process model G when i |G, θ j) and is called the average log-likelihood. T is a negative value of the log-likelihood that indicates the goodness of fit between the process model G and the event log L, which is the observed data, and reaches its minimum value when the trace generation probability Pr(σ|G, θ) of the process model G matches the appearance rate of the trace included in the event log. Therefore, T reaches its minimum value when the process model G generates only the traces included in the event log L, and therefore T is an index that simultaneously evaluates the goodness of fit and accuracy.
[0046] Next, S in the above formula (3) is an index for evaluating the simplicity of the process model G. Pr(G|α G ) is the a priori parameter α G represents the probability that a process model G will be generated from the prior distribution of . Generally, the greater the number of elements in a process model G, the lower the probability of generating the process model G. This is because, for example, there are more process models with three activities than there are process models with two activities. This also applies to the random generation method of a process tree described in Reference 2, for example. Therefore, by using S, which indicates a smaller value as the generation probability of the process model G increases, it becomes possible to highly evaluate a process model G with a small number of elements.
[0047] Finally, V in the above equation (4) is an index used in combination with T to evaluate versatility. Let W'(G):=W(G)-S=T+V / |L|. In this case, W'(G) corresponds to the standard value of the Widely Applicable Information Criteria (WAIC) (Reference 4), which evaluates models based on the theory of Bayesian estimation. WAIC approximates a value called generalization loss, which indicates the closeness between the distribution of unknown true traces and the distribution of traces generated by a process model. WAIC can be said to indicate the predictive ability of unknown observations (i.e., unknown traces). The closeness to the distribution of true traces evaluated by generalization loss is consistent with the definition of versatility in model evaluation of process models, and WAIC can be said to be an index for evaluating versatility.
[0048] From the above, it can be said that W(G) uses the three indices T, S, and V to simultaneously evaluate four aspects: suitability, accuracy, versatility, and simplicity.
[0049] <<Method of Derivation of Formula for Calculating Evaluation Index Value W(G)>> We will now describe how the evaluation index value W(G) calculated by the above formulas (1) to (5) is derived from the theoretical framework of Bayesian estimation.
[0050] Bayesian estimation is a method in which, when observed data L and a process model G are given, a predictive distribution p * (σ|L, G, α G , α θ ) is a general term for methods of predicting unknown σ.
[0051] Hereinafter, the true distribution of the observed value σ is assumed to be q(σ). In this case, the generalization loss g, which is an important index in Bayesian estimation, is expressed by the following equation (7).
[0052] The generalization loss g itself cannot be obtained because it follows the unknown true distribution q(σ). However, it is known that a criterion called WAIC W'(G), which matches the generalization loss with its expected value, can be calculated (Reference 4). W'(G) can be calculated using the following equations (8) to (12).
[0053] Here, Z is a normalization coefficient. In equation (9), f(θ) = Pr(σ|G, θ). In addition, in the first term of the sum in equation (11), f(θ) = (logPr(σ i |θ) 2 , and in the second term, f(θ) = logPr(σ i |θ).
[0054] If an event log L, which is observed data, and a process model G are given and an integral can be calculated, it is possible to calculate the WAIC. However, in practice, analytical integral calculation is often difficult. Therefore, when calculating the WAIC, an approximate calculation based on a sampling method is often used.
[0055] s parameters θ 1 , ..., θ s is the posterior distribution Pr(θ|L, α θ ) can be sampled. In this case, the sampled s parameters θ 1 , ..., θ s Using E θ [f(θ)] can be approximated as in the following equation (13).
[0056] By using the approximation shown in the above formula (13), an approximate value of WAIC W'(G) can be obtained. This approximate value is the evaluation index value W(G) shown in the above formula (5).
[0057] As described above, the evaluation index value W(G) is naturally obtained as an approximation of the generalization loss based on Bayesian estimation, i.e., assuming that the model is evaluated using a predictive distribution. Therefore, the evaluation index value W(G) can be naturally derived without requiring any other assumptions for a process model that satisfies the three conditions 1 to 3 above. As a result, the evaluation index value W(G) that is smaller for a process model that is superior in the four aspects of fitness, accuracy, versatility, and simplicity is obtained. In this way, the evaluation index value W(G) can be used to simultaneously evaluate the four aspects of fitness, accuracy, versatility, and simplicity, simply by assuming that Bayesian estimation is performed.
[0058] A process model evaluation device 10 capable of evaluating the process model G using the above proposed method will be described below.
[0059] <Example of Hardware Configuration of Process Model Evaluation Apparatus 10> An example of the hardware configuration of the process model evaluation apparatus 10 according to this embodiment will be described with reference to Fig. 6. Fig. 6 is a diagram showing an example of the hardware configuration of the process model evaluation apparatus 10 according to this embodiment.
[0060] 6, the process model evaluation apparatus 10 according to this embodiment includes an input device 101, a display device 102, an external I / F 103, a communication I / F 104, a random access memory (RAM) 105, a read only memory (ROM) 106, an auxiliary storage device 107, and a processor 108. These pieces of hardware are connected to each other via a bus 109 so as to be able to communicate with each other.
[0061] The input device 101 is, for example, a keyboard, a mouse, a touch panel, a physical button, etc. The display device 102 is, for example, a display, a display panel, etc. Note that the process model evaluation apparatus 10 does not necessarily have to include at least one of the input device 101 and the display device 102, for example.
[0062] The external I / F 103 is an interface with an external device such as a recording medium 103a. Examples of the recording medium 103a include a CD (Compact Disc), a DVD (Digital Versatile Disk), an SD memory card (Secure Digital memory card), and a USB (Universal Serial Bus) memory card.
[0063] The communication I / F 104 is an interface for connecting to a communication network. The RAM 105 is a volatile semiconductor memory (storage device) that temporarily stores programs and data. The ROM 106 is a non-volatile semiconductor memory (storage device) that can store programs and data even when the power is turned off. The auxiliary storage device 107 is a non-volatile storage device such as a hard disk drive (HDD), a solid state drive (SSD), or a flash memory. The processor 108 is a variety of arithmetic devices such as a central processing unit (CPU) or a graphic processing unit (GPU).
[0064] 6 is just an example, and the hardware configuration of the process model evaluation device 10 is not limited to this. For example, the process model evaluation device 10 may have a plurality of auxiliary storage devices 107 or a plurality of processors 108, may not have some of the hardware shown in the figure, or may have various hardware other than the hardware shown in the figure.
[0065] <Example of Functional Configuration of Process Model Evaluation Apparatus 10> An example of the functional configuration of the process model evaluation apparatus 10 according to this embodiment will be described with reference to Fig. 7. Fig. 7 is a diagram showing an example of the functional configuration of the process model evaluation apparatus 10 according to this embodiment. The process model evaluation apparatus 10 is provided with a process model G to be evaluated and an event log L which is observed observation data.
[0066] 7 , the process model evaluation device 10 according to this embodiment includes an input unit 201, a sampling unit 202, an evaluation index value calculation unit 203, and an output unit 204. These units are realized, for example, by a process in which one or more programs installed in the process model evaluation device 10 are executed by the processor 108 or the like. The process model evaluation device 10 according to this embodiment also includes a hyperparameter storage unit 205. The hyperparameter storage unit 205 is realized, for example, by a storage area of the auxiliary storage device 107 or the like. However, the hyperparameter storage unit 205 may also be realized, for example, by a storage area of a storage device (e.g., a database server) communicably connected to the process model evaluation device 10.
[0067] The input unit 201 receives a given process model G and event log L.
[0068] The sampling unit 202 calculates the number of samplings s and the prior parameter α stored in the hyperparameter storage unit 205. θ and the event log L input by the input unit 201, and s parameters θ are calculated by the above equation (1). 1 , ..., θ s Sample the following.
[0069] The evaluation index value calculation unit 203 calculates the number of samplings s and the prior parameter α G the process model G and the event log L input by the input unit 201, and the parameters θ sampled by the sampling unit 202. 1 , ..., θ s The evaluation index value W(G) is calculated using the above equations (2) to (5).
[0070] The output unit 204 outputs the evaluation index value W(G) calculated by the evaluation index value calculation unit 203 to a predetermined output destination. Examples of the output destination include the display device 102, a storage area of the auxiliary storage device 107, and other devices or apparatuses connected in a communicable manner.
[0071] The hyperparameter storage unit 205 stores the number of sampling times s and the prior parameter α θ , α G Hyperparameters such as
[0072] <Process Model Evaluation Processing> The process model evaluation processing according to this embodiment will be described with reference to Fig. 8. Fig. 8 is a flowchart showing an example of the process model evaluation processing according to this embodiment. In the following, it is assumed that a process model G and an event log L are provided to the process model evaluation device 10.
[0073] First, the input unit 201 receives a given process model G and event log L (step S101).
[0074] Next, the sampling unit 202 calculates the number of samplings s and the prior parameter α θ and the event log L input in step S101, and s parameters θ are calculated by the above equation (1). 1 , ..., θ s is sampled (step S102).
[0075] Next, the evaluation index value calculation unit 203 calculates the number of sampling times s and the prior parameter αG the process model G and the event log L input in step S101, and the parameters θ sampled in step S102. 1 , ..., θ s Using these, the evaluation index value W(G) is calculated using the above formulas (2) to (5) (step S103).
[0076] Then, the output unit 204 outputs the evaluation index value W(G) calculated in step S103 to a predetermined output destination (step S104), thereby obtaining the evaluation index value W(G) that simultaneously evaluates the process model G from four viewpoints: suitability, accuracy, versatility, and simplicity.
[0077] For example, two process models G 1 , G 2 If you want to compare which of the two is better, 1 and the event log L are provided to the process model evaluation device 10 to calculate the evaluation index value W(G 1 ) is calculated, and then the process model G 2 and the event log L are provided to the process model evaluation device 10 to calculate the evaluation index value W(G 2 ) can be calculated. 1 ) and the evaluation index value W(G 2 ) and the two process models G 1 , G 2 You can find out which one is better.
[0078] <Summary> As described above, the process model evaluation device 10 according to this embodiment can calculate the evaluation index value W(G) that can simultaneously and fairly evaluate four perspectives (goodness of fit, accuracy, versatility, and simplicity) based on the theory of Bayesian estimation. Therefore, by using the process model evaluation device 10 according to this embodiment, it becomes possible to quantitatively evaluate how well a process model that represents a business process represents an actual business process from these four perspectives. Therefore, it becomes possible to support the analysis and understanding of business processes with the aim of improving the efficiency of operations in various businesses, for example.
[0079] The present invention is not limited to the above-described specifically disclosed embodiments, and various modifications, changes, and combinations with known technologies are possible without departing from the scope of the claims.
[0080] [References] Reference 1: Watanabe, A., Takahashi, Y., Ikeuchi, H., & Matsuda, K. (2022, October). Grammar-Based Process Model Representation for Probabilistic Conformance Checking. In 2022 4th International Conference on Process Mining (ICPM) (pp. 88-95). IEEE. Reference 2: Generating Artificial Data for Empirical Analysis of Control-flow Discovery Algorithms - A Process Tree and Log Generator. Bus. Inf. Syst. Eng. 61(6): 695-712 (2019). Reference 3: Rogge-Solti, A., van der Aalst, W. M., & Weske, M. (2014). Discovering stochastic petri nets with arbitrary delay distributions from event logs. In Business Process Management Workshops: BPM 2013 International Workshops, Beijing, China, August 26, 2013, Revised Papers 11 (pp. 15-27). Springer International Publishing. Reference 4: Watanabe, S., & Opper, M. (2010). Asymptotic equivalence of Bayes cross validation and widely applicable information criterion in singular learning theory. Journal of machine learning research, 11(12).
[0081] REFERENCE SIGNS LIST 10 Process model evaluation device 101 Input device 102 Display device 103 External I / F 103a Recording medium 104 Communication I / F 105 RAM 106 ROM 107 Auxiliary storage device 108 Processor 109 Bus 201 Input unit 202 Sampling unit 203 Evaluation index value calculation unit 204 Output unit 205 Hyperparameter storage unit
Claims
1. An evaluation device for evaluating a process model representing a business process, comprising: - an input unit that inputs a process model to be evaluated and an event log that is a log of traces representing one or more operations observed in the business process; - an evaluation value calculation unit that calculates, as an evaluation value of the process model, an approximation of the generalization loss between the true distribution of the trace when predicting the distribution of the trace from the process model by Bayesian estimation based on the process model and the event log.
2. The evaluation device according to claim 1, further comprising a sampling unit that samples s predetermined parameters based on the event log and a first prior parameter that defines a probability distribution followed by the parameters defining the generation probability of the trace by the process model, wherein the evaluation value calculation unit calculates an approximation of the generalization loss using the s parameters sampled by the sampling unit and a second prior parameter that defines the generation probability of the process model.
3. The evaluation device according to claim 1 or 2, wherein the evaluation value calculation unit calculates, as an evaluation value of the process model, an approximation of WAIC (Widely-Applicable Information Criteria) in which the generalization loss and the expected value match.
4. An evaluation method for evaluating a process model representing a business process, the method comprising: - an input procedure for inputting a process model to be evaluated and an event log that is a log of traces representing one or more operations observed in the business process; - an evaluation value calculation procedure for calculating, as an evaluation value of the process model, an approximation of the generalization loss between the true distribution of the trace when predicting the distribution of the trace from the process model by Bayesian estimation based on the process model and the event log.
Citation Information
Patent Citations
Process evaluation device and process evaluation program
JP2016122332A
Log Data Compliance
JP2023530104A
Semantic-aware rule-based recommendation for process modeling
US20230281484A1
Work process classification device, work process classification method and program
WO2022215205A1