Approximate consistency detection method for event stream and process model

Through the approximate alignment consistency detection method for event flow, the trajectory editing distance and frequency method sampling is used to construct a model trajectory set, reducing the computational complexity, and achieving efficient and accurate online consistency detection, solving the problem of large computing resources or large errors in existing methods.

CN120336864APending Publication Date: 2025-07-18GUILIN UNIV OF ELECTRONIC TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510031510.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

When the existing online consistency detection method is oriented towards event flow, there are problems such as high computing resource consumption or large error in detection results, making it difficult to achieve efficient and accurate detection.

Method used

The approximate alignment consistency detection method for event flow is used to sample representative sampling event logs through the trajectory editing distance and frequency method, and the model trajectory set is constructed using the optimal alignment algorithm, and the calculation complexity is reduced through the trajectory editing distance cost function, and the consistency is evaluated in combination with the approximate fit and trajectory integrity.

Benefits of technology

It significantly improves the accuracy and detection efficiency of approximate fitting, and can achieve efficient and accurate online consistency detection in the event stream, reducing computing resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336864A_ABST
    Figure CN120336864A_ABST
Patent Text Reader

Abstract

With increase of enterprise business complexity, business process consistency detection is more important. An existing detection method pays more attention to static event logs, has the problem of large calculation amount or large error, and cannot complete consistency detection of a dynamic event stream and a predefined model. Therefore, the approximate alignment detection method for the event stream is provided. The method comprises the following steps: firstly, acquiring a sampling log of a recorded log from an event stream; thirdly, calculating the optimal alignment of the sampling log trajectory and the process model, establishing a model trajectory set, and replacing the original model; and finally, designing an approximate consistency detection method based on prefix alignment, comparing the event flow trajectory with the model trajectory set, and obtaining a comparison result, an approximate fitting degree and trajectory integrity with less calculation amount, thereby evaluating the consistency of the event flow trajectory and the model trajectory set. Experiments show that on five real data sets, compared with an existing approximate consistency detection method, the method has the advantage that the accuracy of the approximate fitting degree and the calculation efficiency are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of business process mining, and is an approximate consistency detection method for event streams and process models. Background Art

[0002] With the continuous increase in the business scale and the gradual complexity of business requirements of modern enterprises and organizations, their business processes have become increasingly complex, and the diversity of business process execution behaviors has also increased. Effective business process management (BPM) can help enterprises and organizations better manage various business processes, reduce operation errors, and lower risks. Process mining is an interdisciplinary field of business process management and data mining, aiming to extract valuable business information from the log data of information systems to optimize the business process management of enterprises. Process mining mainly assists enterprises or organizations in optimizing business processes and improving process management and operation efficiency through three applications: discovering process models, performing consistency detection, and process enhancement. Among them, business process consistency detection, as a key link to ensure the running quality of business processes, is becoming increasingly important.

[0003] Consistency detection compares the observed business process behaviors with the predefined business process model, finds the differences and commonalities between the two, to determine whether the observed behaviors are consistent with their process model, and discovers existing problems.

[0004] In recent years, scholars have proposed many consistency detection methods, including rule detection methods, Token replay methods, and alignment methods. The rule detection method uses the trace rules in the event log extracted from the business process model to test the consistency between the event log and the process model. This method can detect the consistency of a certain segment of a trace, but is not applicable to the entire trace. In order to perform consistency detection on the entire business process trace, the Token replay method replays the trace of the event log in the business process model based on Petri nets, counts the consumed and remaining Tokens, and thus evaluates the consistency between the event log and the process model. This method can perform consistency detection on the entire business process trace, but cannot accurately locate the deviation between the trace of the event log and the process model, resulting in a relatively large fitness obtained from the consistency detection. In order to perform more accurate consistency detection, the alignment method discovers the deviation between the event log trace and the process model by comparing (Align) the event log trace with the process model, and then accurately locates the activities and positions where the deviation occurs. The detection of this deviation can be transformed into a comparison between the trace of the event log and the trace generated by replaying the process model. Given that the alignment method can obtain accurate consistency detection results, it is recognized as the standard method for consistency detection.

[0005] In recent years, researchers have pointed out that the research object of business process mining urgently needs to shift from event logs to stream events. However, traditional business process consistency detection methods are all offline, that is, first determine the events within the time range of interest to construct an event log, and then detect the event log against the process model. In an offline environment, the longer the time required from event log data extraction to analysis and then to decision-making, the lower the value of the event log data. In some cases, such as fraud detection, autonomous driving, or health monitoring scenarios, it is inappropriate to make decisions solely based on the analysis of stagnant data. Therefore, real-time detection of event streams is required.

[0006] Currently, online consistency detection for event streams has become an important research direction in the field of business process mining. In the consistency detection of business processes oriented to event streams, the data to be detected is an unbounded and dynamic event stream, rather than a static event log. When there are differences between the event sequence in the event stream and the predefined business process model, online consistency detection needs to promptly capture these differences. In this way, business process managers can timely discover and correct the adverse effects brought by these differences. To achieve accurate online consistency detection between the event stream and the business process model, researchers use the prefix alignment method to evaluate the consistency between the traces in the event stream and the process model. This method only considers the event sequence (i.e., trace prefix) in the event stream from the start of the execution of the business process to the current moment, rather than the entire event log. This method can calculate the trace alignment when the business process starts to execute, without having to wait until the end of the entire event sequence to start the calculation, thus enabling real-time detection of the differences between the event stream and the process model. However, when pursuing accurate alignment results, online consistency detection methods inevitably require a lot of computational costs, which limits its application in consistency detection. In addition, the detection results of other online consistency detection methods except prefix alignment have a large gap with the alignment detection results, so that accurate deviation analysis cannot be provided. Generally speaking, the current online consistency detection methods mainly face two problems. One is that although the prefix alignment method can obtain accurate results, it consumes a large amount of computing resources. The other is that although the approximate consistency detection method can obtain results in a short time, the error is large. Therefore, the detection efficiency of these methods needs to be further improved. Summary of the Invention

[0007] In view of the above problems, the present invention proposes an approximate alignment consistency detection method for event streams. This method compares the selected trajectories in the event stream with the trajectory set of a predefined business process model to obtain an approximate online prefix alignment result. This method uses a representative trajectory set instead of the business process model, thereby effectively reducing the time complexity of the online consistency detection algorithm while ensuring relatively accurate online consistency detection results.

[0008] The technical content of the present invention is as follows:

[0009] Step 1: Sample the event log L by calculating the normalized edit distance of trajectories. s , and then, use the frequency method for sampling to obtain the final sampled event log. The frequency method improves the quality of the sampled event log by selecting the most frequent behaviors in the event log to ensure the representativeness of the sampled event log. The definition of the normalized edit distance between σ1 and σ2 in the event log trajectory is shown in Equation (1). Generally, a reasonable sampling method can make the sampled event log as representative of the original event log as possible. This is because a reasonable sampling method generally selects trajectories according to the feature distribution in the original event log, such as retaining the diversity and corresponding proportions of different trajectories, activity frequencies, and trajectory lengths. This way can ensure that the sampled event log is consistent with the original event log in statistical characteristics.

[0010] To implement Step 1 of the present invention example and realize the sampled event log, Algorithm 1 is designed, and the pseudocode is as follows:

[0011] Step 2: Based on the obtained sampled event log L s , use the optimal alignment algorithm to obtain the model trajectory set M s of L T and the process model N. That is, calculate the best alignment of the trajectories of the business process model and the sampled event log, and construct the corresponding model trajectory set from the process model trajectories mapped by the alignment result. Among them, the alignment method is as follows. For the alignment of the event log and the process model, let A be a set of activities, and σ ∈ A * be a trajectory on A, and M = (P, T, F, α, m i , m f ) be a Petri net on A. The alignment γ ∈ (A >> × T 》 ) between the trajectory σ and the net N is a movement sequence that satisfies the following conditions. (1) The alignment between the trace in the event log and the process model is a sequence of moves γ. For any move (a, t) ∈ γ in the alignment, the present invention defines (a, t) as follows: 1) If a ∈ A and t = >>, then (a, t) is called a log move; 2) If a = >> and t ∈ T, then (a, t) is called a model move; 3) If a ∈ A, t ∈ T and α(t) = a, or a = >> and t = τ, then (a, t) is called a synchronous move; 4) Otherwise, it is called an illegal move. Denote Γ σ,N to represent the set of all alignments between the trace σ and the Petri net model N. (2) π1(γ) ↓A = σ. The projection (ignoring >>) of the sequence formed by the first elements of each move in the move sequence γ onto A produces the trace; (3) The projection (ignoring >>) of the sequence formed by the second elements of each move in the move sequence γ onto the model T produces a complete firing sequence, which causes the marking of the model T to change from m i to m f .

[0012] To implement step two of the example of the present invention, to obtain the model trace set, Algorithm 2 is designed, and the pseudocode is as follows:

[0013] Step three: Obtain the model trace set M T After that, by using the trace edit distance cost function d(σ1, σ2) to replace the standard cost function c(b), the calculation overhead of the fitness between the event log and the process model is reduced. The standard cost function c(b) and the trace edit distance cost function d(σ1, σ2) are introduced as follows. Standard cost function c(b): Let A be a set of activities, and A* represent all possible sequences of activities. σ ∈ A* is a trace on A, and N = (P, T, F, α, m i , m f ) is a Petri net on A. The standard cost function c(b) can be expressed as the cost value of move assignment, where b = (a, t) is an arbitrary move in the alignment between the trace and the process model. (1) c(b) = 1. If a ∈ A and t = >>, then (a, t) is called a log move; (2) c(b) = 1. If a = >>, t ∈ T and α(t) ≠ τ, then (a, t) is called a model move; (3) \(c(b)=0\). If \(a\in A\), \(t\in T\) and \(\alpha(t)=a\) or \(a=\gg\) and \(\alpha(t)=\tau\) or \(a = \tau\) and \(t=\gg\), then \((a,t)\) is called a synchronous move. (4) \(c(b)=+\infty\). If \(a=\gg\) and \(t=\gg\) or \(a\in A\) τ , \(t\in T\) and \(\alpha(t)\neq a\), then \((a,t)\) is called an illegal move. The trajectory edit distance cost function \(d(\sigma_1,\sigma_2)\): Assume that \(\sigma_1\in A^*\) is a trajectory, and the corresponding activity sequence of the transition sequence \(\langle t_1,t_2,\cdots,t\) n \((n\in N\) * ) is denoted as \(\langle\alpha(t_1),\alpha(t_2),\cdots,\alpha(t\) n )\rangle\) which is denoted as the trajectory \(\sigma_2\). The edit distance cost function of the trajectory \(td(\sigma_1,\sigma_2)\to N\) represents the minimum number of edit operations required to convert \(\sigma_1\) into \(\sigma_2\), where \(N\) is the set of natural numbers. The edit operations here only include insertion and deletion operations, not replacement operations. In the alignment of the trajectory and the process model, each asynchronous move is equivalent to an edit operation, where the insertion operation corresponds to the model move and the deletion operation corresponds to the log move. Therefore, \(d()\) returns the same cost value as the control flow alignment. For example, \(d(\langle a,b\rangle,\langle a,d\rangle)=2\), and its edit operations are one insertion operation \((b,\gg)\) and one deletion operation \((\gg,d)\).

[0014] Step 4: Calculate the approximate online fitness of the trajectory prefix in the event stream and obtain an approximate online fitness through weighted summation. As the business process instance is executed, the evolution of the trajectory prefix tends to be stable, so the number of model trajectories participating in the online consistency detection gradually decreases. When the execution of the business process instance is approaching the end, the event sequence in the trajectory prefix has been basically determined. At this time, only 1 model trajectory is required to meet the calculation needs of the approximate online consistency detection. Therefore, as the business process is executed, the number of model trajectories used to calculate the online fitness in the consistency detection will gradually decrease from multiple at the beginning to 1. The calculation of the online fitness is shown in formula (2). where \(n\) is the ratio of the length of \(\sigma\) to the length of the sampled trajectory with the minimum edit distance corresponding to it and coming from the sampled event log, and \(\minEdDis(\sigma,M\) T , \(n)\) represents the \(n\) model trajectories with the minimum edit distance between \(M\) T and \(\sigma\), and \(len(\sigma)\) represents the length of the trajectory \(\sigma\).

[0015] However, in the scenario of dynamic and unbounded event streams, the consistency detection of business processes not only requires real-time detection of the fitness of the current event trace, but also needs to judge the completeness of the current trace. Using only the fitness to evaluate the online consistency detection results is relatively one-sided and may lead to misleading consequences. Therefore, when conducting the consistency detection of event stream traces, in order to make a more accurate comparison of event stream traces, it is also necessary to judge the progress of the trace and predict the events that may be missing in the current event stream trace. Thus, based on obtaining the fitness between the event stream and the process model, the present invention adds an evaluation of the completeness of the event stream trace.

[0016] The trace completeness can predict the number of events that are currently unobserved in the event stream trace and, at the same time, judge the progress of the executing business process instance represented by this trace. For example, Figure 2 the business process shown in case2 has ended during the detection, so it can be determined that the trace represented by case2 is complete, that is, the completeness is 1. On the contrary, the business process shown in case3 has not ended at the current time point, which means that there may still be events that have not occurred in the instance trace of this business process, and these events need to be inferred during the detection process. Naturally, the completeness of case3 is less than 1. The following gives the relevant definitions of trace completeness.

[0017] Definition (Trace Completeness) Given a trace σ of an event stream log and a set M of model traces T , an integer n ∈ N + , then the completeness of σ is calculated as follows.

[0018]

[0019] To implement Step 4 of the instance of the present invention and use the set of model traces to achieve online approximate consistency detection between the event stream and the business process model, the present invention designs Algorithm 3 for online consistency detection, and the pseudocode is as follows.

[0020] The purpose of the present invention is to solve the problem that the existing online consistency detection methods need to consume a large amount of computing resources to obtain accurate detection results, and there is a large error in the detection results for fast online consistency detection. The present invention aims at the above technical problems existing in the prior art and brings some creative technical effects after solving the problems. The specific description is as follows:

[0021] (1) The method CCAES proposed by the present invention has achieved a significant improvement in the accuracy of the approximate fitness and can achieve online consistency detection for event streams.

[0022] (2) Meanwhile, compared with the mainstream prefix-alignment-based consistency detection methods, the method CCAES has achieved significant improvements in detection efficiency and performance.

[0023] The present invention proposes an approximate alignment consistency detection method for event streams. First, an event stream log sampling method that combines trace edit distance and frequency is used to obtain a sampled event log; then, an alignment method is used to obtain the optimal alignment between the sampled event log trace and the business process model, and the model traces obtained from the optimal alignment constitute the model trace set. Finally, based on the model trace set, online approximate consistency detection is performed on the event log and the process model to obtain the approximate online fitness and trace integrity. Experiments on 5 real event log datasets show that, compared with the existing online approximate consistency detection methods, the method CCAES of the present invention is more efficient than the current online alignment methods. Therefore, the method CCAES of the present invention provides a practical solution for the existing online consistency detection of business process management. Description of the Drawings

[0024] Figure 1 It is the schematic diagram of the approximate alignment consistency method of the present invention.

[0025] Figure 2 It is the online consistency detection diagram of the event stream of the business process of the present invention.

[0026] Figure 3 It is the error diagram of the approximate fitness between the event log and the process model in the present invention.

[0027] Figure 4 It is the Spearman correlation coefficient diagram of the fitness in the present invention.

[0028] Figure 5 It is the change diagram of the Spearman correlation coefficient of the integrity in the present invention.

[0029] Figure 6 It is the average time diagram of the consistency detection events in the present invention.

[0030] Figure 7 It is the memory consumption diagram of the online consistency detection in the present invention. Detailed Embodiment

[0031] To make the objectives, technical solutions, and advantages of the present invention clearer, the following further elaborates on the present invention in detail with reference to specific examples and the accompanying drawings.

[0032] To implement step one of the example of the present invention and obtain the sampled event log, Algorithm 1 is designed, and the pseudocode is as follows:

[0033] In Algorithm 1, lines 2 - 17 calculate the edit distance between the new trajectory and each trajectory in L in the event stream. If the obtained minimum normalized edit distance is greater than the set edit distance threshold he, the new trajectory is added to L s , otherwise the next trajectory is calculated. When the sampling of the event stream stops, all events in the event stream are recorded in the event log L. Lines 18 and 19 select the n trajectory variants with the highest occurrence probability in the event log L as the frequency - based sampling event log, and then merge it with Ls. After analysis, the time complexity of selecting the log candidate traces in lines 2 - 17 is O(m + m / k), where m represents the number of events in the sampled event stream, and k represents the average number of activities of the event stream trajectories. The time complexity of using the frequency - based sampling in line 19 is O(nr), where n represents the number of trajectories sampled using the frequency - based method, and r represents the number of trajectory variants in L|TV s |. Thus, the time complexity of Algorithm 1 is O((m + m / k + nr). L |.

[0034] To implement Step 2 of the example of the present invention, to obtain the model trajectory set, Algorithm 2 is designed, and the pseudocode is as follows:

[0035] In Algorithm 2, the process model N is obtained by mining the event log L using the inductive mining algorithm. Lines 2 - 5 traverse the sampled event log L s , and obtain the model trajectory corresponding to each sampled trajectory. All the model trajectories form the model trajectory set M T . After analysis, the time complexity of this algorithm is O(n). Where n = |L s |.

[0036] In Step 3 of the example of the present invention, in order to reduce the calculation overhead of the fitness calculation between the event log and the process model, the present invention uses the trajectory edit distance cost function d(σ1, σ2) to replace the standard cost function c(b). This is because in approximate consistency detection, by using the model trajectory set to replace the complex process model, the originally complex alignment calculation is simplified to the calculation of the trajectory edit distance in Definition 9. This method significantly reduces the overhead required for calculating the fitness.

[0037] To more accurately evaluate the consistency between the event log and the process model, the present invention defines the concept: approximate fitness. This fitness is measured by the matching degree between the business process model behavior characterized by the model trajectory set and the event log trajectories. This approximate fitness can reduce the computational complexity while providing sufficiently accurate evaluation results.

[0038] For each σ ∈ L s, first compute its optimal alignment γ opt , and then project γ opt onto the model T, π2(γ opt ) ↓T Add the obtained model trace to the model trace set M T . Next, for the remaining traces σ ∈ L c (i.e., L c = L - L s ), use M T to calculate their approximate fitness appFitness(σ, M T ).

[0039] Since the edit distance cost function d(σ, σ m ) in Definition 9 is used in the approximate alignment instead of the standard cost function c(b), the alignment cost calculation can be transformed into the calculation of the minimum edit distance between σ and σ m . Let K(γ opt ) be the optimal alignment cost between σ and the process model N. If π2(γ opt ) ↓A = σ m , it means that d(σ, σm) = K(γ opt ), and further appFitness(σ, MT) = fitness. If π2(γ opt ) ↓A ≠ σ m , it means that d(σ, σm) > K(γ opt ), then appFitness(σ, MT) < fitness. Therefore, in the approximate alignment process using M T , if M T contains more model traces of the optimal alignment with the process model in the event log, it means that M T can represent the process model more effectively , will improve the accuracy of the approximate fitness in consistency detection.

[0040] According to the above analysis, setting a higher sampling ratio and a smaller edit distance threshold can generate a more appropriate model trace set, obtain a more accurate approximate alignment result, and thus improve the accuracy of approximate consistency detection. (However) However, although increasing the sampling ratio and decreasing the edit distance threshold can improve the accuracy of approximate consistency detection, it usually brings higher computational overhead. Therefore, setting a reasonable sampling ratio and edit distance threshold is very important for improving the computational efficiency of the approximate consistency detection method.

[0041] In step four of the embodiments of the present invention, online consistency detection evaluates the online fitness by comparing the difference degree between the current trace in the event stream and the expected process model behavior. Figure 2 Fig. Figure 2 shows a schematic diagram of online consistency detection for business processes oriented to event streams.

[0042] For example, Figure 2 in case 1, the case has been completed at the current moment, so the obtained online consistency detection result can represent the fitness of the complete trace. Although case 2 starts only from the beginning of the detection, the process has ended at the current moment, so its online consistency detection can also obtain the fitness of the complete trace. However, case 3 has not ended, and its online fitness will be continuously updated according to the changes in the event stream.

[0043] Currently, accurate online alignment methods use the trace prefix in the event stream as the detection object, calculate the optimal alignment by finding the differences between the trace prefix and the process model, and then obtain the online fitness. Although this method can accurately locate the differences between the trace prefix and the process model, when dealing with a large number of events in the event stream and multiple trace prefixes need to be detected simultaneously, it requires a large amount of computational effort and running memory. At the same time, since the event stream changes continuously with the execution of the business process, in order to more accurately handle the evolution of the trace prefixes included in the event stream, in the real-time consistency detection process, n traces closest to the current trace prefix σ are selected from the model trace set, the approximate online fitness between these traces and the trace prefix is calculated, and an approximate online fitness is obtained through weighted summation. As the execution of the business process instance progresses, the evolution of the trace prefix becomes more stable, so the number of model traces required for online consistency detection gradually decreases. When the execution of the business process instance is approaching the end, the event sequence in the trace prefix is basically determined, and at this time, only 1 model trace is required to meet the computational needs of approximate online consistency detection. Therefore, as the business process is executed, the model traces used to calculate the online fitness in the consistency detection will gradually decrease from multiple at the beginning to 1. In order to efficiently obtain the online fitness of the trace prefix in the event stream of the business process, an approximate online fitness is defined. The relevant definitions of the approximate online fitness are given below.

[0044]

[0045] However, in the scenario of dynamic and unbounded event streams, the consistency detection of business processes not only needs to detect the fitness of the current event trace in real time, but also needs to judge the completeness of the current trace. Using only the fitness to evaluate the results of online consistency detection is relatively one-sided and may lead to misleading consequences. Therefore, when conducting the consistency detection of event stream traces, in order to make a more accurate comparison of event stream traces, it is also necessary to judge the progress of the trace and predict the events that may be missing in the current event stream trace. Thus, based on obtaining the fitness between the event stream and the process model, the present invention adds an evaluation of the completeness of the event stream trace.

[0046] The trace completeness can predict the number of events that are currently unobserved in the event stream trace, and at the same time judge the progress of the executing business process instance represented by this trace. For example, Figure 2 the business process shown in case2 in has ended during the detection, so it can be determined that the trace represented by case2 is complete, that is, the completeness is 1. On the contrary, the business process shown in case3 has not ended at the current time point, which means that there may still be events that have not occurred in the instance trace of this business process, and these events need to be inferred during the detection process. Naturally, the completeness of case3 is less than 1. The following gives the relevant definition of trace completeness.

[0047] Definition (Trace Completeness) Given a trace σ of an event stream log and a set M of model traces T , an integer n ∈ N + , then the calculation of the completeness of σ is as follows.

[0048]

[0049] To implement step four of the instance of the present invention, an online approximate consistency detection of the event stream and the business process model is realized by using the set of model traces. The present invention designs algorithm 3 for online consistency detection, and the pseudocode is as follows.

[0050] In algorithm 3, line 1 obtains the event e in the event stream S. Lines 3-5 update the corresponding case trace trace according to e. Lines 6-9 traverse the set L of traces of the sampled event log s , and obtain the n model traces that are most similar to the case trace trace. Lines 10-11 calculate the approximate fitness and completeness of trace. Line 12 returns the online fitness fitness and completeness completeness of the trace corresponding to the current event e, and the algorithm ends. After analysis, the time complexity of selecting similar model traces in lines 1-9 is O(m). Where m = |L s|The time complexity of calculating the fitness and completeness in lines 10-11 is O(n), where n represents the number of similar model traces. Therefore, the time complexity of Algorithm 3 is O(m + n).

[0051] An approximate consistency detection system for event streams and process models provided by an embodiment of the present invention includes:

[0052] An event stream log sampling module for selecting a sampled event log in the event log that can represent the recorded event log;

[0053] An approximate consistency detection module for the prefixed-aligned event log and the process model, which uses a set of model traces to replace the business process model to reduce the computational complexity of consistency detection and reduce the fitness error;

[0054] A calculation module for the approximate fitness and completeness of traces in the event stream to more accurately evaluate the consistency between the event stream and the process model;

[0055] To prove the creativity and technical value of the technical solution of the present invention, this part is an application embodiment of the technical solution of the claims on a specific product or related technology.

[0056] Figure 1 Taking the overall schematic diagram of the present invention as an example, the consistency detection method of the present invention can be implemented in two steps: constructing a set of model traces through the optimal alignment of the sampled event log and the process model, and performing online consistency detection of the trace prefix in the event stream and the process model. The specific process is as follows. Denote the event stream being observed as S, the predefined business process model as N, and the sampled event log as L s , s and the set of model traces obtained from the optimal alignment of L T and N as M s . Suppose the currently observed event is e(cid, a), where cid is the case ID and a is the activity executed by e. First, add the event e to the trace trace with the case ID cid (initially trace is empty); then, calculate the edit distance between the trace trace and each sampled trace in the log L s , and use the sampled trace in L T with the minimum edit distance from trace to predict the progress of trace; then, select n model traces in M to calculate the approximate fitness and completeness of the trace trace, where n is determined by the ratio of the length of trace to the length of the sampled trace with the minimum edit distance corresponding to it. Finally, use the approximate fitness and completeness as the approximate consistency detection result of the trace prefix represented by the current event e and N to analyze the consistency between the traces in the event stream and the business process model.

[0057] An approximate consistency detection method for an event stream and a process model provided by an application embodiment of the present invention is applied to a computer device. The computer device includes a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor executes the steps of the approximate consistency detection method for the event stream and the process model.

[0058] An approximate consistency detection method for an event stream and a process model provided by an application embodiment of the present invention is applied to an information data processing terminal, and the information data processing terminal is used to implement the approximate consistency detection system for the event stream and the process model. It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware part can be implemented using dedicated logic; the software part can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated design hardware. Those of ordinary skill in the art can understand that the above devices and methods can be implemented using computer-executable instructions and / or included in processor control code, for example, such code is provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuits of programmable hardware devices such as very large scale integrated circuits or gate arrays, semiconductors such as logic chips and transistors, or field programmable gate arrays and programmable logic devices, can also be implemented by software executed by various types of processors, or can be implemented by a combination of the above hardware circuits and software, such as firmware. The embodiments of the present invention have achieved some positive effects during the research and development or use process, and indeed have great advantages compared with the prior art. The following content will be described in combination with the data, charts, etc. of the experimental process.

[0059] (1) Experimental setup

[0060] The present invention uses 5 publicly available business process event logs: RequestForPayment[https: / / data.4tu.nl / datasets / a6f651a7-5ce0-4bc6-8be1-a7747effa1cc / 1], PermitLog[https: / / data.4tu.nl / datasets / db35afac-2133-40f3-a565-2dc77a9329a3 / 1], DomesticDeclarations[https: / / data.4tu.nl / datasets / 6a0a26d2-82d0-4018-b1cd-89afb0e8627f / 1], BPIC2012[https: / / www.win.tue.nl / bpi / 2012 / challenge.html] and CCC19[https: / / doi.org / 10.4121 / uuid:c923af09-ce93-44c3-ace0-c5508cf103ad.(2019-02-14).] are used as datasets to conduct experiments. Among them, RequestForPayment (abbreviated as RP), PermitLog (abbreviated as PL), and DomesticDeclarations (abbreviated as DD) belong to the event logs of BPIC2020. CCC19 is a publicly available event log dataset in the Conformance Checking Challenge in 2019. The characteristic information of each dataset is shown in Table 1.

[0061] Table 1 Characteristic information of real event log datasets

[0062] In order to obtain an event stream that can better simulate the actual business scenario than a static event log, the present invention designs a simple event stream simulator that uses a static event log to generate an event stream similar to that generated by a stream processing engine (such as Apache Kafka or Flink). When in use, multiple traces are randomly selected from each event log dataset in Table 2 first, and then the events in these traces are read to construct multiple event queues. At each moment, the event stream emulator randomly selects a queue and executes the head event of the queue until all the events in all event queues are executed. In this way, the event stream simulator can generate a dynamic and continuous event stream, which can more accurately simulate the execution behavior of business processes in the actual business scenario.

[0063] The experiment of the present invention requires setting three parameters: the behavior filter threshold ht, the trace sampling ratio sr, and the edit distance threshold he. The process discovery algorithm used is the Inductive Miner-InFrequency (abbreviated as IMi). Compared with the ordinary inductive mining algorithm (IM), IMi introduces infrequent behavior filters to filter out infrequent behaviors. Among them, IMi uses the set ht, sr is the proportion of trace variants sampled from the event log, and the edit distance threshold he is used to determine the similarity between the sampled traces.

[0064] The error of the approximate consistency detection method is the difference between the approximate fitness and the actual fitness. The smaller the error value, the closer the approximate fitness is to the actual fitness, and the higher the accuracy of the approximate alignment. The calculation method is as shown in Equation (4), where accFitness is the accurate fitness using the alignment method, and fitness is the approximate fitness calculated by the approximate alignment.

[0065] error = |accFitness - fitness| (4)

[0066] Since the fitness and completeness of the online consistency detection change with the change of the event stream, the present invention uses the Spearman's rank correlation coefficient to evaluate the relationship between the detection results. Among them, a Spearman's rank correlation coefficient of 1 indicates a perfect positive correlation, -1 indicates a perfect negative correlation, and 0 indicates no correlation. That is, the closer the Spearman's rank correlation coefficient is to 1, the more similar the detection result records of the two online consistency detections are. The calculation method of the Spearman's rank correlation coefficient ρ is as shown in Equation (5)

[0067]

[0068] where d i is the difference in the values of the i-th event representing the detection result, and n represents the number of events in the event stream.

[0069] (2) Experimental Results and Analysis

[0070] Error Analysis of Approximate Consistency Detection

[0071] To verify the effectiveness of the proposed approximate online fitting degree method of the present invention, five datasets in Table 2 are used in this experiment to conduct an offline consistency detection experiment. By comparing the approximate fitting degree errors error between the method CCAES of the present invention and the approximate alignment, the effectiveness of the approximate consistency detection method of the present invention is evaluated. To avoid sampling overly similar sampling trajectories and control the number of sampling trajectories, the present invention will set the edit distance threshold he = 0.5, the trajectory sampling ratio sr = 0.05, and use the sampling event log combining the trajectory edit distance and frequency to construct the model trajectory set. For each consistency detection algorithm, the present invention runs 5 times to obtain the corresponding 5 approximate fitting degree errors error, and then averages them. The obtained average value is used as the final error error of the algorithm as the experimental result. The experimental results are as Figure 3 shown. The abscissa in the figure is the 5 datasets, and the ordinate is the error error of the approximate fitting degree between the event log and the process model.

[0072] As Figure 3 shown, in the 5 datasets in Table 2, when using the same model trajectory set, the method CCAES of the present invention has a smaller error value than the approximate alignment, and its corresponding approximate fitting degree is more accurate, thus verifying the effectiveness of the method CCAES of the present invention in calculating the fitting degree between the event log and the process model. In the event log DD, the error values of the method CCAES and the approximate alignment are very close, so the difference in their approximate fitting degrees is small. This is because the number of trace variants in the event log DD is small and the similarity between trace variants is high, and the fitting degrees between each trace and the process model do not differ much, so that the approximate fitting degree can be accurately obtained only by selecting the most similar trace. Compared with the approximate alignment, the method CCAES can obtain more accurate consistency detection results when calculating the approximate fitting degree.

[0073] From Figure 4 it can be seen that the approximate fitting degree errors of the method CCAES proposed by the present invention on the datasets RP and DD are smaller than those on the datasets PL, CCC19, and BPIC2012. The main reason for this phenomenon is that the number of activity types in the datasets PL, CCC19, and BPIC2012 is more than that in the datasets RP and DD, and the average trace length is greater than that in the datasets RP and DD. Compared with the datasets RP and DD, the corresponding business process models of the datasets PL, CCC19, and BPIC2012 are more complex, so it is more difficult to include as many optimal alignment model trajectory sets as possible when sampling the event log. Therefore, the approximate fitting degree errors obtained by the method CCAES proposed by the present invention on the datasets PL, CCC19, and BPIC2012 are relatively large.

[0074] Verifying the effectiveness of the event log sampling method

[0075] To verify the effectiveness of considering both the edit distance and the trace frequency in event log sampling, the present invention uses an efficient event stream sampling method that takes both factors into account. By setting different trace sampling ratios sr and edit distance thresholds he, different sampled event logs are obtained. Then, the corresponding model trace sets are obtained through alignment, and the approximate fitness error error between the event log and each model trace set is calculated using an approximate online alignment algorithm. Finally, the effectiveness of the event log sampling method is evaluated by comparing the magnitudes of the errors error. The smaller the error, the more effective the sampling method. In this experiment, each method undergoes five offline approximate alignment detections in all event logs to obtain 5 errors error, and their average value is taken as the final error. The experimental results are shown in Table 2.

[0076] Table 2 Errors of the approximate values of the consistency detection

[0077] The rightmost column of the data in Table 2 shows that, except for CCC19, in the remaining 4 publicly available datasets, the sampling method of the event stream that considers both the edit distance and the trace frequency can more effectively obtain a more accurate approximate fitness. This is because the trace frequency plays a key role in sampling the event log. However, when comparing the data in the rightmost column of Table 2 with the experimental results obtained by the trace frequency sampling method, the method that considers both the edit distance and the trace frequency yields a smaller error than the method that only considers the trace frequency. It can be seen that relying solely on the trace frequency for sampling may cause traces with low frequencies and large differences from each other to be ignored, thereby making the sampled event log unable to effectively reflect the characteristics of the original log. By calculating the edit distance, these low-frequency traces can be effectively sampled, thus avoiding the loss of important information. In the CCC19 dataset, since the number of traces is 20 and the number of trace variants is also 20, the frequencies of all 20 traces are 1. Therefore, using a sampling ratio from 0.03 to 0.07 is actually equivalent to randomly selecting one trace for sampling. Therefore, in this case, the method that comprehensively considers the edit distance and the trace frequency does not achieve the desired effect as expected. Compared with the event stream sampling methods that only consider the edit distance or the trace frequency separately, the model trace set obtained by the sampling method of the present invention can achieve a more accurate fitness. This indicates that considering both the trace edit distance and the trace frequency can more effectively sample the event log and obtain a more accurate approximate alignment result. Thus, it shows that the trace sampling method proposed in the present invention, which considers both the trace frequency distribution and the individual differences of traces, is effective in event stream sampling.

[0078] To obtain a reasonable trajectory sampling ratio and edit distance threshold for facilitating subsequent experiments, this experiment uses the grid method to determine these two parameters. First, construct parameter combinations of the trajectory sampling ratio sr and the edit distance threshold he

[0079] As can be seen from Table 2, the fitting error errors obtained using a trajectory sampling ratio of 0.03 are relatively large. If the trajectory sampling ratio continues to be decreased, it will not be suitable for consistency detection. If the sampling ratio exceeds 0.1, more additional computational effort will be required. At the same time, the error obtained using only an edit distance threshold of 0.7 is relatively large. If a higher edit distance threshold is selected, the error will further increase; while if a threshold lower than 0.3 is selected, although the error may be small, it will select redundant sampled event logs, resulting in an increase in the computational effort for sampling and consistency detection. Therefore, the present invention selects one from {0.03, 0.05, 0.08, 0.1} as the numerical trajectory sampling ratio sr, and thus selects one numerical value from {0.3, 0.5, 0.8} as the edit distance threshold he to form different parameter combinations, and performs event log sampling to obtain different sampled event logs. Then, calculate the set of model trajectories obtained by aligning these sampled event logs with the process model. Next, use the approximate online alignment algorithm to calculate the approximate fitting error error between the event log and each set of model trajectories. Finally, evaluate the effectiveness of different parameter combinations by comparing the magnitudes of the errors. The smaller the error, the more reasonable the parameter combination of the sampling ratio and the edit distance threshold. Since the number of trajectories in the dataset CCC19 is small, in the parameter evaluation experiment of the event log sampling method, the method CCAES of the present invention does not use the dataset CCC19. The method CCAES of the present invention first uses other datasets in Table 2 except CCC19, performs five consistency detections to obtain 5 errors error, and then takes their average value as the final error. The experimental results are shown in Table 3.

[0080] Table 3 Errors of different combinations of the parameter sampling rate and the edit distance threshold in CCA

[0081] As can be seen from Table 3, as the trajectory sampling ratio sr increases and the edit distance threshold he decreases, the approximate fitting error error becomes smaller and smaller, indicating that the approximate fitting becomes more and more accurate. It can be concluded that a higher trajectory sampling ratio and a lower edit distance threshold can effectively improve the quality of the constructed set of model trajectories and ultimately improve the approximate alignment result. However, a higher trajectory sampling ratio and a lower edit distance threshold often require a greater computational effort, so it is very necessary to select a compromise parameter.

[0082] It can also be seen from Table 3 that on datasets RP, DD, and BPIC2012, when the sampling ratio sr = 0.05 and the edit distance threshold he = 0.5, the error is below 0.03. This indicates that the approximate fitness at this time can accurately measure the consistency between the event log and the process model. Selecting this combination of the trace sampling ratio and the edit distance threshold can ensure that the sampled event log volume is sufficiently representative while reducing the computational effort when processing the sampled event log, thus well balancing the accuracy of the fitness and the computational effort. The reason is that setting the edit distance threshold he to 0.5 can tolerate small differences between traces and avoid too strict screening criteria resulting in too many sampled traces. Especially for complex processes (such as datasets PL and BPIC2012), when the edit distance threshold he is 0.5, it can capture most of the individual differences in the business process traces. At the same time, setting the trace sampling ratio sr to 0.05 can effectively capture the important traces with higher frequencies in the event log. Therefore, in the subsequent online consistency detection experiment, the present invention sets the trace sampling ratio sr = 0.05 and the edit distance threshold he = 0.5.

[0083] Approximate Online Fitness Evaluation

[0084] In this study, the accuracy of the fitness of each method is evaluated by comparing the Spearman correlation coefficients of the online fitness obtained by the CCAES method, the HMM method, and the OCC method with a sliding window of 1 (OCC1) of the present invention. Since the IASR method can provide an accurate online consistency detection fitness, in the experiment, the present invention only performs a Spearman correlation coefficient analysis on the online fitness of the methods CCAES, OCC1, and HMM. The higher the Spearman correlation coefficient, the closer the detection result is to the result of the IASR method, that is, the more accurate the result of the online consistency detection.

[0085] In the experiment, first, each event log is simulated as an unbounded event stream data corresponding to the business process, and then the online fitness obtained by performing online consistency detection on the event stream from the start to the end with a preset business process model is recorded. In the CCAES method of the present invention, the sampling edit distance threshold he = 0.5 and the trace sampling ratio sr = 0.05 are set. The results are as Figure 4 shown. The abscissa in the figure is 5 event logs, and the ordinate represents the Spearman correlation coefficient between the online fitness of the three methods and the IASR online fitness.

[0086] As Figure 4As shown, the Spearman correlation coefficient corresponding to the CCAES method is better than that of the OCC1 method and the HMM method in the PL, DD, CCC19, and BPCI2012 datasets. The goodness of fit obtained by the CCAES method is closest to the accurate online alignment goodness of fit. Therefore, it can be concluded that the goodness of fit obtained by using the CCAES method in these four event logs is more accurate than that of the OCC1 method and the HMM method. In the RP event log, since the trace composition of this event log is simple and the trace length is small, only the OCC method and the HMM method with 1 window can obtain accurate goodness of fit.

[0087] In all the event log datasets of the experiments, under the condition that the CCAES method only sets the edit distance threshold he = 0.5 and the trace sampling ratio sr = 0.05, its corresponding Spearman correlation coefficient exceeds 0.85. This result shows that the CCAES method has a high correlation with the accurate online detection results. Based on this, it can be inferred that the CCAES method can provide relatively accurate online approximate consistency detection values. Although there will be certain errors when using the CCAES method for consistency detection, the ultimate goal of online consistency detection is to detect inconsistent execution traces. Therefore, some of the errors that occur in the CCAES method of the present invention are acceptable in most application scenarios of online consistency detection.

[0088] Evaluation of the completeness of trace prefixes

[0089] Since the accurate OCC and IASR methods currently cannot obtain the detection results of the trace completeness in the event stream, in the completeness analysis, the present invention only compares the CCAES method, the HMM method with the actual completeness of the trace prefixes in the event stream. Given that there are few traces in CCC19, only the event streams simulated by the four datasets RP, PL, DD, and BPCI2012 are selected for experiments in the integrity analysis experiment. Record the completeness of the trace prefixes in the event stream calculated by using the CCAES method and the HMM method. Figure 5 The change graph of the Spearman correlation coefficient of the completeness of the method of the present invention and the HMM method from the beginning to the end for simulating the event stream using the event log.

[0090] From Figure 5 (a), it can be seen that the Spearman coefficient of the HMM method is relatively high at the beginning stage of the event stream. However, as the event stream progresses, the Spearman coefficient calculated by the HMM method gradually changes and tends to be stable. The HMM method can obtain more accurate completeness at the beginning of the event stream, but as the event stream progresses, its accuracy gradually decreases. From Figure 5(b) It can be seen that the Spearman coefficients calculated by the method CCAES on the four datasets range from 0.965 to 1, indicating a high correlation between the completeness calculated by the method CCAES and the actual completeness. Among them, the actual completeness of the trajectories in the event stream is the ratio of the length of the prefix of the trajectory being executed to the length of the complete trajectory. As the event stream progresses, the Spearman coefficients of each event log tend to stabilize and eventually remain above 0.98. Compared with the method HMM, the completeness calculated by the method CCAES is more accurate for most of the time in the event stream. The completeness data calculated by the method CCAES has a high positive correlation with the actual completeness, and the Spearman coefficient is stable above 0.98, indicating that the fitness obtained by the method CCAES can effectively judge the progress of the event stream trajectory, and the method CCAES can be better applied to the online consistency detection of business process management.

[0091] Computational Efficiency Evaluation of Online Approximate Consistency Detection

[0092] To evaluate the performance of the method of the present invention, the present invention compares the average time for detecting events of the method CCAES, the method OCC1 and the method IASR. Since the fitness of the online approximate consistency detection of the method HMM is not accurate, the present invention does not compare with the HMM method in the performance experiment, but compares with the fastest method OCC (OCC1) and the method IASR to evaluate the performance of the online consistency detection method. In the experiment, the present invention records the time taken for each online consistency detection algorithm to perform online consistency detection with a preset business process model from start to end, and calculates the average time for detecting one event as the experimental result. The method of the present invention is to set the edit distance threshold he = 0.5 and the trajectory sampling ratio sr = 0.05 in the experiment. The results are as Figure 6 shown. The abscissa in the figure is each event log, and the ordinate is the average time for detecting the execution of an event.

[0093] As Figure 6 shown, compared with the consistency detection method based on prefix alignment, the CCAES method proposed by the present invention has achieved significant improvement in time performance. Compared with the IASR method, the approximate online detection method CCAES method and the OCC1 method in the OCC framework can obtain online detection results faster. Compared with the OCC1 method, the detection time of the method CCAES for processing a single event is further reduced. In the experiments on the three event log datasets of RFP, PL, and DD, the detection time of the method CCAES is significantly lower than that of the OCC1 method. This is because the method CCAES does not need to explore the space of the business process model in actual detection, thus saving more time and memory for detection. Therefore, the method CCAES is superior to the existing more accurate online consistency detection methods in terms of the efficiency of online approximate consistency detection.

[0094] To evaluate the memory consumption of the CCAES method of the present invention, the present invention compared the memory consumption of the CCAES method with that of OCC, OCC1, and IASR when performing online consistency detection in the event stream. This experiment used the tracemalloc tool of Python 3.7 to record the memory consumption of the four methods when performing online consistency detection, and used the recorded memory consumption as the experimental result. The comparison results are shown in Table 4.

[0095] Table 4 Memory Consumption (MB) of Four Online Consistency Detection Methods

[0096] As can be seen from Table 5, for the four online consistency detection methods, as the scale of the event stream simulated by the event stream simulator increases sequentially from top to bottom, the memory consumption of the four methods also increases accordingly. Compared with the existing accurate prefix alignment methods OCC, OCC1, and IASR methods, the CCASE method of the present invention significantly reduces the memory consumption. The reason is that the CCASE method uses a set of model traces instead of a complex process model in consistency detection, converting the original complex calculation of aligning trace prefixes with the process model into the calculation of edit distance and the length of the longest common sub-trace with less computational complexity, thus effectively reducing the memory consumption.

[0097] To further evaluate the memory consumption of the CCASE method proposed by the present invention in a larger-scale event stream, the present invention first used the designed event stream simulator to generate an event stream with a scale of 1 million events using the datasets DD and BPIC2012. Then, the CCASE method was used for online consistency detection, and the memory consumption of the CCASE method was recorded. Since the business process of the dataset DD is relatively simple, while the business process of the dataset BPIC2012 is relatively complex, the present invention selected them as two typical representatives of large-scale event streams to carry out the memory consumption evaluation of online consistency detection.

[0098] In the online consistency detection carried out, the memory consumption of the CCASE method was recorded every 1000 events detected until all events in the event stream were detected. The recorded memory consumption was used as the experimental result. The experimental results are as Figure 7 shown, where the abscissa is the number of events detected, and the ordinate is the memory space consumed by the detected events.

[0099] From Figure 7It can be seen that as the number of detection events increases, for the datasets DD and BPIC2012, the memory consumption of the method CCASE proposed by the present invention during online consistency detection generally increases linearly. And for the same number of detection events, on the datasets DD and BPIC2012, the memory consumption of the method CCASE is not much different. The main reason is that when calculating the approximate online fitness, the method CCAES uses the edit distance and the length of the longest common sub-trajectory, and its space complexity is O(n), where n is the number of online consistency detection events. Therefore, the memory consumption for each detected event is roughly the same, and the overall memory consumption increases with the increase in the number of events, making its memory usage approximately linearly related to the scale of the event stream. From Figure 7 It can also be seen that when performing online consistency detection in an event stream with 1 million events, the memory consumption of the method CCAES proposed by the present invention is about 120MB. Therefore, it can be concluded that for the online consistency detection of large-scale event streams, the memory consumption of the method CCAES proposed by the present invention is within an acceptable range.

[0100] In summary, compared with the existing approximate consistency detection methods, the method CCAES proposed by the present invention has significantly improved the accuracy of the approximate fitness and can achieve online consistency detection for event streams. At the same time, compared with the mainstream prefix alignment-based consistency detection methods, the method CCAES has significantly improved the detection efficiency and performance.

[0101] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, any modification, equivalent replacement, and improvement made within the spirit and principle of the present invention shall be covered by the protection scope of the present invention.

Claims

1. The present invention provides an approximate alignment consistency detection method for event streams, characterized in that, Including: Step 1: Sample the event log L by calculating the normalized edit distance of the traces; Step 2: Based on the obtained sampled event log L s , use the optimal alignment algorithm [12] to obtain the model trace set M of L s and the process model N T; Step 3: Obtain the model trace set M T After that, by using the trace edit distance cost function d(σ1,σ2) to replace the standard cost function c(b), the calculation overhead of the fitness between the event log and the process model is reduced; Step 4: Calculate the approximate online fitness of the trace prefixes in the event stream and obtain an approximate online fitness through weighted summation.

2. The approximate alignment consistency detection method for event streams according to claim 1, characterized in that, Define the normalized edit distance between σ1 and σ2 in the event log trace For the above-mentioned Step 1, in order to sample the event log, Algorithm 1 is designed, and the pseudo-code is as follows: With the help of Algorithm 1, the sampled event log can be obtained from the event log.

3. The approximate alignment consistency detection method for event streams according to claim 1, characterized in that: In the second step described above, based on the obtained sampled event log L s , the optimal alignment algorithm is used [12] to obtain L s and the model trace set M of the process model N T; That is, calculate the best alignment of the traces of the business process model and the sampled event log, and construct the corresponding model trace set from the process model traces mapped by the alignment result. Among them, the alignment method is as follows: Alignment of event logs and process models. Let A be a set of activities, and σ ∈ A * be a trace over A, and M = (P, T, F, α, m i , m f ) be a Petri net over A. An alignment γ ∈ (A 》 × T 》 ) between the trace σ and the net N is a sequence of moves that satisfies the following conditions; (1) The alignment between the trace in the event log and the process model is a movement sequence γ. For any movement (a, t) ∈ γ in the alignment, the present invention defines (a, t) as follows: 1) If a ∈ A and t =》, then (a, t) is called a log movement; 2) If a =》 and t ∈ T, then (a, t) is called a model movement; 3) If a ∈ A, t ∈ T and α(t) = a, or a =》 and t = τ, then (a, t) is called a synchronous movement; 4) Otherwise, it is called an illegal movement; Denote Γ σ,N represents the set of all alignments between the trace σ and the Petri net model N; (2)π1(γ) ↓A = σ, the projection of the sequence consisting of the first element of each move in the move sequence γ on A (ignoring ») produces a trajectory; (3) The projection (ignoring >>) on model T of the sequence formed by the second elements of each move in the move sequence γ produces a firing sequence that causes the marking of model T to change from m i to m f ; To obtain the model trace set, Algorithm 2 is designed, and the pseudo-code is as follows: With the help of Algorithm 2, the model behavior set can be obtained from the sampled event log.

4. An approximate alignment consistency detection method for event streams according to claim 1, characterized in that: For the above-mentioned Step 3, use the trace edit distance cost function d(σ1, σ2) to replace the standard cost function c(b). The standard cost function c(b) and the trace edit distance cost function d(σ1, σ2) are introduced as follows: Standard cost function c(b): Let A be a set of activities, A* denote all possible sequences of activities, σ ∈ A* be a trace over A, and N = (P, T, F, α, m i , m f ) be a Petri net over A. The standard cost function c(b) can be expressed as the cost value of a move assignment, where b = (a, t) is an arbitrary move in the alignment of the trace with the process model: (1) c(b) = 1. If a ∈ A and t =》, then (a, t) is called a log movement; (2) c(b) = 1. If a =》, t ∈ T and α(t) ≠ τ, then (a, t) is called a model movement; (3) c(b) = 0. If a ∈ A, t ∈ T and α(t) = a, or a =》 and α(t) = τ, or a = τ and t =》, then (a, t) is called a synchronous movement; (4)c(b) = +∞, if a => and t => or a ∈ A τ , t ∈ T and α(t) ≠ a, then (a, t) is called an illegal move; Trajectory Edit Distance Cost Function d(σ1, σ2): Suppose σ1 ∈ A* is a trajectory, and the activity sequence <t1, t2,..., t n (n ∈ N * ) corresponding to <α(t1), α(t2),..., α(t n )> is denoted as trajectory σ2. The trajectory edit distance cost function td(σ1, σ2) → N represents the minimum number of edit operations required to convert σ1 into σ2, where N is the set of natural numbers. Here, the edit operations only include insertion and deletion operations and do not include replacement operations. In the alignment of trajectories and process models, each asynchronous move is equivalent to one edit operation, where the insertion operation corresponds to model movement and the deletion operation corresponds to log movement; Therefore, d() returns the same cost value as the control flow alignment. For example, d(<a, b>, <a, d>) = 2, and its edit operation is one insertion operation (b,》) and one deletion operation (》, d).

5. The approximate alignment consistency detection method for event streams according to claim 1, wherein: The above-mentioned Step 4: Calculate the approximate online fitness of the trace prefixes in the event stream and obtain an approximate online fitness through weighted summation; As the business process instance is executed, the evolution of the trace prefix tends to be stable, so the number of model traces participating in the online consistency detection also gradually decreases; When the execution of the business process instance is approaching the end, the event sequence in the trace prefix has been basically determined. At this time, only 1 model trace is required to meet the calculation needs of the approximate online consistency detection. As the business process is executed, the model traces used to calculate the online fitness in the consistency detection will gradually decrease from multiple at the beginning to 1. The calculation of the online fitness is shown in formula (2): where n is the ratio of the length of σ to the length of the sampled trace with the minimum edit distance corresponding to it and coming from the sampled event log, and minEdDis(σ, M T , n) represents the n model traces with the minimum edit distance between M T and σ, and len(σ) represents the length of the trace σ In the scenario of a dynamic unbounded event stream, the consistency detection of the business process not only needs to detect the fitness of the current event trace in real time, but also needs to judge the completeness of the current trace; Using only the fitness to evaluate the online consistency detection result is one-sided and may lead to misleading consequences. When conducting the consistency detection of the event flow trajectory, in order to compare the event flow trajectories more accurately, it is also necessary to judge the progress of the trajectory and predict the events that may be missing from the current event flow trajectory; Based on obtaining the fitness between the event flow and the process model, the present invention adds an evaluation of the integrity of the event flow trajectory. The trajectory integrity can predict the number of events that are currently unobserved in the event flow trajectory, and at the same time judge the progress of the executing business process instance represented by this trajectory; On the contrary, the business process shown in case3 has not ended at the current time point, which means that there may still be events that have not occurred in the instance trace of this business process. These events need to be inferred during the detection process. The integrity of case3 is less than 1. The following gives the relevant definition of trace integrity. Given the trace σ of the event stream log and the model trace set M T , the integer n ∈ N + , then the integrity of σ is calculated as follows: In order to implement step four of the present invention's example, an online approximate consistency detection of the event flow and the business process model is realized using the model trajectory set. The present invention designs algorithm 3 for online consistency detection, and the pseudocode is as follows: Finally, the consistency detection results fitbness and completeness are obtained through algorithm 3.

6. An approximate consensus detection system implementing the trace frequency cost matrix according to any one of claims 1-5, characterized in that, The approximate consistency detection system based on the trace frequency cost matrix includes: An event flow log sampling module to select a sampled event log in the event log that can represent the recorded event log; An approximate consistency detection module for the prefixed-aligned event log and the process model, which uses the model trajectory set to replace the business process model to reduce the computational complexity of the consistency detection and reduce the fitness error; A calculation module for the approximate fitness and integrity of the trajectories in the event flow to more accurately evaluate the consistency between the event flow and the process model.