Process mining method, system and equipment and computer storage medium
By obtaining event logs in process mining, reducing self-cycle processes, filtering noise events and computing dependencies, the problems of noise interference and complex process processing in the prior art are solved, and the accuracy of the process model is improved.
Patent Information
- Application Number
- CN202510020894.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art is difficult to resist noise interference and handle complex loop and nested process structures in process mining, resulting in the impact of the accuracy of the process model.
By obtaining event logs, determining event order and frequency, reducing self-loop flow, filtering noise events and noise loops, calculating event dependence and deleting low-dependence downstream events, and generating process models.
Effectively filter noise events and noise loops to improve the accuracy of process models, especially when dealing with complex loops and nested structures.
Smart Images

Figure CN119940895A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of process mining, and in particular relates to a process mining method, system, device and computer storage medium. Background Art
[0002] With the development of enterprises, the issue of process consistency has received more and more attention. Process consistency, in layman's terms, refers to whether the actual workflow is consistent with the process model. If it is inconsistent, it means that there is a problem in the work process. For example, in the material supply management link of power companies, the integrated management of the supply chain is achieved by integrating logistics, information flow, capital flow and other elements. However, in terms of process consistency analysis, it often relies on manual audits and experience judgments, lacks overall layout and global judgment, and leads to various problems in the process of purchasing and applying materials, affecting the stability of material supply. Therefore, an accurate process model is crucial, but the actual workflow and the workflow perceived by the management are prone to differences. Therefore, the existing technology uses process mining technology to determine the process model.
[0003] Process mining is to obtain the real process of the entire business execution by analyzing event logs. However, the execution process is not equivalent to the process model, because errors or abnormal events are prone to occur during the execution process. Therefore, the current existing technology still has certain shortcomings. First, it cannot resist noise interference; second, it is difficult to handle complex process structures such as loops and nesting. Both of these problems are likely to cause model deviations, which ultimately affect the accuracy of the process model.
[0004] The information disclosed in this background technology section is only intended to enhance the understanding of the overall background of the invention and should not be regarded as an acknowledgment or any form of suggestion that the information constitutes the prior art already known to a person skilled in the art. Summary of the invention
[0005] The purpose of the present invention is to solve the problem of inaccurate acquisition of process models and to provide a process mining method, system, device and computer storage medium.
[0006] A first aspect of the present invention provides a process mining method, comprising: acquiring data, and determining the order and frequency of each event based on an event log in the data;
[0007] Reduce the self-looping process of events that are repeated multiple times to a single self-looping process, but keep its number of loops;
[0008] Filter noise events and noise loops;
[0009] Calculate the dependency of adjacent upstream and downstream events. If the dependency value is less than the preset dependency threshold, delete the downstream event and assign the frequency of the deleted event to the associated parallel event.
[0010] Generate a process model based on the remaining events and the logical relationships between them.
[0011] In one embodiment of the present invention, the process of filtering noise events is: deleting events whose frequency is lower than a preset noise threshold, and assigning the frequency quantity of the deleted events to the associated parallel events.
[0012] In one embodiment of the present invention, the process of filtering noise events is: identifying abnormally frequent and / or abnormally sparse events based on statistical methods and deleting them, and assigning the frequency quantity of the deleted events to the associated parallel events.
[0013] In one embodiment of the present invention, the first cycle factor is defined as
[0014]
[0015] Where N(XX) represents the number of times event X repeats itself; N(X) represents the total number of times event X occurs;
[0016] If the value of the first cycle factor is less than the preset cycle threshold, the self-cycle of event X is considered to be a noise cycle;
[0017] Delete the downstream events in the self-loop, and assign the frequency quantity of the deleted events to the associated parallel events.
[0018] In one embodiment of the present invention, the second cycle factor is defined as
[0019]
[0020] Among them, N(XYX) represents the number of times event X goes through event Y and then executes event X; N(XY) represents the number of times event X goes to event Y;
[0021] If the value of the second cycle factor is less than the preset cycle threshold, the two-step cycle of event X is considered to be a noise cycle.
[0022] Delete the downstream event in the two-step loop, and assign the frequency quantity of the deleted event to the associated parallel event.
[0023] In one embodiment of the present invention, the calculation formula of the dependency is:
[0024] W(X,Y)=α·F(X,Y)+β·CF1(X)+γ·CF2(X,Y);
[0025] Among them, α, β and γ are weight distribution parameters; F(X,Y) is the dependency calculation formula of the Heuristic Miner algorithm.
[0026] In one embodiment of the present invention, the source of the data is an enterprise management system.
[0027] The second aspect of the present invention provides a process mining system, including: an input module for acquiring data; a processing module for determining the order and frequency of each event based on the event log in the data; reducing a self-looping process in which events are repeated multiple times in succession to a single self-looping process, but retaining its number of cycles; filtering noise events and noise cycles; calculating the dependency of adjacent upstream and downstream events, and if the value of the dependency is less than a preset dependency threshold, deleting the downstream event, and assigning the frequency quantity of the deleted event to the associated parallel event; generating a process model based on the remaining events and the logical relationships between the events; and a display module for visually expressing the process model.
[0028] A third aspect of the present invention provides a device, including a processor and a memory communicatively connected to the processor, wherein the memory stores computer instructions executable by the processor, and the computer instructions implement the above-mentioned process mining method when executed.
[0029] A fourth aspect of the present invention provides a computer storage medium storing computer instructions, which implement the above-mentioned process mining method when executed.
[0030] Compared with the prior art, the technical effects achieved by the present invention are as follows:
[0031] 1. It can filter noise events and noise cycles, thereby reducing the impact of random events or abnormal behaviors on the process model and improving the accuracy of the process model;
[0032] 2. In the process of calculating the dependency, the influence of the self-loop and the two-step loop is fully considered. That is, when there are complex loops in the process, the dependency calculation in the present invention can better reflect the actual dependency situation, thereby improving the accuracy of the process model. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 is a flow chart of a process mining method according to an embodiment of the present invention;
[0034] Figure 2 is a Petri diagram after reducing the self-loop process according to the process mining method of one embodiment of the present invention;
[0035] Figure 3 is a Petri diagram of a process mining method after filtering noise events according to an embodiment of the present invention;
[0036] Figure 4 is a Petri diagram after filtering out noise self-loops according to a process mining method of an embodiment of the present invention;
[0037] Figure 5 It is a Petri diagram of the process mining method according to an embodiment of the present invention after filtering the two-step cycle of noise. DETAILED DESCRIPTION
[0038] Unless explicitly stated otherwise, throughout the specification and claims, the term “comprise” or variations such as “include” or “comprising”, etc., will be understood to include the stated elements or components but not to exclude other elements or components.
[0039] The technical solution of the present invention is described below by specific embodiments. It should be understood that one or more steps mentioned in the present invention do not exclude the existence of other methods and steps before and after the combination step, or other methods and steps can be inserted between these explicitly mentioned steps. It should also be understood that these examples are only used to illustrate the present invention and are not used to limit the scope of the present invention. Unless otherwise specified, the numbering of each method step is only for the purpose of identifying each method step, and does not limit the order of arrangement of each method or limit the scope of implementation of the present invention. The change or adjustment of the relative relationship thereof can also be regarded as the scope of implementation of the present invention without substantial changes in the technical content.
[0040] The sources of the raw materials and instruments used in the examples are not particularly limited and can be purchased from the market or prepared according to conventional methods known to those skilled in the art.
[0041] like Figure 1 As shown, the process mining method according to the preferred embodiment of the present invention includes acquiring data and determining the order and frequency of each event based on the event log in the data.
[0042] The source of data is usually the EPR system, SCM system or other enterprise management system, but data can also be generated by manual entry. The data includes the most important event log, which records the event number, execution content, execution time, upstream and downstream events of an event. By analyzing the event log, the logical relationship between events and the frequency of events can be obtained.
[0043] Since there are many events in the whole process that will be self-looping, that is, event X will be executed after event X is executed, and the number of self-loops is unlimited, it is necessary to reduce the self-looping process of events that are repeated many times to a single self-looping process, but keep its number of loops. For example Figure 2 From event G to event G, there is a self-loop, but the frequency is written 5 times.
[0044] Noise events and noise cycles then need to be filtered out.
[0045] There are two ways to filter noise events. One is to delete events whose frequency is lower than the preset noise threshold. The noise threshold is an empirical parameter that needs to be determined by considering the overall number of events in the data. The second is to identify and delete abnormally frequent and abnormally sparse events based on statistical methods. The statistical method uses quantile statistics, which regards events below the 10th percentile as abnormally sparse events and events above the 90th percentile as abnormally frequent events. Abnormally frequent and abnormally sparse events are both abnormal events that deviate from the normal range. The purpose of filtering noise events is to reduce the impact of random events or abnormal behaviors on the process model.
[0046] There are two processes of filtering cycle, one is filtering of self-loop, the other is filtering of two-step loop. Regarding self-loop filtering, it is necessary to define the first loop factor, and its formula is
[0047]
[0048] Where N(XX) represents the number of times event X repeats itself; N(X) represents the total number of times event X occurs.
[0049] If the value of the first cycle factor is less than the preset cycle threshold, the self-loop of event X is considered to be a noise loop. Then, the downstream events in the noise loop need to be deleted, and the frequency of the deleted events is assigned to the associated parallel events.
[0050] For two-step cyclic filtration, it is necessary to define the second cyclic factor, whose formula is:
[0051]
[0052] Among them, N(XYX) represents the number of times event X goes through event Y and then executes event X; N(XY) represents the number of times event X goes to event Y;
[0053] If the value of the second cycle factor is less than the preset cycle threshold, the two-step cycle of event X is considered to be a noise cycle. Event Y in the noise cycle needs to be deleted, and the frequency of the deleted event is assigned to the associated parallel event.
[0054] Noise cycles are different from noise events. Whether an event is a noise event depends mainly on the frequency of the event itself, while noise cycles mainly look at the ratio of downstream events to upstream events. If you only look at the frequency of the event itself, you will easily ignore the noise cycle, which will affect the accuracy of the process model.
[0055] The X and Y in this article are just parameterized representations and do not refer to any specific event.
[0056] The so-called associated parallel events refer to multiple events with the same upstream event and the same location level. When an event is deleted, its frequency quantity needs to be assigned to the associated parallel events. If there are multiple parallel events, they are assigned according to the frequency weights of the parallel events.
[0057] refer to Figure 2 and Figure 3 ,After filtering the noise events, event F is deleted, and the frequency quantity of event F is assigned to event E.
[0058] refer to Figure 2 and Figure 4 After the noise cycle of the first cycle factor is filtered, the downstream event G is deleted. Although event G and event J are parallel events, since the downstream event of event G is also event J, the frequency quantity of event J does not change after the frequency quantity of event G is assigned to event J.
[0059] refer to Figure 2 and Figure 5 After the noise cycle about the second cycle factor is filtered, the downstream event I in the two-step cycle of event I through event K and then to event I is deleted. The frequency quantity of event I is assigned to event K.
[0060] Then the dependency of adjacent upstream and downstream events is calculated, such as the dependency of event A and event B, the dependency of event A and event C, and the dependency of event I and event K. The dependency of all adjacent upstream and downstream events is calculated. If the value of the dependency is less than the preset dependency threshold, the downstream event is deleted, and the frequency of the deleted event is assigned to the associated parallel event.
[0061] The calculation formula for dependency is:
[0062] W(X,Y)=α·F(X,Y)+β·CF1(X)+γ·CF2(X,Y);
[0063] Among them, α, β and γ are weight distribution parameters; F(X,Y) is the dependency calculation formula of the Heuristic Miner algorithm. The Heuristic Miner algorithm is an existing technology, and its dependency calculation does not take into account the impact of the self-loop and two-step loop of the event. That is, when there are complex loops in the process, the dependency calculation of the existing technology cannot reflect the actual dependency situation, which ultimately affects the accuracy of the process model.
[0064] If the calculated dependency is lower than the dependency threshold, it is considered that the downstream event has a low degree of relevance to the upstream event. The lower the degree of relevance, the worse the logic of sequential execution, and the downstream event should be deleted.
[0065] Then, a process model is generated based on the remaining events and the logical relationships between them. Figure 5 Content shown.
[0066] As an implementation of the above method, the present invention discloses an embodiment of a process mining system, which corresponds to the above method embodiment.
[0067] The process mining system includes an input module, a processing module and a display module.
[0068] Input module, which is used to obtain data. Processing module, which determines the order and frequency of each event based on the event log in the data; reduces the self-loop process of events that are repeated multiple times to a single self-loop process, but retains its number of loops; filters noise events and noise loops; calculates the dependency of adjacent upstream and downstream events. If the value of the dependency is less than the preset dependency threshold, the downstream event is deleted, and the frequency of the deleted event is assigned to the associated parallel event; generates a process model based on the remaining events and the logical relationship between events. Display module, used to visualize the process model.
[0069] As an implementation of the above method, the present invention discloses an embodiment of a device, which corresponds to the above method embodiment. The device includes a processor and a memory connected to the processor in communication, the memory stores computer instructions that can be executed by the processor, and the above process mining method is implemented when the computer instructions are executed.
[0070] As an implementation of the above method, the present invention discloses an embodiment of a computer storage medium, which corresponds to the above method embodiment. The computer storage medium stores computer instructions, which implement the above process mining method when executed.
[0071] The foregoing description of specific exemplary embodiments of the present invention is for the purpose of illustration and demonstration. These descriptions are not intended to limit the present invention to the precise form disclosed, and it is clear that many changes and variations can be made based on the above teachings. The purpose of selecting and describing the exemplary embodiments is to explain the specific principles of the present invention and its practical application, so that those skilled in the art can realize and utilize various different exemplary embodiments of the present invention and various different selections and changes. The scope of the present invention is intended to be limited by the claims and their equivalents.
Claims
1. A process mining method, characterized in that: include: Obtain data and determine the order and frequency of events based on the event logs in the data; Reduce the self-looping process of events that are repeated multiple times to a single self-looping process, but keep its number of loops; Filter noise events and noise loops; Calculate the dependency of adjacent upstream and downstream events. If the dependency value is less than the preset dependency threshold, delete the downstream event and assign the frequency of the deleted event to the associated parallel event. Generate a process model based on the remaining events and the logical relationships between them.
2. The process mining method according to claim 1, characterized in that: The process of filtering noise events is to delete events whose frequency is lower than the preset noise threshold, and assign the frequency of the deleted events to the associated parallel events.
3. The process mining method according to claim 1, characterized in that: The process of filtering noise events is: identifying abnormally frequent and / or abnormally sparse events based on statistical methods and deleting them, and assigning the frequency quantity of the deleted events to the associated parallel events.
4. The process mining method according to claim 1, characterized in that: Define the first cycle factor as Where N(XX) represents the number of times event X repeats itself; N(X) represents the total number of times event X occurs; If the value of the first cycle factor is less than the preset cycle threshold, the self-cycle of event X is considered to be a noise cycle; Delete the downstream events in the self-loop, and assign the frequency quantity of the deleted events to the associated parallel events.
5. The process mining method according to claim 4, characterized in that: Define the second cycle factor as Among them, N(XYX) represents the number of times event X goes through event Y and then executes event X; N(XY) represents the number of times event X goes to event Y; If the value of the second cycle factor is less than the preset cycle threshold, the two-step cycle of event X is considered to be a noise cycle; Delete the downstream event in the two-step loop, and assign the frequency quantity of the deleted event to the associated parallel event.
6. The process mining method according to claim 5, characterized in that: The calculation formula for dependency is: W(X,Y)=αF(X,Y)+βCF1(X)+γCF2(X,Y); Among them, α, β and γ are weight distribution parameters; F(X,Y) is the dependency calculation formula of the Heuristic Miner algorithm.
7. The process mining method according to claim 1, characterized in that: The source of the data is the enterprise management system.
8. A process mining system, characterized in that: include: Input module, used to obtain data; A processing module determines the order and frequency of events based on the event logs in the data; Reduce the self-looping process of events that are repeated multiple times to a single self-looping process, but keep its number of loops; Filter noise events and noise cycles; calculate the dependency of adjacent upstream and downstream events. If the dependency value is less than the preset dependency threshold, delete the downstream event and assign the frequency of the deleted event to the associated parallel event; Generate a process model based on the remaining events and the logical relationships between them; The display module is used to visually express the process model.
9. A device, characterized in that: It comprises a processor and a memory in communication with the processor, wherein the memory stores computer instructions executable by the processor, and when the computer instructions are executed, the process mining method according to any one of claims 1 to 6 is implemented.
10. A computer storage medium storing computer instructions, characterized in that: When the computer instructions are executed, the process mining method as described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Double-granularity noise log filtering method based on incidence relation
CN110032494A
Process mining system based on causal concurrent network
CN113947374A
Incremental event log-oriented process model mining method and system
CN115525693A
Process model discovery method based on decomposition loop structure
CN115617877A