One-time plot rule mining method and device for process event logs

By introducing time interval constraints and one-time conditions in episode mining, generating a set of frequent one-time episodes, iteratively executing the episode connection strategy, and screening the support and confidence thresholds, the problem of meaningless episode mining caused by not considering time intervals in existing technologies is solved, and efficient and accurate rule mining is achieved.

CN120724401APending Publication Date: 2025-09-30HEBEI UNIV OF TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510831897.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

Existing episode mining methods do not consider time interval constraints, which leads to overestimation of episode frequency and mining of a large number of meaningless episodes.

Method used

By adopting time interval constraints and one-time conditions, a set of frequent one-time episodes is generated, the episode connection strategy is iteratively executed, and the support and confidence thresholds are screened to generate efficient and accurate one-time episode rules.

Benefits of technology

It achieves efficient and accurate mining of strong one-time plot rules with time interval constraints, avoids the mining of meaningless plots, and improves the effectiveness of plot mining.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120724401A_ABST
    Figure CN120724401A_ABST
Patent Text Reader

Abstract

The invention discloses a one-time plot rule mining method and device for a process event log, and belongs to the field of plot mining. The distance between two events is limited by adopting a time interval constraint, and repeated use of the events is avoided by adopting a one-time condition so as to avoid mining excessive meaningless plots. Moreover, when a strong one-time plot rule (i.e., a one-time plot rule with confidence greater than or equal to a minimum confidence threshold) is mined, a triple mechanism of candidate plot generation, support degree calculation and rule generation is provided. In a candidate plot generation stage, eliminating redundant candidate plot extension by adopting a plot connection strategy; in the support degree calculation stage, the support degree is calculated based on the position index; in a rule generation stage, firstly, all frequent one-time plots are mined, and then strong one-time plot rules are screened according to a minimum support degree threshold value and a minimum confidence coefficient threshold value. The technical effect of efficiently and accurately mining the strong one-time plot rule is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of scenario mining, and in particular to a one-time scenario rule mining method and device for process event logs. Background Art

[0002] The development of the Internet of Things has enhanced the ability to generate and store data from various sources. A large number of connected devices are embedded in the surrounding environment, completely changing the way people interact with things. The increase in connected devices means that a large amount of valuable process event log data is now stored. However, due to the explosive growth in the amount of data, obtaining valuable information remains challenging. Data mining is a useful method to extract meaningful information from large amounts of data. Sequential pattern mining (SPM) is a research branch of data mining that aims to find patterns of interest to users from sequence databases. Various SPM methods have been developed to solve different problems, including spatial co-location SPM, three-partition SPM, contrast SPM, high-utility SPM, negative SPM, and interval-constrained SPM.

[0003] However, SPM only considers the order of characters and not timestamps. A process event log is a list of events with many attributes, of which the timestamp (indicating the time when the event occurred) is particularly important. Episode mining (EM) can process event sequences with continuous timestamps and has become a popular research area in data mining. EM techniques have been widely applied to real-world data across various industries, such as customer transactions, web navigation logs, alarm sequences in telecommunications networks, and timestamped automobile breakdown reports.

[0004] Discovering hidden episodic rules from frequent episodes is a fundamental problem in EM. Given two frequent episodes α and β, one can generate valid episodic rules of the form α→β, where β is a prefix of α. In many practical applications, discovered frequent episodes are used to form episodic rules for decision making and prediction.

[0005] Example 1: Figure 1 An example of plot rule mining is shown, where events are represented by capital letters and timestamps are represented by numbers. The plot (D, T, T) appears twice in the event sequence. If T is taken as the consequent, the plot rule (D, T) → (D, T, T) is generated. This rule states that after the plot (D, T) occurs, event T will occur. However, although the plot rule (D, T) → (D, T, T) states that T will occur, it does not specify when T will occur. Therefore, in practical applications, if there is no upper limit on the time interval between the occurrence of T, such a rule is useless. For example, suppose Figure 1The sequence of events in describes the production information of a product order, where each event corresponds to a processing machine. Then the above rule means that after the operation of machine D is completed, machine T will be reworked. However, without an upper limit on the time interval of T, it is difficult to determine whether rework occurs in T. If the upper limit of the time interval is set to 1, the second occurrence (the second rectangle) does not exist, and the support (number of occurrences) is 1. This can be regarded as a random phenomenon and no maintenance is required. If the upper limit of the time interval is set to 5, the support is 2. Assuming the support threshold is 2, the plot (D, T, T) is frequent, indicating that there is a rework problem on machine T and the production line needs to be stopped for inspection to eliminate potential risks.

[0006] Currently, no effective solution has been proposed to the technical problem that the plot mining methods in the above-mentioned prior art generally do not consider time interval constraints and overestimate the frequency of plots, resulting in the mining of a large number of meaningless plots. Summary of the Invention

[0007] The embodiments of the present disclosure provide a one-time scenario rule mining method and apparatus for process event logs, to at least address the technical problem that the scenario mining methods in the prior art generally fail to consider time interval constraints and overestimate the frequency of scenarios, resulting in the mining of a large number of meaningless scenarios.

[0008] According to one aspect of an embodiment of the present disclosure, a one-time plot rule mining method for process event logs is provided, comprising: parsing the process event logs, generating an event sequence s, s = s1s2 ... s n =(e1,t1)(e2,t2)…(e n ,t n ), where s i Indicates an event, e i Indicates event activity, t i Timestamps indicating when the event occurred, t1 and t n are the start time and end time of s respectively; mine the frequent one-time episodes with episode length m in s and obtain the frequent one-time episode set F m , and create the F based on the timestamps of all events contained in each frequent one-time episode m Position index; wherein, the frequent one-time plot is a plot with a support degree in the s greater than or equal to the minimum support threshold, requiring that all events involved in multiple occurrences have time interval constraints and no overlap in occurrence time; with the F m As the initial input frequent set, iteratively perform the following steps until the preset stopping condition is met: Based on the current input frequent set, generate a one-time candidate episode set C with an episode length of m+1 according to the preset episode connection strategy m+1; Wherein, the plot connection strategy is: if the prefix subplot of the frequent one-time plot α is the same as the suffix subplot of the frequent one-time plot β, then generate a candidate one-time plot r with a plot length of m+1, r=γ⊕α; based on the position index, filter the C m+1 The candidate one-time episodes with medium support ≥ the minimum support threshold are used to obtain the frequent one-time episode set F m+1 ; Screening the F m+1 A one-time plot rule with confidence ≥ minimum confidence threshold; wherein the implication of the one-time plot rule is β→α, α and β are two frequent one-time plots, and β is a prefix subplot of α, and the confidence of the one-time plot rule is the ratio of the support of α to the support of β; the F m+1 As the new current input frequent set; wherein, the stopping condition is the C m+1 Empty.

[0009] According to another aspect of an embodiment of the present disclosure, a storage medium is further provided, the storage medium including a stored program, wherein when the program is run, a processor executes any one of the above methods.

[0010] According to another aspect of the embodiment of the present disclosure, a one-time plot rule mining device for process event logs is provided, comprising: an event sequence generation module for parsing the process event logs and generating an event sequence s, s = s1s2 ... s n =(e1,t1)(e2,t2)…(e n ,t n ), where s i Indicates an event, e i Indicates event activity, t i Timestamps indicating when the event occurred, t1 and t n are the start time and end time of s respectively; the frequent one-time episode mining module is used to mine the frequent one-time episodes with an episode length of m in s, and obtain the frequent one-time episode set F m , and create the F based on the timestamps of all events contained in each frequent one-time episode m Position index of the frequent one-time plot; wherein, the support degree of the frequent one-time plot in the s is greater than or equal to the minimum support threshold, requiring that all events involved in multiple occurrences have time interval constraints and no overlap in the occurrence time; a strong one-time plot rule mining module is used to extract the frequent one-time plot rule based on the F m As the initial input frequent set, iteratively perform the following steps until the preset stopping condition is met: Based on the current input frequent set, generate a one-time candidate episode set C with an episode length of m+1 according to the preset episode connection strategy m+1Wherein, the plot connection strategy is: if the prefix subplot of the frequent one-time plot α is the same as the suffix subplot of the frequent one-time plot β, then generate a candidate one-time plot r with a plot length of m+1, Based on the position index, filter the C m+1 The candidate one-time episodes with medium support ≥ the minimum support threshold are used to obtain the frequent one-time episode set F m+1 ; Screening the F m+1 A one-time plot rule with confidence ≥ minimum confidence threshold; wherein the implication of the one-time plot rule is β→α, α and β are two frequent one-time plots, and β is a prefix subplot of α, and the confidence of the one-time plot rule is the ratio of the support of α to the support of β; the F m+1 As the new current input frequent set; wherein, the stopping condition is the C m+1 Empty.

[0011] According to another aspect of the embodiment of the present disclosure, there is also provided a one-time plot rule mining device for process event logs, comprising: a processor; and a memory connected to the processor, for providing the processor with instructions for processing the following processing steps: parsing the process event log to generate an event sequence s, s = s1s2 ... s n =(e1,t1)(e2,t2)…(e n ,t n ), where s i Indicates an event, e i Indicates event activity, t i Timestamps indicating when the event occurred, t1 and t n are the start time and end time of s respectively; mine the frequent one-time episodes with episode length m in s and obtain the frequent one-time episode set F m , and create the F based on the timestamps of all events contained in each frequent one-time episode m Position index; wherein, the frequent one-time plot is a plot with a support degree in the s greater than or equal to the minimum support threshold, requiring that all events involved in multiple occurrences have time interval constraints and no overlap in occurrence time; with the F m As the initial input frequent set, iteratively perform the following steps until the preset stopping condition is met: Based on the current input frequent set, generate a one-time candidate episode set C with an episode length of m+1 according to the preset episode connection strategy m+1 Wherein, the plot connection strategy is: if the prefix subplot of the frequent one-time plot α is the same as the suffix subplot of the frequent one-time plot β, then generate a candidate one-time plot r with a plot length of m+1, Based on the position index, filter the C m+1The candidate one-time episodes with medium support ≥ the minimum support threshold are used to obtain the frequent one-time episode set F m+1 ; Screening the F m+1 A one-time plot rule with confidence ≥ minimum confidence threshold; wherein the implication of the one-time plot rule is β→α, α and β are two frequent one-time plots, and β is a prefix subplot of α, and the confidence of the one-time plot rule is the ratio of the support of α to the support of β; the F m+1 As the new current input frequent set; wherein, the stopping condition is the C m+1 Empty.

[0012] This application proposes a one-time plot rule mining scheme with time interval constraints, which uses time interval constraints to limit the distance between two events and uses one-time conditions to avoid the reuse of events, so as to avoid mining too many meaningless plots. In addition, when mining strong one-time plot rules (i.e., one-time plot rules with confidence ≥ minimum confidence threshold), a triple mechanism of candidate plot generation, support calculation and rule generation is proposed. In the candidate plot generation stage, a plot connection strategy is adopted to eliminate redundant candidate plot extensions; in the support calculation stage, support is calculated based on position index; in the rule generation stage, all frequent one-time plots are first mined, and then strong one-time plot rules are screened according to the minimum support threshold and the minimum confidence threshold. The technical effect of efficiently and accurately mining strong one-time plot rules is achieved. This solves the technical problem that the plot mining methods in the prior art generally do not consider time interval constraints and overestimate the frequency of plots, resulting in the mining of a large number of meaningless plots. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The drawings described herein are used to provide a further understanding of the present disclosure and constitute a part of this application. The illustrative embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute an improper limitation of the present disclosure. In the drawings:

[0014] Figure 1 is a running example of a sequence of events;

[0015] Figure 2 is a hardware structure block diagram of a computing device for implementing the method according to embodiment 1 of the present disclosure;

[0016] Figure 3 This is a framework flow chart of the one-time plot rule mining method for process event logs according to Example 1 of the present application;

[0017] Figure 4 is a schematic diagram of a candidate scenario tree generated by the enumeration strategy according to Example 1 of the present application;

[0018] Figure 5 1 is a schematic diagram of the Nettree nodes of the scenario α on the sequence s and the matching process thereof according to Example 1 of the present application;

[0019] Figure 6 This is a comparison chart of operating times according to Example 1 of the present application;

[0020] Figure 7 This is a comparison chart of the number of candidate plots according to Example 1 of the present application;

[0021] Figure 8 This is a memory usage comparison chart according to Example 1 of the present application;

[0022] Figure 9 This is a comparison chart of the running times for different minsup values ​​according to Example 1 of the present application;

[0023] Figure 10 This is a comparison chart of the number of candidate plots with different minsup values ​​according to Example 1 of the present application;

[0024] Figure 11 This is a comparison chart of memory usage for different minsup values ​​according to Example 1 of the present application;

[0025] Figure 12 This is a comparison chart of the running time of different tgap values ​​according to Example 1 of the present application;

[0026] Figure 13 This is a comparison chart of the number of candidate plots with different tgap values ​​according to Example 1 of the present application;

[0027] Figure 14 This is a comparison chart of memory usage for different tgap values ​​according to Example 1 of the present application;

[0028] Figure 15 This is a comparison chart of the running time of different log sizes according to Example 1 of the present application;

[0029] Figure 16 This is a comparison chart of memory usage for different log sizes according to Example 1 of the present application;

[0030] Figure 17 This is a comparison chart of the number of rules of ELog1-ELog8 described in Example 1 of the present application;

[0031] Figure 18 This is a comparison chart of the confidence of the OER and strong OER of Elog8 described in Example 1 of the present application;

[0032] Figure 19This is a comparison chart of the number of strong OERs under different minconf values ​​described in Example 1 of the present application;

[0033] Figure 20 is a schematic diagram of the Petri net of the "Adjusting-nut" product described in Example 1 of the present application;

[0034] Figure 21 2 is a schematic diagram of a one-time plot rule mining device for process event logs according to Example 2 of the present application;

[0035] Figure 22 This is a schematic diagram of a one-time plot rule mining device for process event logs according to Example 3 of the present application. DETAILED DESCRIPTION

[0036] In order to enable those skilled in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present disclosure.

[0037] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0038] Example 1

[0039] According to this embodiment, a method embodiment of a one-time plot rule mining method for process event logs is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0040] The method embodiment provided in this embodiment can be executed in a server or similar computing device. Figure 2 FIG. 1 shows a hardware structure block diagram of a computing device for implementing a one-time plot rule mining method for process event logs. Figure 2 As shown, the computing device may include one or more processors (the processor may include but is not limited to a microprocessor MCU or a programmable logic device FPGA, etc.), a memory for storing data, a transmission device for communication functions, and an input / output interface. The memory, transmission device, and input / output interface are connected to the processor via a bus. In addition, it may also include: a display, a keyboard, and a cursor control device connected to the input / output interface. It will be understood by those skilled in the art that Figure 2 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 2 More or fewer components than shown, or with Figure 2 Different configurations shown.

[0041] It should be noted that the one or more processors and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry." The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuitry may be a single, independent processing module, or may be incorporated in whole or in part into any of the other components of the computing device. As described in the embodiments of the present disclosure, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0042] The memory can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the one-time plot rule mining method for process event logs in the embodiment of the present disclosure. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, realizing the one-time plot rule mining method for process event logs of the above-mentioned application. The memory may include a high-speed random access memory and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely located relative to the processor, and these remote memories may be connected to the computing device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0043] The transmission device is used to receive or send data via a network. Specific examples of the aforementioned network may include a wireless network provided by a communications provider of the computing device. In one embodiment, the transmission device includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0044] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computing device.

[0045] It should be noted that, in some optional embodiments, the above Figure 2 The computing device shown may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware and software elements. Figure 2 This is merely one example of a particular embodiment and is intended to illustrate the types of components that may be present in the computing device described above.

[0046] In the embodiment of the present invention, combined with Figure 1 As shown, considering that not all events recorded in the event sequence are core activities, they are usually accompanied by secondary activities. Since the focus is on core activities, the mining of meaningless events can be avoided by setting a lower limit on the time interval between events. For example, in the processing of metal parts, after the aluminum alloy parts are quenched, precision milling will not start immediately, and it will take 24 hours to eliminate internal stress. During this interval, secondary activities such as part transfer, labeling, and sampling inspection may occur. By setting a lower limit on the time interval, we can focus on the key processes between core activities (such as quenching and precision milling) and filter out trivial events. Therefore, this application studies the mining of plot rules with time interval constraints.

[0047] Furthermore, most frequency definitions used in episode mining (EM) count an event multiple times, leading to an overestimation of the event's frequency. For example, suppose an event sequence is (A,1)(B,2)(A,3)(C,4)(C,5). Traditional EM algorithms consider event A before C to have occurred four times, with timestamps <1,4>, <3,4>, <1,5>, and <3,5>. However, because each event is used multiple times, the number of event occurrences is overestimated. To address this issue, some algorithms calculate disjoint occurrences, where the minimum position of the next occurrence must be greater than the maximum position of the previous occurrence. Therefore, if occurrence <1,4> is chosen, occurrence <3,5> becomes illegal because 3 is less than 4. Consequently, this definition is overly strict and leads to the omission of important events, which can result in a lower episode frequency.

[0048] To avoid mining episodes with large intervals, this application employs an interval constraint. Furthermore, to prevent each event from being reused and to avoid excessive strictness, this application employs a one-time condition, which prohibits the same event from being used twice. Thus, this application proposes a mining scheme for strong one-time episode rules based on frequent one-time episodes with time interval constraints.

[0049] In the above operating environment, according to the first aspect of this embodiment, a one-time scenario rule mining method for process event logs is provided. Figure 3 A schematic diagram showing the framework of this method is shown in FIG. Figure 3 As shown, the method includes:

[0050] Step 1: Parse the process event log to generate the event sequence s, s = s1s2…s n =(e1,t1)(e2,t2)…(e n ,t n ), where s i Indicates an event, e i Indicates event activity, t i Timestamps indicating when the event occurred, t1 and t n are the start time and end time of s respectively;

[0051] Step 2: Mining the frequent one-time episodes with episode length m in the s to obtain the frequent one-time episode set F m , and create the F based on the timestamps of all events contained in each frequent one-time episode m Position index; wherein, the frequent one-time plot is a plot whose support in s is greater than or equal to the minimum support threshold, requiring that all events involved in multiple occurrences have time interval constraints and no overlap in occurrence time;

[0052] Step 3: Take the F mAs the initial input frequent set, iteratively execute the following steps until the preset stop condition is met:

[0053] (1) Based on the current input frequent set, generate a candidate episode set C with an episode length of m + 1 according to the preset episode connection strategy m+1 ; where the episode connection strategy is: if the prefix sub - episode of the frequent single - episode α is the same as the suffix sub - episode of the frequent single - episode β, then generate a candidate single - episode r with an episode length of m + 1

[0054] (2) Based on the position index, filter the candidate single - episodes in C m+1 with a support degree ≥ the minimum support degree threshold to obtain a frequent single - episode set F m+1 ;

[0055] (3) Filter the single - episode rules in F m+1 with a confidence degree ≥ the minimum confidence degree threshold; where the implicative form of the single - episode rule is β → α, α and β are two frequent single - episodes, and β is the prefix sub - episode of α, and the confidence degree of the single - episode rule is the ratio of the support degree of α to the support degree of β

[0056] (4) Use F m+1 as the new current input frequent set;

[0057] where the stop condition is that C <00​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​s = (D,1)(T,2)(T,3)(P,4)(E,5)(D,6)(T,7)(D,8)(E,9)(S,10)(D,11)(T,12)(D,13)(E,14).

[0063] In the embodiments of the present invention, the definition of an episode with a time interval is as follows:

[0064] An episode with a time interval can be expressed as α = e1[M,N]e2…e j [M,N]…[M,N]e m (1 < j ≤ m - 1, 0 ≤ M ≤ N), and it can also be abbreviated as α = e1e2…e m , where tgap = [M,N], and M and N are two non - negative integers, representing the minimum and maximum time - interval constraints respectively.

[0065] In the embodiments of the present invention, the definition of an occurrence and a single - occurrence is as follows:

[0066] Given an episode α = e1[M,N]e2…e j [M,N]…[M,N]e m , t = <t1,t2,…,t m > is an occurrence of α in the event sequence s if and only if for all j ∈ [1,m], e j occurs at time t j , satisfying 1 ≤ t1 < t2 < … < t m ≤ n, and M ≤ t j+1 - t j - 1 ≤ N. Given another occurrence t' = <t1',t2',…,t m '>, t and t' are two single - occurrences of the episode α if and only if 1 ≤ i,j ≤ m, t i ≠ t j '.

[0067] In the embodiments of the present invention, the following Example 3 is provided:

[0068] In Figure 1 , <1,3,6> is an occurrence of the episode α = D[0,2]T[0,2]D in s because the events at times t1, t3, t6 are D, T, D respectively, satisfying 0 ≤ 3 -​​In the embodiment of the present invention, the support and frequent one-time episodes (OOE) are defined as follows:

[0070] The number of one-time occurrences represents the support of episode α in s, denoted as sup(α,s). For example, the support of episode α in s in Example 3 is 2. If the support of an episode is greater than or equal to the minimum support threshold minsup given by the user, the episode is called a frequent one-time episode (OOE).

[0071] In the embodiment of the present invention, the OOE mining target is defined as follows:

[0072] Given an event sequence s, a time interval threshold tgap, and a support threshold minsup, the goal of OOE mining is to mine all frequent episodes.

[0073] In an embodiment of the present invention, the following Example 4 is provided:

[0074] exist Figure 1 In

[15] , when tgap = [0,2] and minsup = 3, OOE mining is to mine all frequent OOEs: {D,T,E,D[0,2]T,T[0,2]D,T[0,2]E,D[0,2]T[0,2]E}.

[0075] In the embodiment of the present invention, the definitions of prefix and suffix are as follows:

[0076] Suppose there is a scenario β=e1[M,N]e2…[M,N]e m and events c and d. If α = β[M,N]c, then β is called the prefix subplot of α, denoted by prefix(α) = β. Similarly, if γ = d[M,N]β, then β is called the suffix subplot of γ, denoted by suffix(γ) = β.

[0077] In this embodiment of the present invention, the definition of the plot rule is as follows:

[0078] An episodic rule is an implication of the form β→α, where α and β are two frequent one-time episodes (OOEs), and β is a prefix of α. The confidence of the rule β→α is denoted as conf(β→α), which is defined as the ratio of the support of α to the support of β, that is, conf(β→α) = sup(α,s) / sup(β,s). If the confidence of β→α is at least a minimum confidence threshold minconf, the episodic rule is called a strong OER (strong one-time episode rule).

[0079] In this embodiment of the present invention, OER mining is defined as follows:

[0080] The problem of this application is to mine all strong OERs in frequent OOEs based on the support threshold minsup and the minimum confidence threshold minconf.

[0081] In an embodiment of the present invention, the following Example 5 is provided:

[0082] In Example 4, D is a prefix OOE of D[0,2]T. Both D and D[0,2]T are frequent OOEs, with supports of 5 and 3, respectively. Therefore, the confidence of the rule D→D[0,2]T is conf(D→D[0,2]T) = 3 / 5 = 0.6. If minconf = 0.7, then D→D[0,2]T is not a strong OER. However, conf(T→T[0,2]D) = 3 / 4 = 0.75, which is greater than minconf, making it a strong OER. The set of strong OERs in Example 4 is R = {T→T[0,2]D, T→T[0,2]E, D[0,2]T→D[0,2]T[0,2]E}.

[0083] The main symbols used in the embodiments of the present invention are shown in Table 1.

[0084] Table 1

[0085] symbol describe s Sequence of Events α plot M,N Minimum and maximum time interval constraints sup(α,s) The support of plot α in event sequence s minsup Predefined minimum support threshold minconf Predefined minimum confidence threshold prefix(α) Prefix subplot of plot α suffix(α) Suffix subplot of plot α

[0086] In OER mining, the main task of the present invention is to discover strong OERs. The following first introduces the plot connection strategy for generating candidate plots, then describes the one-time plot support calculation method based on position index, and then proposes the OER-Miner algorithm for mining strong OERs based on frequent OOEs. Figure 3 The framework of the OER-Miner algorithm is presented.

[0087] The detailed steps of the present invention for generating candidate plots are as follows:

[0088] Many episode mining methods use an enumeration tree strategy to generate candidate episodes because this approach can generate all feasible candidate episodes. However, this strategy may generate a large number of unpromising candidate episodes. To overcome this shortcoming, this application proposes an episode connection strategy based on the Apriori property.

[0089] In the embodiment of the present invention, the following Theorem 1 is proposed:

[0090] The support satisfies anti-monotonicity, that is, sup(β,s)≥sup(α,s), where the plot β is a subplot of α.

[0091] The proof of Theorem 1 is as follows: Assume that the scenario α=e1[M,N]e2…[M,N]e m Since the plot β is a subplot of α, β can be written as e a[M,N]e a+1 …[M,N]e k , where 1 ≤ a ≤ k ≤ m. Assume sup(α, s) = n, which means that episode α has n single occurrences in s, i.e., <t 1,1 , t 1,2 , …, t 1,m >, <t 2,1 , t 2,2 , …, t 2,m >, …, <t n,1 , t n,2 , …, t n,m >. Therefore, <t 1,a , t 1,a+1 , …, t 1,k >, <t 2,a , t 2,a+1 , …, t 2,k >, …, <t n,a , t n,a+1 , …, t n,k > are n single occurrences of β in s. Therefore, the number of single occurrences of episode β in s is not less than n. Hence, sup(β, s) ≥ sup(α, s).

[0092] To prune unpromising candidate episodes, this application proposes a pruning strategy based on Theorem 1.

[0093] In an embodiment of the present invention, the pruning strategy is defined as: for an episode α, if sup(α, s) < minsup (the minimum support threshold), then α is not a single episode OOE, and neither is any of its super-episodes. Therefore, they can be pruned.

[0094] In an embodiment of the present invention, the following Example 6 is provided:

[0095] Consider Figure 1 the event sequence s therein, and use tgap = [0, 2], minsup = 3 to introduce the principle of the pruning strategy. The candidate episode enumeration tree generated using the enumeration strategy is as Figure 4 shown.

[0096] Calculate the support of episodes with length 1 using the enumeration strategy: sup('D', s) = 5, sup('T', s) = 4, sup('P', s) = 1, sup('E', s) = 3, sup('S', s) = 1. According to the pruning strategy, episodes 'P' and 'S' are pruned because their supports are lower than minsup. Therefore, episodes D[0, 2]P and D[0, 2]S can be pruned because they are super-episodes of 'P' and 'S' respectively.

[0097] From Figure 4As can be seen, the pruning strategy eliminates the extension of non-promising episodes. However, this candidate episode generation strategy still produces a lot of redundancy. This application adopts an episode connection strategy based on the Apriori property, which significantly reduces the number of generated candidate episodes.

[0098] In an embodiment of the present invention, the definition of plot connection is as follows:

[0099] Suppose there is a scenario β=e1[M,N]e2…[M,N]e m and event activities c, d. The plots α = β[M,N]c and γ = d[M,N]β are superplots of β. Since prefix(α) = suffix(γ) = β, a new superplot r can be generated by plot connection, that is,

[0100] This application uses Case 7 to demonstrate the principle of the plot connection strategy.

[0101] Example 7: Consider two plots γ = T[0,2]D[0,2]T and α = D[0,2]T[0,2]D. The prefix subplot of α and the suffix subplot of γ is D[0,2]T. Therefore, the superplot

[0102] In an embodiment of the present invention, the following Theorem 2 is provided:

[0103] The set F of all frequent episodes of length m+1 m+1 All of them are connected through the plot strategy from F of length m m Generated candidate episode set C m+1 , which means that the plot connection strategy is complete.

[0104] The proof of Theorem 2 is: Using the proof by contradiction, assuming the plot is a one-shot OOE of length m+1, but Does not belong to Candidate Episode Set C m+1 Plot The prefix subplot and suffix subplot of can be expressed as and and and There are also two one-time plot OOE. According to the definition of plot connection, and Can be generated through plot connection strategy Therefore, In set C m+1 In contradiction with the assumption.

[0105] In an embodiment of the present invention, Theorem 3 is provided: the episode connection strategy generates fewer candidate episodes than the enumeration tree strategy.

[0106] Theorem 3 is proved as follows: It is known that the enumeration tree strategy is complete. Now, we prove that the episode connection strategy generates fewer candidate episodes than the enumeration tree strategy. Assume that dβ is frequent and event activity c is frequent. According to the enumeration tree strategy, regardless of whether βc is frequent or not, a super-episode r = dβc can be generated. It can be determined that β is frequent because dβ is frequent. Therefore, one-shot episode mining satisfies anti-monotonicity. However, although β and c are frequent, it cannot be determined that βc is also frequent. Assuming that βc is infrequent, the episode connection strategy cannot be used to generate a super-episode r = dβc. Therefore, the episode connection strategy is stricter than the enumeration tree strategy because the episode connection strategy can only generate dβc when both dβ and βc are frequent, while the enumeration tree strategy only requires dβ to be frequent. Therefore, the episode connection strategy generates fewer candidate episodes than the enumeration tree strategy.

[0107] An illustrative example (ie, Example 8) is given below to illustrate the superiority of the plot connection strategy over the enumeration tree strategy.

[0108] Example 8: Consider Figure 1 Consider the event sequence s in [0,2], for tgap = [0,2] and minsup = 3. There are three length-2 OOEs, {D[0,2]T, T[0,2]D, T[0,2]E}. According to the enumeration tree strategy, there will be 3 × 3 = 9 candidate episodes of length 3, because 3 candidate episodes can be generated based on each length-2 OOE, and there are 3 length-2 OOEs, resulting in 9 candidate episodes of length 3. For example, three candidate episodes can be generated based on T[0,2]D: T[0,2]D[0,2]D, T[0,2]D[0,2]T, and T[0,2]D[0,2]E. Furthermore, since D[0,2]E is not a frequent OOE, according to the pruning strategy, the super-episode T[0,2]D[0,2]E is not a frequent OOE and can be pruned. However, based on the episode connection strategy, there are three candidate episodes of length 3, {D[0,2]T[0,2]D,D[0,2]T[0,2]E,T[0,2]D[0,2]T}. Therefore, the episode connection strategy outperforms the enumeration tree strategy.

[0109] In the embodiment of the present invention, the detailed steps of calculating the support are as follows:

[0110] To calculate support, some traditional algorithms first need to create a Nettree with multiple roots and multiple parent-child relationships, and then iteratively search for root-leaf paths, where each root-leaf path corresponds to an occurrence. However, Nettrees have many useless parent-child relationships. To solve this problem, this application proposes a Position Index (POE) algorithm for one-shot episodes. This algorithm uses position indexes to create nodes at each layer of the Nettree, and then iteratively discovers root-leaf paths using a depth-first search and backtracking strategy.

[0111] In this embodiment of the present invention, the location index is defined as follows:

[0112] Assume that the event sequence s has l different events. The position index I can be expressed as {k1:v1,k2:v2,…,k l :v l}, where key k a (1≤a≤l) stores an event, value v a Store event k a A list of locations where this occurred.

[0113] In an embodiment of the present invention, the following Example 9 is provided:

[0114] Consider Figure 1 Consider the event sequence s in [1]. Taking event 'D' as an example, it is easy to find that 'D' occurs at positions 1, 6, 8, 11, and 13. Therefore, the position index of 'D' is [1, 6, 8, 11, 13]. The corresponding position index of s is I = {'D': [1, 6, 8, 11, 13], 'T': [2, 3, 7, 12], 'P': [4], 'E': [5, 9, 14], 'S':

[10] }.

[0115] The pseudo code of the POE support calculation process is shown in Algorithm 1. The detailed steps of POE are summarized as follows:

[0116] Step 1: Create m layers of nodes based on position index, where m is the length of the episode.

[0117] Step 2: Find the first unused root node in the first layer (corresponding to lines 3-8 in the pseudocode).

[0118] Step 3: Assume that the node is found at layer j POE then searches for the first unused node at layer (j+1) If t j <t j+1 , and t j and t j+1 Satisfy the time interval constraint [M,N], that is, M≤t j+1 -tj -1≤N, then find the node at the (j+1)th layer And at the node and Create a parent-child relationship between them (corresponding to line 10 in the pseudocode).

[0119] Step 4: If the node is found in step 3 Then POE iterates step 3 until a node is found at the mth layer Otherwise, no Therefore, POE backtracks from the (j+1)th layer to the (j-1)th layer and finds a new node at the jth layer (corresponding to lines 11-18 in the pseudocode).

[0120] Step 5: If the node is found Then POE finds a <t1,t2,…,t m >, according to the one-time condition, the events in this occurrence cannot be reused. Iterate steps 2, 3, and 4 to find new occurrences until no new occurrences are found (corresponding to lines 20-23 in the pseudocode).

[0121]

[0122]

[0123] In the embodiment of the present invention, the operation process of the POE algorithm is explained using Example 10, which is provided as follows:

[0124] With Figure 1 Take the event sequence s and the plot α=e1[M,N]e2[M,N]e3=D[0,2]T[0,2]D in the example, the matching process is as follows Figure 5 shown.

[0125] According to Example 9, all the layer nodes of the Nettree of episode α in s are 'D':[1,6,8,11,13], 'T':[2,3,7,12] and 'D':[1,6,8,11,13], as shown in Figure 5 shown.

[0126] POE finds the root-leaf path from the first layer to the third layer as the emergence of plot α. is the first unused node in the first layer. POE then finds the node in the second layer. The first unused child node that satisfies the time interval constraint, i.e. the node of the node Because the nodes of the nodes There is no child node, POE backtracks to the first layer and finds the node of the node The second child node of Similarly, the nodes found As a node . Therefore, POE finds the first occurrence of <1,3,6>. Now, POE searches for the next unused node in the first layer. Cannot be selected because position 6 has already been used in occurrence <1,3,6>. POE then selects the node But it has no child nodes. Therefore, POE selects the node of the node It is easy to see that <11,12,13> is another occurrence.

[0127] In an embodiment of the present invention, an OER-Miner algorithm is proposed. The pseudo code of the OER-Miner mining process is shown in Algorithm 2. The OER-Miner mining process can be divided into the following steps:

[0128] Step 1: Traverse the event sequence, mine all frequent episodes (OOEs of length 1) stored in F1, and create their position indexes (corresponding to line 1 in the pseudocode).

[0129] Step 2: Use OOE set F m Generate a candidate episode set C of length m+1 m+1 (corresponding to lines 2-3 in the pseudocode).

[0130] Step 3: Calculate the set C m+1 If sup(α)≥minsup(minimum support threshold), then store the plot α in F m+1 (corresponding to lines 6-11 in the pseudocode).

[0131] Step 4: If the support of episode α is not less than the product of its prefix support and minconf (minimum confidence threshold), that is, α.support ≥ prefix(α).support × minconf, then prefix(α) → α is a strong OER (corresponding to lines 12-14 in the pseudocode).

[0132] Step 5: Repeat steps 2, 3, and 4 until C m+1 is empty (corresponding to lines 17-19 in the pseudocode).

[0133]

[0134]

[0135] In an embodiment of the present invention, the following Example 11 is provided to introduce the key steps of OER-Miner.

[0136] Example 11: Use Figure 1 The event sequence s is shown, and the time interval tgap is set to [0, 2], the minimum support minsup is set to 3, and the minimum confidence minconf is set to 0.7.

[0137] Step 1: Traverse the sequence s, find all frequent episodes (OOE of length 1), store them in F1 and create their position index. In Example 9, F1 = {D, T, E}, and its position index is: {'D': [1, 6, 8, 11, 13], 'T': [2, 3, 7, 12], 'E': [5, 9, 14]}.

[0138] Step 2: Generate a set of candidate episodes C2 of length 2 from F1 using the episode connection strategy. Note that the prefix and suffix sub-episodes of the episodes in F1 are both NULL. Therefore, event D can be expanded to D[0,2]D, D[0,2]T, and D[0,2]E. Similarly, event T can be expanded to T[0,2]T, T[0,2]E, and T[0,2]D; event E can be expanded to E[0,2]E, E[0,2]D, and E[0,2]T. The number of candidate episodes generated in this step is 3×3=9, which are stored in C2. C2 = {D[0,2]D, D[0,2]T, D[0,2]E, T[0,2]T, T[0,2]E, T[0,2]D, E[0,2]E, E[0,2]D, E[0,2]T}

[0139] Step 3: Calculate the support of the episodes in C2 and store the frequent OOEs. For example, the support of episode D[0,2]T is 3, which is no less than minsup, so D[0,2]T is a frequent OOE. Similarly, we obtain a set of frequent OOEs F2 = {D[0,2]T, T[0,2]D, T[0,2]E} of length 2.

[0140] Step 4: D[0,2]T is a frequent episode (sup = 3), and its prefix episode is D (sup = 5). Therefore, the confidence of the rule D→D[0,2]T is conf(D→D[0,2]T) = 3 / 5 = 0.6. Since the minimum confidence threshold minconf = 0.7, this rule is not a strong OER. The confidence of the rule T→T[0,2]D is conf(T→T[0,2]D) = 3 / 4 = 0.75, which is greater than minconf and is therefore a strong OER. The confidence of the rule T→T[0,2]E is conf(T→T[0,2]E) = 3 / 4 = 0.75, which is greater than minconf and is therefore a strong OER.

[0141] Step 5: Iterate steps 2, 3, and 4, using F2 to generate a set of candidate episodes C3 of length 3 using the episode connection strategy. C3 = {D[0,2]T[0,2]D, D[0,2]T[0,2]E}. Calculate the support of the episodes in C3 and store the frequent OOEs. For example, since suffix(D[0,2]T) = prefix(T[0,2]D) = T, the episode D[0,2]T[0,2]D is generated, with a support of 1, which is less than minsup and is not a frequent OOE. For another example, if suffix(D[0,2]T) = prefix(T[0,2]E) = T, the episode D[0,2]T[0,2]E is generated, with a support of 3, which is not less than minsup and is a frequent OOE. D[0,2]T[0,2]E is a frequent episode (sup = 3), and its prefix episode is D[0,2]T (sup = 3). Therefore, the confidence of the rule D[0,2]T→D[0,2]T[0,2]E is conf(D[0,2]T→D[0,2]T[0,2]E) = 3 / 3 = 1. Since the minimum confidence threshold minconf = 0.7, this rule is a strong OER. This gives us a frequent OER set F3 = {D[0,2]T[0,2]E} of length 3.

[0142] Step 6: Iterate steps 2, 3, and 4 again, using F3 to generate a candidate episode set C4 of length 4 using the episode connection strategy. Since F3 = {D[0,2]T[0,2]E}, the candidate episodes of length 4 generated by the episode connection strategy are empty, and thus the candidate episode set C4 is empty. Stop iteration;

[0143] Through iteration, we finally found 7 frequent OOEs: D, T, E, D[0,2]T, T[0,2]D, T[0,2]E, D[0,2]T[0,2]E. In addition, the set of strong one-shot episode rules R = {T→T[0,2]D, T→T[0,2]E, D[0,2]T→D[0,2]T[0,2]E}.

[0144] In an embodiment of the present invention, the following Theorem 4 is provided:

[0145] The time complexity of OER-Miner is O(n+n×K+K×log(K)), where n and K represent the length of the event sequence and the number of candidate episodes, respectively.

[0146] The proof of Theorem 4 is as follows: The time complexity of OER-Miner comes from two steps: creating the position index and calculating the support. Creating the position index requires scanning the event sequence, so the time complexity is O(n). For POE, each event is used at most once, which means that the time complexity of POE is O(n). Since there are K episodes in total, the time complexity is O(n × K). The time complexity of generating all candidate episodes is O(K × log(K)).

[0147] Therefore, the time complexity of OER-Miner is O(n+n×K+K×log(K)).

[0148] In an embodiment of the present invention, the following Theorem 5 is provided:

[0149] The space complexity of OER-Miner is O(n+m×K), where m is the length of the maximum candidate episode.

[0150] The proof of Theorem 5 is as follows: First, OER-Miner needs to scan the event sequence to create a position index, which is used to store the position of each event. Therefore, the space complexity of creating the position index is O(n). Furthermore, OER-Miner needs to store candidate episodes, which has a space complexity of O(m×K). Therefore, the total space complexity of OER-Miner is O(n+m×K).

[0151] In this embodiment of the present invention, the performance of OER-Miner is evaluated and its application is analyzed through experiments. The experiments are conducted on a Windows 10 64-bit computer equipped with a 3.70 GHz Intel(R) Core(TM) i7-8700K CPU and 16.0 GB of memory.

[0152] This application uses seven real process event logs and two simulation logs as experimental datasets. The detailed information of each log is shown in Table 2.

[0153] Table 2: Log details

[0154]

[0155]

[0156] OER-Miner can mine both frequent OOEs and strong OERs, and the time it takes to find strong OERs is almost negligible. Therefore, this application uses the time it takes to mine frequent OOEs to evaluate its efficiency. To comprehensively evaluate OER-Miner's mining capabilities, this application introduces the following algorithm (implemented using PyCharm 2020.2.3Pro, the code can be obtained from https: / / github.com / wuc567 / PatternMining / tree / master / OER-Miner):

[0157] 1. OER-NoPrun: This algorithm is designed to verify whether the pruning strategy significantly reduces the generation of candidate episodes with no prospects. The only difference from OER-Miner is that the pruning strategy is not applied.

[0158] 2. OER-Df and OER-Bf: To analyze the plot connection strategy, we introduce algorithms that use depth-first and breadth-first strategies respectively.

[0159] 3. Sow-H and MatchDB-O: To verify the computational efficiency of the POE algorithm, Sow-H (using one-way scanning to calculate the support of candidate episodes) and MatchDB-O (using MatchDB

[16] to calculate the support) are introduced.

[0160] 4. FEM-DFS: To compare the mining capabilities, the current advanced plot mining algorithm FEM-DFS

[45] is introduced.

[0161] 5.All-Rule: To explore the confidence of strong rules mined by OER-Miner, this algorithm is used to generate all OERs based on all frequent OOEs.

[0162] In an embodiment of the present invention, event logs are mainly collected by sensors and stored in log data format. Data needs to be converted from the current format to a process event log for further processing. A process event log is a simple collection of events, each of which can be represented as a triple (case, activity, timestamp), where case, activity, and timestamp represent a specific process instance, process behavior, and its occurrence time, respectively. For example, if someone performs the "Drinking" activity at 17:22:16 on October 3, 2023, the event is stored in the process event log as e=(Case1, Drinking, 10 / 03 / 2023

[0163] 17:22:16). To simplify event representation, "Case 1" is represented by 1, the "Drinking" activity is represented by 'D', and the time is represented by the timestamp t1. The event can be recorded as (1, D, t1). Table 3 shows a partial excerpt of the process event log.

[0164] Table 3: Process event log snippet

[0165] Case Activity Timestamp Case Activity Timestamp 1 D <![CDATA[10 / 03 / 2023 17:22:16(t1)]]> 1 P <![CDATA[10 / 03 / 2023 17:25:16(t4)]]> 2 D <![CDATA[11 / 03 / 2023 18:32:16(t1)]]> 2 T <![CDATA[11 / 03 / 2023 18:33:16(t2)]]> 1 T <![CDATA[10 / 03 / 2023 17:23:16(t2)]]> 1 E <![CDATA[10 / 03 / 2023 17:26:16(t5)]]> 1 T <![CDATA[10 / 03 / 2023 17:24:16(t3)]]> 2 E <![CDATA[11 / 03 / 2023 18:34:16(t3)]]>

[0166] According to Table 3, the event sequence is s1 = (D, t1)(T, t2)(T, t3)(P, t4)(E, t5) and s2 = (D, t'1)(T, t'2)(E, t'3).

[0167] To verify the effectiveness of OER-Miner, this application compares it with six comparison algorithms on ELog1-ELog8. Since OER-NoPrun, OER-Df, OER-Bf, Sow-H, MatchDB-O, and OER-Miner are all complete algorithms, they mine the same number of frequent OOEs. On ELog1-ELog8, when tgap = [0, 7200] and minsup is 20, 100, 134, 200, 230, 903, 170, and 310, respectively, the six algorithms mine 98, 36, 62, 74, 115, 17, 72, and 117 frequent OOEs, respectively. Since the definition of episode support in FEM-DFS is different from that in this application, to ensure fairness, its minsup is set to 8, 1100, 130, 160, 190, 240, 190, and 1210, respectively, which corresponds to mining 72, 21, 72, 65, 143, 14, 74, and 100 frequent episodes. Figure 6-Figure 8 The running time, number of candidate episodes, and memory usage of different algorithms on ELog1-ELog8 are shown respectively.

[0168] This application has the following observations:

[0169] 1. OER-Miner is faster than OER-NoPrun, which verifies the effectiveness of the pruning strategy. Figure 7 It can be seen that the number of candidate episodes generated by OER-Miner is less than that of OER-NoPrun; Figure 6 It shows a shorter runtime. For example, on ELog4, OER-NoPrun takes 12.59 seconds to generate 58,836 candidate episodes, while OER-Miner only takes 1.41 seconds to generate 1,428 candidate episodes. This is because the pruning strategy avoids the expansion of redundant episodes.

[0170] 2. OER-Miner outperforms OER-Df and OER-Bf, verifying that the plot connection strategy is better than the enumeration tree strategy. Figure 7 It shows that OER-Miner generates fewer candidate episodes. Figure 6It demonstrates faster runtimes. For example, on ELog2, OER-Df takes 62.12 seconds to generate 1,205 candidate episodes, OER-Bf takes 7.43 seconds to generate 227 candidate episodes, and OER-Miner takes only 1.27 seconds to generate 77 candidate episodes. This is because the depth-first and breadth-first strategies based on tree enumeration generate more candidate episodes than the strategy based on episode connections.

[0171] 3. OER-Miner is faster than Sow-H and MatchDB-O on all logs, indicating that the POE algorithm is more efficient in support calculation. For example, on ELog3, OER-Miner takes 0.79 seconds, while Sow-H and MatchDB-O take 9.27 seconds and 3.42 seconds, respectively. This is because Sow-H and MatchDB-O use linear search to find the next event location and perform multiple judgments, which is less efficient and has a higher time complexity than OER-Miner. Figure 8 When the number of events is large, OER-Miner's memory consumption is slightly higher than Sow-H and MatchDB-O (e.g., 264.08MB, 261.80MB, and 262.68MB, respectively, on ELog5). This is because OER-Miner needs to create a position index, which increases memory usage. Overall, OER-Miner significantly improves speed at the expense of slightly increasing memory usage.

[0172] 4. OER-Miner outperforms FEM-DFS. Despite using different support definitions, OER-Miner achieves lower runtime and memory usage when mining a similar number of episodes. For example, on ELog8, FEM-DFS requires 157.50 seconds and 188.73MB of memory, while OER-Miner only requires 6.26 seconds and 111.10MB of memory. This is because FEM-DFS needs to store occurrence positions and compare them one by one, resulting in longer runtime and greater memory usage.

[0173] To verify the impact of different minsup values ​​on OER-Miner's mining performance, this application selected OER-NoPrun, OER-Df, OER-Bf, Sow-H, and MatchDB-O as comparison algorithms and conducted experiments on ELog1. Setting tgap = [0, 7200] and minsup to 35, 30, 25, 20, 15, and 10, respectively, resulted in 33, 47, 66, 98, 162, and 390 episodes mined. Figure 9-11 The running time, number of candidate episodes, and memory usage for different minsup values ​​are shown.

[0174] This application has the following observations. As the minsup value decreases, the number of frequent OOEs, the number of candidate episodes, and the running time all increase. For example, when minsup = 25, the number of frequent OOEs mined by OER-Miner, the number of candidate episodes generated, and the running time are 66, 399, and 0.27 seconds, respectively; and when minsup = 15, the corresponding number of frequent OOEs, the number of candidate events, and the running time are 162, 976, and 1.05 seconds, respectively. Similar situations also occur in other algorithms. This is because when the minsup value decreases, according to the definition of support and frequent one-time episodes (OOE) and OOE mining, the number of frequent episodes will increase, and the number of candidate episodes generated will also increase accordingly, so the running time for calculating the support of these candidate events will also increase.

[0175] To verify the impact of different time gaps (tgaps) on OER-Miner's mining performance, this application conducted experiments on ELog7 using OER-NoPrun, OER-Df, OER-Bf, Sow-H, and MatchDB-O as comparison algorithms. Setting minsup = 170 and tgap to [0,7200], [0,7800], [0,8400], [0,9000], [0,9600], and [0,10200], respectively, resulted in 72, 93, 115, 138, 171, and 192 episodes mined, respectively. Figure 11-13 The running time, number of candidate episodes, and memory usage for different tgap values ​​are shown.

[0176] from Figure 13-14 The present application has the following observations:

[0177] As the time interval range increases, the number of frequent OOEs, the number of candidate episodes, and the runtime all increase. For example, when tgap = [0, 7200], OER-Miner mined 72 frequent OOEs and generated 277 candidate episodes in 1.17 seconds; while when tgap = [0, 7800], OER-Miner mined 93 frequent OOEs and generated 338 candidate episodes in 1.52 seconds. Similar results were observed in other algorithms. This is because, as the time interval range increases, according to the definition of episodes with time intervals and the definition of OOE mining, more episodes may become frequent OOEs and candidate episodes, and thus the runtime for calculating the support of these candidate episodes also increases. More importantly, OER-Miner outperforms the other five compared algorithms at all tgap values.

[0178] To verify the scalability of OER-Miner, this application compared OER-NoPrun, OER-Df, OER-Bf, Sow-H, and MatchDB-O algorithms, using ELog7 as the experimental data. ELog7_1 through ELog7_6 were created, with lengths 1, 2, 3, 4, 5, and 6 times that of ELog7, respectively. To eliminate the impact of frequent OOEs on runtime and memory usage, tgap = [0, 9000] was set for ELog7_1 through ELog7_6, and minsup was set to 170, 340, 510, 680, 850, and 1020, respectively. All algorithms mined 138 events. Figure 15 and Figure 16 Shows the runtime and memory usage for different log sizes.

[0179] The runtime and memory usage of OER-Miner grow slower than the size of the logs. For example, Elog7_6 is six times the size of Elog7_1. OER-Miner runs in 14.47 seconds and uses 114.22MB of memory on Elog7_6, which are 5.399 times the runtime (2.68 seconds) and 1.049 times the memory usage (108.90MB) of Elog7_1, respectively. Similar phenomena occur in all other logs. More importantly, OER-Miner outperforms other competing algorithms in terms of scalability. Figure 15 It can be seen that the runtime growth rates of OER-NoPrun, OER-Df, OER-Bf, Sow-H, and MatchDB-O are 15.93 / 2.77=5.751, 110.00 / 20.33=5.411, 35.67 / 6.51=5.479, 45.36 / 7.81=5.808, and 21.31 / 3.30=6.458, respectively, which are all higher than the growth rate of OER-Miner (14.47 / 2.68=5.399). Figure 16 As can be seen, OER-Miner's memory usage growth rate (114.22 / 108.90 = 1.049) is higher than that of the comparison algorithm. In short, OER-Miner has better scalability.

[0180] In this embodiment of the present invention, All-Rule was selected as the comparison algorithm and experiments were conducted on ELog1-ELog8. tgap was set to [0, 7200], and the minsup values ​​for ELog1-ELog8 were 20, 100, 134, 200, 230, 903, 170, and 310, respectively, and minconf values ​​were 0.35, 0.75, 0.8, 0.9, 0.9, 0.5, 0.65, and 0.75, respectively. Figure 17 and Figure 18 The comparison of the number of rules generated on ELog8 and the OER confidence is shown respectively.

[0181] Experimental results show that OER-Miner effectively reduces the number of rules. Figure 17 The results show that on ELog8, OER-Miner and All-Rule generate 52 and 110 rules respectively. This is because OER-Miner sets a minimum confidence threshold to retain only strong rules with high confidence, thus avoiding the generation of invalid rules. Figure 18 It shows that the rules mined by OER-Miner have higher confidence, indicating that the rules mined by it are more practical than those of All-Rule.

[0182] To explore the impact of different minconf values ​​on the number of rules and runtime, this application was compared with All-Rule on ELog8. Tgap was set to [0,7200], minsup to 310, and minconf to 0.60, 0.65, 0.70, 0.75, 0.80, and 0.85, respectively. Figure 19 The number of strong OERs under different minconf values ​​is shown.

[0183] The runtimes of OER-Miner and All-Rule are almost identical (6.310 seconds each) because strong OER determination consumes negligible time. As minconf increases, the number of strong OERs decreases, but the number of candidate events, frequent events, and total OERs remains constant. For example, when minconf = 0.60, the number of strong OERs is 83; when minconf = 0.85, the number of strong OERs decreases to 19. This is because the number of strong OERs is correlated with the minconf value, while the number of candidate events, frequent events, and total OERs are independent of minconf and therefore do not change with minconf.

[0184] In this case study, this application applies OER-Miner to complex process logs in real industrial scenarios. A company typically produces multiple products, and each product may involve multiple activities that perform the same function, that is, there are multiple production paths. Some production paths are relatively stable, while others have higher risks due to rework. It is worth noting that problematic activities are often not caused by themselves, but by their predecessor activities. Therefore, how to effectively identify risky paths and discover the optimal production process has become an urgent problem to be solved.

[0185] This application uses ELog9 as the experimental log, which contains 43 products, 55 activities, and 4543 events. To simplify the explanation, when mining OER, we only focus on events related to the product "Adjusting-nut", which involves 7 activities and 27 events. The Python tool PM4Py is used to draw a Petri net to describe the production process of this product, as shown in the following figure: Figure 20 As shown (each activity is coded as an integer). Figure 20 In the example, activity 3 or activity 6 and activity 2 will be selectively executed, and the rest of the process is the same.

[0186] Table 4 shows the OER set mined for the “Adjusting-nut” product using OER-Miner.

[0187] Table 4: OERs of “Adjusting nut” products

[0188] Strong OERs Rework (2)→(2,6),(2,6)→(2,6,6),(2,6,6)→(2,6,6,6),(2,6,6,6)→(2,6,6,6,6) Rework (2,7,4)→(2,7,4,5),(2,7,4,5)→(2,7,4,5,5),(2,7,4,5,5)→(2,7,4,5,5,5) No rework (3)→(3,5),(3,5)→(3,5,7),(3,5,7)→(3,5,7,1)

[0189] Table 4 shows that all OER rules can be divided into two categories: rules with rework and rules without rework (rework refers to repeated execution of an activity). For example, the first row in the table shows that after activity 2 is completed, activity 6 is repeated multiple times, indicating that activity 6 has rework. The second row shows that after activities 2, 7, and 4 are completed, activity 5 is repeated multiple times, indicating that activity 5 has rework. The rule sequence in the third row contains no repeated activities, indicating no rework. These results indicate that the rework is caused by the execution of activities 6 and 2, rather than activity 3, in the production path. This also indicates that activity 3 is preferred over activities 6 and 2 because it reduces the repetition and rework of subsequent activities. Furthermore, activities 6 and 2 require special attention as they may trigger adverse events in the process. Therefore, OER-Miner can discover and identify optimal paths and production paths that require optimization from process event logs, providing a basis for scientific production and decision-making in enterprises.

[0190] In summary, this application proposes an OER mining method to discover strong episodic rules from process event logs. To avoid mining excessive and meaningless episodes, OER mining uses a time interval constraint to limit the distance between two events. Furthermore, to prevent event duplication, this application uses a one-time condition when calculating event occurrences. This application proposes the OER-Miner algorithm for mining strong OERs based on frequent one-time episodes (OOEs). OER-Miner consists of three parts: (i) candidate episode generation, (ii) support calculation, and (iii) rule generation. In the candidate episode generation phase, an episode connection strategy based on the Apriori property is employed to eliminate redundant candidate episode extensions. In the support calculation phase, a POE algorithm is proposed, which calculates support using a depth-first search and backtracking strategy on a position index. In the rule generation phase, all frequent OOEs are first mined, and then strong OERs are selected based on minimum support (minsup) and minimum confidence (minconf). To validate the mining efficiency of OER-Miner, this application selected seven real process event logs and two simulated logs and compared them with seven comparison algorithms. Experimental results show that OER-Miner outperforms other comparison algorithms on all logs. More importantly, case studies show that OER-Miner can be applied to real industrial logs to identify risk paths and discover optimal production processes. Figure 1 As shown, according to a second aspect of this embodiment, a storage medium is provided, wherein the storage medium includes a stored program, wherein when the program is run, a processor executes any one of the above methods.

[0191] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present invention is not limited by the order of the actions described, because according to the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present invention. Through the description of the above embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus a necessary general hardware platform, and of course, by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, disk, CD), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0192] Example 2

[0193] Figure 21 FIG. 4 shows a one-time plot rule mining device for process event logs according to this embodiment, which corresponds to the method according to embodiment 1. Figure 21 As shown, the device includes: an event sequence generating module 2110, which is used to parse the process event log and generate an event sequence s, s = s1s2 ... s n =(e1,t1)(e2,t2)…(e n ,t n ), where s i Indicates an event, e i Indicates event activity, t i Timestamps indicating when the event occurred, t1 and t n are the start time and end time of s respectively; the frequent one-time episode mining module 2120 is used to mine the frequent one-time episodes with an episode length of m in s, and obtain the frequent one-time episode set F m , and create the F based on the timestamps of all events contained in each frequent one-time episode m Position index of the frequent one-time plot; wherein, the support degree of the frequent one-time plot in the s is greater than or equal to the minimum support threshold, requiring that all events involved in multiple occurrences have time interval constraints and no overlap in the occurrence time; the strong one-time plot rule mining module 2130 is used to extract the frequent one-time plot rule with the F m As the initial input frequent set, iteratively perform the following steps until the preset stopping condition is met: Based on the current input frequent set, generate a one-time candidate episode set C with an episode length of m+1 according to the preset episode connection strategy m+1 Wherein, the plot connection strategy is: if the prefix subplot of the frequent one-time plot α is the same as the suffix subplot of the frequent one-time plot β, then generate a candidate one-time plot r with a plot length of m+1, Based on the position index, filter the C m+1 The candidate one-time episodes with medium support ≥ the minimum support threshold are used to obtain the frequent one-time episode set F m+1 ; Screening the F m+1 A one-time plot rule with confidence ≥ minimum confidence threshold; wherein the implication of the one-time plot rule is β→α, α and β are two frequent one-time plots, and β is a prefix subplot of α, and the confidence of the one-time plot rule is the ratio of the support of α to the support of β; the F m+1 As the new current input frequent set; wherein, the stopping condition is the C m+1 Empty.

[0194] Therefore, according to this embodiment, a time interval constraint is used to limit the distance between two events, and a one-time condition is used to avoid the reuse of events, so as to avoid mining too many meaningless plots. In addition, when mining strong one-time plot rules (i.e., one-time plot rules with confidence ≥ minimum confidence threshold), a triple mechanism of candidate plot generation, support calculation and rule generation is proposed. In the candidate plot generation stage, a plot connection strategy is adopted to eliminate redundant candidate plot extensions; in the support calculation stage, support is calculated based on the position index; in the rule generation stage, all frequent one-time plots are first mined, and then strong one-time plot rules are screened according to the minimum support threshold and the minimum confidence threshold. The technical effect of efficiently and accurately mining strong one-time plot rules is achieved. This solves the technical problem that the plot mining methods in the prior art generally do not consider time interval constraints and overestimate the frequency of plots, resulting in the mining of a large number of meaningless plots.

[0195] Example 3

[0196] Figure 22 FIG. 4 shows a one-time plot rule mining device for process event logs according to this embodiment, which corresponds to the method according to embodiment 1. Figure 22 As shown, the device includes: a processor 2210; and a memory 2220, connected to the processor 2210, for providing the processor 2210 with instructions for processing the following processing steps: parsing the process event log to generate an event sequence s, s = s1s2 ... s n =(e1,t1)(e2,t2)…(e n ,t n ), where s i Indicates an event, e i Indicates event activity, t i Timestamps indicating when the event occurred, t1 and t n are the start time and end time of s respectively; mine the frequent one-time episodes with episode length m in s and obtain the frequent one-time episode set F m , and create the F based on the timestamps of all events contained in each frequent one-time episode m Position index; wherein, the frequent one-time plot is a plot with a support degree in the s greater than or equal to the minimum support threshold, requiring that all events involved in multiple occurrences have time interval constraints and no overlap in occurrence time; with the F m As the initial input frequent set, iteratively perform the following steps until the preset stopping condition is met: Based on the current input frequent set, generate a one-time candidate episode set C with an episode length of m+1 according to the preset episode connection strategy m+1; Wherein, the plot connection strategy is: if the prefix subplot of the frequent one-time plot α is the same as the suffix subplot of the frequent one-time plot β, then generate a candidate one-time plot r with a plot length of m+1, r=γ⊕α; based on the position index, filter the C m+1 The candidate one-time episodes with medium support ≥ the minimum support threshold are used to obtain the frequent one-time episode set F m+1 ; Screening the F m+1 A one-time plot rule with confidence ≥ minimum confidence threshold; wherein the implication of the one-time plot rule is β→α, α and β are two frequent one-time plots, and β is a prefix subplot of α, and the confidence of the one-time plot rule is the ratio of the support of α to the support of β; the F m+1 As the new current input frequent set; wherein, the stopping condition is the C m+1 Empty.

[0197] Therefore, according to this embodiment, a time interval constraint is used to limit the distance between two events, and a one-time condition is used to avoid the reuse of events, so as to avoid mining too many meaningless plots. In addition, when mining strong one-time plot rules (i.e., one-time plot rules with confidence ≥ minimum confidence threshold), a triple mechanism of candidate plot generation, support calculation and rule generation is proposed. In the candidate plot generation stage, a plot connection strategy is adopted to eliminate redundant candidate plot extensions; in the support calculation stage, support is calculated based on the position index; in the rule generation stage, all frequent one-time plots are first mined, and then strong one-time plot rules are screened according to the minimum support threshold and the minimum confidence threshold. The technical effect of efficiently and accurately mining strong one-time plot rules is achieved. This solves the technical problem that the plot mining methods in the prior art generally do not consider time interval constraints and overestimate the frequency of plots, resulting in the mining of a large number of meaningless plots.

[0198] The serial numbers of the above embodiments of the present invention are for description only and do not represent the advantages or disadvantages of the embodiments.

[0199] In the above embodiments of the present invention, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0200] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0201] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0202] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0203] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk, etc. Various media that can store program codes.

[0204] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.

Claims

1. A one-time plot rule mining method for process event logs, characterized in that: include: Parse the process event log to generate the event sequence s, s = s1s2…s n =(e1,t1)(e2,t2)…(e n ,t n ), where s i Indicates an event, e i Indicates event activity, t i Timestamps indicating when the event occurred, t1 and t n are the start time and end time of s respectively; Mining the frequent one-time episodes with episode length m in the s, and obtaining the frequent one-time episode set F m , and create the F based on the timestamps of all events contained in each frequent one-time episode m Position index; wherein, the frequent one-time plot is a plot whose support in s is greater than or equal to the minimum support threshold, requiring that all events involved in multiple occurrences have time interval constraints and no overlap in occurrence time; With the F m As the initial input frequent set, iteratively perform the following steps until the preset stopping condition is met: Based on the current input frequent set, generate a one-time candidate episode set C with an episode length of m+1 according to the preset episode connection strategy m+1 ; Wherein, the plot connection strategy is: if the prefix subplot of the frequent one-time plot α is the same as the suffix subplot of the frequent one-time plot β, then generate a candidate one-time plot r with an episode length of m+1, r=γ⊕α; Based on the position index, filter the C m+1 The candidate one-time episodes with medium support ≥ the minimum support threshold are used to obtain the frequent one-time episode set F m+1 ; Screening the F m+1 A one-shot plot rule with medium confidence ≥ minimum confidence threshold; wherein the implication of the one-shot plot rule is β→α, α and β are two frequent one-shot plots, and β is a prefix subplot of α, and the confidence of the one-shot plot rule is the ratio of the support of α to the support of β; The F m+1 As the new current input frequent set; Wherein, the stopping condition is the C m+1 Empty.

2. The method according to claim 1, characterized in that Mining the frequent one-time episodes with episode length m in the s, and obtaining the frequent one-time episode set F m Operations include: Define the scenario α as a scenario with time intervals, α = e1[M,N]e2…e j [M,N]…[M,N]e m (1 < j ≤ m - 1, 0 ≤ M ≤ N), where M and N are two non - negative integers, representing the minimum and maximum time interval constraints respectively; Define the scenario α=e1[M,N]e2…e j [M,N]…[M,N]e m , t= <t1,t2,…,t m > is an occurrence of α in the event sequence s if and only if for all j∈[1,m], e j In t j Occurs at the moment, satisfying 1≤t1 <t2<…<t m ≤n, and M≤t j+1 -t j -1≤N; Another occurrence of defining the plot α is t'= <t1',t2',…,t m '>, if and only if 1≤i,j≤m, t i ≠t j ', t and t' are two one-time occurrences of episode α; Define the support of plot α in the s as the total number of one-time occurrences of plot α; Traverse all events in s and calculate the total number of one-time occurrences of each episode of length m as the support of the corresponding episode in s; and Filter the episodes with support ≥ minimum support threshold from all episodes of length m as frequent one-time episodes, and get the frequent one-time episode set F m .

3. The method according to claim 2, characterized in that Based on the current input frequent set, generate a one-time candidate episode set C with an episode length of m+1 according to the preset episode connection strategy m+1 Operations include: Define a scenario β=e1[M,N]e2…[M,N]e m and event activities c and d, the plots α = β[M,N]c and γ = d[M,N]β are superplots of β, and since prefix(α) = suffix(γ) = β, a new superplot r is generated by plot connection, that is, r = γ ⊕ α = d[M,N]β[M,N]c; where prefix(α) is the prefix subplot of plot α, and suffix(γ) is the suffix subplot of plot γ; and Determine whether the prefix subplot of the first plot and the suffix subplot of the second plot in the current input frequent set are the same. If they are the same, generate a candidate one-time plot r with a plot length of m+1 according to the definition r=γ⊕α, otherwise, generate a candidate one-time plot r with a plot length of m+1.

4. The method according to claim 3, characterized in that Based on the position index, filter the C m+1 The candidate one-time episodes with medium support ≥ the minimum support threshold are used to obtain the frequent one-time episode set F m+1 Operations include: Based on the position index, calculate each candidate one-time episode in C m+1 Support in According to the calculated support, the C m+1 The candidate one-time episodes with medium support ≥ the minimum support threshold are used to obtain the frequent one-time episode set F m+1 .

5. The method according to claim 4, characterized in that Based on the position index, calculate each candidate one-time episode in C m+1 The operations on support in include: Step 1: Create m+1 layer nodes based on the position index, where m+1 is the C m+1 the length of the candidate one-shot episode; Step 2: Find the first unused root node in the first layer Step 3: Assume that the node is found at layer j Search for the first unused node in the (j+1)th layer If t j <t j+1 , and t j and t j+1 Satisfy the time interval constraint [M,N], that is, M≤t j+1 -t j -1≤N, then find the node at the (j+1)th layer And at the node and Create a parent-child relationship between them; Step 4: If the node is found in step 3 Then iterate step 3 until a node is found at the mth layer Otherwise, no Backtrack from the (j+1)th layer to the (j-1)th layer and find a new node in the jth layer; Step 5: If the node is found Then find an occurrence <t1,t2,…,t m >, according to the one-time condition, the events in this occurrence cannot be reused; iterate steps 2, 3, and 4 to find new occurrences until no new occurrences are found; Step 6: Place each candidate one-shot scenario in the C m+1 The total number of occurrences of the candidate one-shot episode in the C m+1 The support in .

6. The method according to claim 5, characterized in that Define a scenario β=e1[M,N]e2…[M,N]e m and event activities c and d, if α=β[M,N]c, then β is called the prefix subplot of α, denoted by prefix(α)=β, if γ=d[M,N]β, then β is called the suffix subplot of γ, denoted by suffix(γ)=β.

7. The method according to claim 6, characterized in that Screening the F m+1 The operations for one-shot plot rules with medium confidence ≥ minimum confidence threshold include: The implication of the plot rule is defined as β→α, where α and β are two frequent one-time plots, and β is a prefix subplot of α. The confidence of the one-time plot rule β→α is recorded as conf(β→α), and conf(β→α) is defined as the ratio of the support of α to the support of β, that is, conf(β→α)=sup(α,s) / sup(β,s), where sup(α,s) is the support of α in s, and sup(β,s) is the support of β in s; According to the definition conf(β→α)=sup(α,s) / sup(β,s), calculate the F m+1 The confidence level of all one-shot plot rules in ; According to the calculated confidence, the F m+1 One-shot episode rule with medium confidence ≥ minimum confidence threshold.

8. A storage medium, characterized in that: The storage medium includes a stored program, wherein when the program is executed, the processor executes the method according to any one of claims 1 to 7.

9. A one-time plot rule mining device for process event logs, characterized in that: include: The event sequence generation module is used to parse the process event log and generate the event sequence s, s = s1s2…s n =(e1,t1)(e2,t2)…(e n ,t n ), where s i Indicates an event, e i Indicates event activity, t i Timestamps indicating when the event occurred, t1 and t n are the start time and end time of s respectively; Frequent one-time episode mining module, used to mine the frequent one-time episodes with episode length m in the s, and obtain the frequent one-time episode set F m , and create the F based on the timestamps of all events contained in each frequent one-time episode m Position index; wherein, the frequent one-time plot is a plot whose support in s is greater than or equal to the minimum support threshold, requiring that all events involved in multiple occurrences have time interval constraints and no overlap in occurrence time; Strong one-time plot rule mining module for the F m As the initial input frequent set, iteratively perform the following steps until the preset stopping condition is met: Based on the current input frequent set, generate a one-time candidate episode set C with an episode length of m+1 according to the preset episode connection strategy m+1 ; Wherein, the plot connection strategy is: if the prefix subplot of the frequent one-time plot α is the same as the suffix subplot of the frequent one-time plot β, then generate a candidate one-time plot r with an episode length of m+1, r=γ⊕α; Based on the position index, filter the C m+1 The candidate one-time episodes with medium support ≥ the minimum support threshold are used to obtain the frequent one-time episode set F m+1 ; Screening the F m+1 A one-shot plot rule with medium confidence ≥ minimum confidence threshold; wherein the implication of the one-shot plot rule is β→α, α and β are two frequent one-shot plots, and β is a prefix subplot of α, and the confidence of the one-shot plot rule is the ratio of the support of α to the support of β; The F m+1 As the new current input frequent set; Wherein, the stopping condition is the C m+1 Empty.

10. A one-time plot rule mining device for process event logs, characterized in that: include: processor; as well as A memory, connected to the processor, configured to provide the processor with instructions for processing the following processing steps: Parse the process event log to generate the event sequence s, s = s1s2…s n =(e1,t1)(e2,t2)…(e n ,t n ), where s i Indicates an event, e i Indicates event activity, t i Timestamps indicating when the event occurred, t1 and t n are the start time and end time of s respectively; Mining the frequent one-time episodes with episode length m in the s, and obtaining the frequent one-time episode set F m , and create the F based on the timestamps of all events contained in each frequent one-time episode m Position index; wherein, the frequent one-time plot is a plot whose support in s is greater than or equal to the minimum support threshold, requiring that all events involved in multiple occurrences have time interval constraints and no overlap in occurrence time; With the F m As the initial input frequent set, iteratively perform the following steps until the preset stopping condition is met: Based on the current input frequent set, generate a one-time candidate episode set C with an episode length of m+1 according to the preset episode connection strategy m+1 ; Wherein, the plot connection strategy is: if the prefix subplot of the frequent one-time plot α is the same as the suffix subplot of the frequent one-time plot β, then generate a candidate one-time plot r with an episode length of m+1, r=γ⊕α; Based on the position index, filter the C m+1 The candidate one-time episodes with medium support ≥ the minimum support threshold are used to obtain the frequent one-time episode set F m+1 ; Screening the F m+1 A one-shot plot rule with medium confidence ≥ minimum confidence threshold; wherein the implication of the one-shot plot rule is β→α, α and β are two frequent one-shot plots, and β is a prefix subplot of α, and the confidence of the one-shot plot rule is the ratio of the support of α to the support of β; The F m+1 As the new current input frequent set; Wherein, the stopping condition is the C m+1 Empty.

Citation Information

Cited By

  • Intelligent optimization method and device for business process of cloud edge collaboration, and storage medium

    CN120979942A

  • A cloud-edge collaborative business process intelligent optimization method and device and storage medium

    CN120979942B