Business process anomaly detection method based on event attribute graph
By defining the event attribute diagram and extracting structural information using Weisfeiler-Lehman algorithm and PV-DBOW model, a prediction model for the next event attribute value is solved, and the problem of difficulty in detecting data flow abnormalities in the business process in the existing technology is solved, and more accurate abnormality detection and positioning is achieved.
Patent Information
- Application Number
- CN202510031636.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art is difficult to effectively detect data flow exceptions in business processes in uncertain and variable business environments, and traditional methods cannot locate the specific event attribute values that trigger the exception.
A business process exception detection method based on event attribute graph embedding is proposed. By defining event attribute graphs, extracting structural information and dependencies using Weisfeiler-Lehman algorithm and PV-DBOW model, building a prediction model for the next event attribute value, and then discovering exception event attributes.
It effectively improves the accuracy of abnormal detection of business processes, can locate abnormal locations in the trajectory, and is suitable for complex and diverse business environments.
Smart Images

Figure CN119939472A_ABST
Abstract
Description
(I) Technical field
[0001] The invention belongs to the field of business process management and is a business process anomaly detection method based on event attribute graph. (II) Background technology
[0002] At present, modern enterprises and organizations have widely used process aware information system (PAIS) to manage business processes. With the increasing complexity of business processes and the increasing diversification of business needs, it is particularly important to ensure the correct operation of business processes. At the same time, the probability of business process anomalies in process aware information systems is gradually increasing. This brings greater challenges to the operation and management of business processes in enterprises and organizations. Business process anomalies refer to the occurrence of business process instances that do not conform to the business process definition, violate business process constraints or are unreasonable. These anomalies may cause business processing delays, resource waste, and reduced customer satisfaction in modern enterprises and organizations, and even cause huge losses. Therefore, it is urgent to introduce efficient business process anomaly detection technology in PAIS to meet current challenges and ensure the safe and efficient operation of business processes in enterprises and organizations. Business process anomaly detection refers to the use of anomaly detection technology to discover and identify anomalies in business process execution in business process management, helping managers to deal with potential problems in a timely manner, optimize resource allocation, and ensure business compliance, thereby providing a solid guarantee for the smooth operation of business processes. With the increasing complexity of business processes, it is particularly important to ensure their correctness.
[0003] Traditional anomaly detection methods based on business process models can no longer meet the needs of modern enterprises and organizations for business process management technology in uncertain and changing business environments. This type of method usually adopts consistency detection technology, with the standard reference business process model as the core basis for anomaly detection. Lu et al. proposed a consistency detection method based on partially ordered events, and Adriansyah et al. proposed a consistency detection method based on cost adaptability analysis. Both methods align the traces in the event log with the standard reference business process model to detect anomalies. This type of method can detect control flow anomalies in the business process, but because it does not use the attribute information of the event, it cannot detect data flow anomalies in the business process.
[0004] In recent years, researchers have introduced machine learning methods, especially deep learning, into the discovery and identification of business process anomalies, and proposed a variety of business process anomaly detection methods. Depending on whether the samples used are labeled, these methods are divided into supervised methods and unsupervised methods. Methods based on association rule mining belong to supervised methods. They capture the typical behavior of business processes by formulating rules, thereby identifying normal business process behaviors and classifying trajectories that do not conform to these rules as anomalies. Sarno et al. used association rule mining algorithms to discover anomalies in business processes. The association rules generated by the training data set show the confidence of the trajectory. The disadvantage is that rules need to be developed manually (for example, users define the expected maximum activity time). Subsequently, Sarno et al. proposed an anomaly detection method based on fuzzy association rule learning, combining process mining, fuzzy multi-attribute decision making and fuzzy association rule learning to detect anomalies in business processes. This method first uses process mining technology to discover the business process model, then applies consistency detection technology and fuzzy multi-attribute decision making methods to calculate the anomaly rate, and finally generates rules for detecting abnormal trajectories through fuzzy association rule learning. These methods can generate a set of rules for detecting abnormal trajectories, but in business process management, overly strict rules may affect the accuracy of anomaly detection results.
[0005] Nolle et al. pointed out that not only the deviation of the control flow of the business process from the predetermined mode will cause business process anomalies, but also data flow errors, that is, abnormal attribute values of events, will cause business process anomalies. Sun Xiaoxiao et al. proposed a multi-angle online anomaly detection method based on context awareness. This method uses the Split Miner algorithm to mine the process model, uses the replay technology on the business process model to generate the behavioral context of the business process, and uses other event attributes as data context to perform anomaly detection from three perspectives: behavior, time, and attributes. HuoS et al. used a GNN-based autoencoder to extract the structural information of the business process instance and proposed a business process anomaly detection method based on a graph autoencoder. This method uses the activities of events in the trajectory as nodes and the dependency relationships between activities in the trajectory as edges to generate a directed graph for representing the trajectory. The methods of Sun Xiaoxiao et al. and HuoS et al. detect anomalies of business process instances from multiple perspectives, which can detect both control flow anomalies and data flow anomalies of business process instances. However, these two methods can only determine whether the business process instance is abnormal, but cannot locate the attribute values of the specific events that cause the business process instance anomalies.
[0006] Since the proportion of abnormal trajectories in event logs is small and the number of normal and abnormal trajectories is unbalanced, supervised business process anomaly detection methods using unbalanced trajectory data are not the best choice. In addition, event logs cannot cover all possible abnormal trajectories of business processes, and supervised methods are limited to identifying abnormal trajectories in existing event logs. Nolle et al. proposed an anomaly detection method based on denoising autoencoders. This method takes the event sequence in the trajectory as input data and trains the autoencoder to reconstruct the sequence. If the reconstruction error of the input and output exceeds the threshold, the trajectory is considered to be an abnormal trajectory. The disadvantage is that it can only detect anomalies in the control flow, but cannot discover and identify abnormal event attribute values. Subsequently, Nolle et al. added data streams to the input of the autoencoder. This method can detect abnormal event attribute values in business processes, but its performance on data sets with long trajectory sequences is poor. Nolle et al. also proposed a multivariate anomaly detection method BINet considering event attributes. This method uses LSTM to build a model to predict the probability of the next event of a business process instance. It further improves the accuracy of anomaly detection, but its performance on event logs with longer trajectories is still poor. Guan et al. proposed a GNN-based anomaly detection method, which constructs the trajectory into multiple event attribute graphs through the control flow perspective and the discrete attribute values, uses the GNN encoder to extract the trajectory feature vector based on the multi-graph, and uses the GRU decoder to reconstruct the difference. Guan et al.'s method can effectively improve the accuracy of anomaly detection by extracting the business process structure information through GNN, and achieves better results in anomaly detection accuracy than the anomaly detection method based on sequence autoencoders. However, this method relies on a manually set activity co-occurrence frequency threshold to discover the structural information of the business process. The size of this threshold has an impact on the performance of anomaly detection; and the GNN encoder does not consider the event sequence information in the trajectory. (III) Summary of the invention
[0007] Since the trajectory in the event log of the business process is not a flat structural data, directly using the attributes of the event to encode the trajectory in a sequence form will ignore the dependency between the structural information of the business process and the event attribute values, so that the behavior of the business process cannot be fully characterized, which in turn affects the accuracy of business process anomaly detection. To this end, the present invention proposes a business process anomaly detection method based on event attribute graph embedding. The concept of event attribute graph is defined, and the trajectory is represented as multiple event attribute graphs. The Weisfeiler-Lehman algorithm and PV-DBOW model are used to extract the dependency between the structural information of the business process and the event attribute values, and a prediction model for the next event attribute value of the business process instance is constructed. The model is used to predict the attribute probability distribution of the next event of the business process instance, and the abnormal event attributes are discovered through the attribute values of the actual event and the predicted probability.
[0008] The technical contents of the present invention are as follows:
[0009] Step 1: Extract business process structure information
[0010] (1) Generate event attribute graph The event attribute graph of the event log trajectory can describe the structural information of the business process from multiple event attribute perspectives. In order to fully explore the dependencies between the structural information of the business process and the event attributes, the present invention uses multiple event attribute graphs to represent a trajectory from multiple event attribute perspectives. For the continuous attributes of the trajectory, its attribute values can be discretized and converted into discrete attributes, and then the event attributes can be converted into event attribute graphs. Therefore, the present invention currently mainly studies the method of representing the discrete attributes of the event into an event attribute graph. The event attribute graph of a trajectory is an undirected graph. For each discrete attribute (including activity attributes) in the trajectory, an event attribute graph can be constructed. The event attribute graph defined in the present invention is as shown in Definition 1. Definition 1 (Event Attribute Graph) The event attribute graph of the event attribute att in the trajectory tr is a binary pair G att =(V att ,E att ).in, V att is a set of attribute values of event attributes att in trajectory tr, indicating that G att Node collection; Represents G att A set of edges, each edge corresponds to an attribute-value pair with a dependency relationship of attribute att. Assume that the set AS is the discrete attribute set of all events in the event log, and the trajectory tr = <e 1 ,…,e n>, the steps for constructing the event attribute graph are as follows. (1) Obtain the node set of the event attribute graph. Initialize the event attribute graph G using the attribute value set of the event attribute att∈AS of the trajectory tr. att The node set V att If event e j If there is no attribute att, it will be filled with 0. (2) Generate attribute-value pairs of event attributes and traverse each event e in the trajectory tr i , the event e i and e i+1 The attribute value of each attribute att Forms a pair of attribute value pairs. (3) Obtain the edge set of the event attribute graph for each attribute value pair In the event attribute graph G att Generate a line from point to edge. Table 1 is a fragment of an event log. Each row represents an event, including case ID, event ID, activity name, time, resources and other attributes. Case ID is used to identify a sequence of events, or trajectory, that occurs sequentially. For example, in the case trajectory with CaseID 142857 in Table 1, the attribute value pairs obtained from the activity attributes of the event are (A, B), (B, C), (C, B), (B, C), (C, D), (D, C), (C, B) and (B, E), and the attribute pairs corresponding to other attributes are obtained in a similar way. According to the above-mentioned event attribute graph construction method, the event attribute graphs of the activity attributes and resource attributes of the event of trajectory #142857 can be obtained respectively, as shown in Figure 1. Figure 2 shown. Table 1 A fragment of an event log
[0011] (2) Constructing a subtree pattern corpus of event attribute graphs According to Definition 1 and the method for constructing event attribute graphs, each track of the event log can be converted into multiple event attribute graphs to form a set of event attribute graphs. Using the WL algorithm to extract subtree patterns of different heights in the event attribute graph, a 1-hop, 2-hop, ..., m-hop subtree pattern corpus can be constructed, where m is the number of iterations of the WL algorithm.
[0012] (3) Generate semantic feature vectors of event attribute graph nodes The PV-DBOW model learns the feature vector of the subtree pattern from the subtree pattern corpus of the event attribute graph, so that the feature vector of the event attribute graph node generated finally contains both the structural information between nodes and the semantic information of the nodes. The subtree pattern in the event attribute graph can be regarded as a word, and the sequence composed of different subtree patterns is similar to a sentence composed of different words. i Training the PV-DBOW model above can get a matrix The matrix Mat i The nth row of the i The semantic embedding of the nth subtree pattern in , 1≤i≤m, 1≤n≤len(tr), len(tr) is the length of the trajectory tr, and dim is the feature dimension. The process of using the PV-DBOW model to generate the semantic feature vector of the event attribute graph node is as follows Figure 3 shown.
[0013] Step 2: Generate the feature vector of the trajectory The event attribute values contained in the trajectory are sometimes repeated, and these repeated values correspond to the same node in the event attribute graph. However, it is worth noting that the same event attribute value in the trajectory has different meanings at different locations, which is crucial for analyzing and reconstructing business process behaviors. After extracting the structural information of the business process, the event attribute values in the trajectory are mapped to the feature vectors of the event attribute graph nodes, and the feature vector of the trajectory is constructed accordingly.
[0014] Step 3: Build a prediction model for the next event attribute value The present invention mainly uses the gated recurrent unit GRU, encoder-decoder structure and dot product attention mechanism in the recurrent neural network to build a prediction model for the attribute value of the next event, and uses the feature vector of the trajectory to train this model. The deep network structure of the prediction model is as follows Figure 4 shown.
[0015] Step 4: Trajectory abnormality determination Trajectory anomaly determination requires anomaly scores of event attribute values in the trajectory. Calculating the anomaly score of event attribute values in a business process requires using the prediction results of the next event attribute value prediction model, that is, the probability distribution P of the event attribute. The present invention uses an anomaly scoring function to calculate the anomaly score of each event attribute value in the trajectory.
[0016] The purpose of the present invention is to solve the problem that most of the existing business process anomaly detection methods based on deep learning only use the event sequence information in the event log trajectory to train the next event prediction model, ignoring the structural information of the business process, and thus cannot fully characterize the business process behavior, thereby affecting the accuracy of business process anomaly detection. The present invention solves the technical problems existing in the above-mentioned prior art and brings about some creative technical effects after solving the problems. The specific description is as follows:
[0017] (1) The present invention defines the concept of event attribute graph and proposes a method of using multiple event attribute graphs to represent a trajectory, so as to more accurately characterize the business process behavior described by the trajectory.
[0018] (2) In order to capture the dependency between business process structure information and event attributes, the Weisfeiler-Lehman algorithm and the PV-DBOW model are introduced to construct a subtree pattern corpus and obtain the semantic feature vectors of event attribute graph nodes to generate event attribute value embeddings containing business process structure information.
[0019] (3) Using deep learning technology, we build a model to predict the attribute value of the next event based on the event attribute graph embedding and the sequence information of events in the trajectory.
[0020] (4) For the encoder network in the next event attribute value prediction model, the dot product attention mechanism is used to calculate the correlation vector between each object and the hidden state in the output sequence of the encoder network, so as to effectively capture the dependency relationship between events in long trajectories and enhance the prediction model's ability to predict the attribute probability distribution of the next event in long trajectories.
[0021] (5) Experimental verification shows that this method effectively improves the performance of business process anomaly detection on public datasets compared with existing business process anomaly detection methods, and can locate the position of anomalies in the trajectory.
[0022] The present invention proposes a business process anomaly detection method (EAGE) based on event attribute graph embedding, which makes full use of the structural information of the business process and the dependency between event attributes, and improves the performance of anomaly detection through deep learning technology. Experimental results show that the EAGE method can effectively improve the accuracy of business process anomaly detection on multiple public data sets. Through comparative experiments and ablation experiments, the present invention verifies the effectiveness of the dot product attention mechanism in the proposed method, as well as the advantages of combining the WL algorithm and the PV-DBOW model to extract features. The method proposed in the present invention is of great significance to improving the detection accuracy of business process anomalies and maintaining the smooth operation of process perception information systems. (IV) Description of the drawings
[0023] Figure 1It is the overall block diagram of the EAGE of the present invention.
[0024] Figure 2 This is the event attribute diagram of trajectory #142857 of the present invention.
[0025] Figure 3 It is a subtree pattern corpus diagram for constructing an event attribute graph of the present invention.
[0026] Figure 4 The semantic feature vector graph of the event attribute graph nodes generated by the present invention.
[0027] Figure 5 This is a prediction model diagram of the next event attribute value of the present invention.
[0028] Figure 6 This is a graph showing how the anomaly detection performance of the present invention changes with the number of WL relabeling iterations iter.
[0029] Figure 7 This is a comparison chart of the F1 scores of various business process anomaly detection methods of the present invention.
[0030] Figure 8 This is the ablation study diagram of the dot product attention module of the present invention. (V) Specific implementation methods
[0031] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below in combination with specific examples and with reference to the accompanying drawings.
[0032] In the complex business environment of modern enterprises and organizations, business process anomaly detection has become one of the important methods to ensure the smooth operation of business processes. With the increasing complexity of business processes, traditional anomaly detection methods have gradually exposed their limitations. For example, it is difficult to discover data flow anomalies in business processes, it is difficult to locate the position of anomalies in the trajectory, and it is impossible to effectively deal with various abnormal situations. To address this problem, the present invention proposes a business process anomaly detection method based on event attribute graph embedding (Business Process Anomaly Detection based on Event-Attribute Graph Embedding, EAGE). Fully extract the dependency relationship between the structural information of the business process and the event attribute value, generate the feature vector of the trajectory in the event log, construct the attribute value prediction model of the next event, and improve the accuracy of business process anomaly detection. The overall framework of the method of the present invention is as follows: Figure 1 shown.
[0033] Step 1: Extract business process structure information
[0034] (1) Generate event attribute graph The event attribute graph of the event log trajectory can describe the structural information of the business process from multiple event attribute perspectives. In order to fully explore the dependencies between the structural information of the business process and the event attributes, the present invention uses multiple event attribute graphs to represent a trajectory from multiple event attribute perspectives. For the continuous attributes of the trajectory, its attribute values can be discretized and converted into discrete attributes, and then the event attributes can be converted into event attribute graphs. Therefore, the present invention currently mainly studies the method of representing the discrete attributes of the event into an event attribute graph. The event attribute graph of a trajectory is an undirected graph. For each discrete attribute (including activity attributes) in the trajectory, an event attribute graph can be constructed. The event attribute graph defined in the present invention is as shown in Definition 1. Definition 1 (Event Attribute Graph) The event attribute graph of the event attribute att in the trajectory tr is a binary pair G att =(V att ,E att ).in, V att is a set of attribute values of event attributes att in trajectory tr, indicating that G att Node collection; Represents G att A set of edges, each edge corresponds to an attribute-value pair with a dependency relationship of attribute att. Assume that the set AS is the discrete attribute set of all events in the event log, and the trajectory tr = <e 1 ,…,e n >, the steps for constructing the event attribute graph are as follows. (1) Obtain the node set of the event attribute graph. Initialize the event attribute graph G using the attribute value set of the event attribute att∈AS of the trajectory tr. att The node set V att If event e j If there is no attribute att, it will be filled with 0. (2) Generate attribute-value pairs of event attributes and traverse each event e in the trajectory tr i , the event e i and e i+1 The attribute value of each attribute att Forms a pair of attribute value pairs. (3) Obtain the edge set of the event attribute graph for each attribute value pair In the event attribute graph G att Generate a line from point to edge. Table 1 is a fragment of an event log. Each row represents an event, including case ID, event ID, activity name, time, resources and other attributes. Case ID is used to identify a sequence of events, or trajectory, that occurs sequentially. For example, in the case trajectory with CaseID 142857 in Table 1, the attribute value pairs obtained from the activity attributes of the event are (A, B), (B, C), (C, B), (B, C), (C, D), (D, C), (C, B) and (B, E), and the attribute pairs corresponding to other attributes are obtained in a similar way. According to the above-mentioned event attribute graph construction method, the event attribute graphs of the activity attributes and resource attributes of the event in trajectory 142857 can be obtained respectively, as shown in the figure. Figure 2 shown. Table 1 A fragment of an event log
[0035] (2) Constructing a subtree pattern corpus of event attribute graphs According to Definition 1 and the method for constructing event attribute graphs, each track of the event log can be converted into multiple event attribute graphs to form a set of event attribute graphs. Using the WL algorithm to extract subtree patterns of different heights in the event attribute graph, a 1-hop, 2-hop, …, m-hop subtree pattern corpus can be constructed, such as Figure 3 As shown, where m is the number of iterations of the WL algorithm. Assume that the event attribute graph set G = {g 1 ,g 2 ,....,g n}, in the i-th iteration of the WL algorithm, each g j ∈G will generate a multi-set consisting of i-hop subtree patterns Where V j is the event attribute graph g j A set of nodes, 1≤i≤m. i Each i-hop subtree pattern in (g) is regarded as a word, and the subtree patterns are sorted according to the feature vector centrality to form a sequence s i (g) Thus, the i-hop subtree pattern corpus Cp of the event attribute graph set G is constructed i . The eigenvector centrality is defined by equation (1). x v is the eigenvector centrality of node v (v∈V j ), Formula 2 indicates that the eigenvector centrality of node v is the sum of the eigenvector centralities of its neighboring nodes u. A[v,u]=1 means that node u and node v are connected, otherwise they are not connected. The eigenvector centrality depends on the number of node neighbors and the importance of their neighbors. The larger the eigenvector centrality, the more important the node. The corpus constructed from the i-hop subtree patterns extracted from the event attribute graph set G is called Cp i =(Ω i ,R i ), where Ω i is the vocabulary of the corpus, i.e., the set of different i-hop subtree patterns that appear in G, R i is the sequence s i Therefore, corpora are constructed for subtree patterns of different heights, thereby generating a corpus set CS = {Cp 1 ,Cp 2 ,...,Cp m}. Where m represents the number of iterations of the WL algorithm.
[0036] (3) Generate semantic feature vectors of event attribute graph nodes The PV-DBOW model learns the feature vector of the subtree pattern from the subtree pattern corpus of the event attribute graph, so that the feature vector of the event attribute graph node generated finally contains both the structural information between nodes and the semantic information of the nodes. The subtree pattern in the event attribute graph can be regarded as a word, and the sequence composed of different subtree patterns is similar to a sentence composed of different words. i Training the PV-DBOW model above can get a matrix The matrix Mat i The nth row of the i The semantic embedding of the nth subtree pattern in , 1≤i≤m, 1≤n≤len(tr), len(tr) is the length of the trajectory tr, and dim is the feature dimension. The process of using the PV-DBOW model to generate the semantic feature vector of the event attribute graph node is as follows Figure 4 shown. For the corpus set CS = {Cp 1 ,Cp 2 ,...,C m}, training an independent PV-DBOW model, can generate a matrix set Mat = {Mat 1 ,Mat 2 ,...,Mat mSince the subtree patterns of different heights of the same node describe the characteristics of the root node and its neighboring nodes at different granularities, the present invention concatenates the semantic feature vectors of multiple subtree patterns with node v as the root node in the event attribute graph to form a multi-scale semantic feature vector of the root node v. Event attribute graph of trajectory g j The multi-scale semantic feature vector X of node v in ∈G can be extracted from the matrix set Mat, and the process is shown in formula (2). X=Concat(ft 1 ,ft 2 ,…,ft m ) (2) where ft 1 ,ft 2 ,...,ft m They are matrices Mat 1 ,Mat 2 ,...,Mat m The feature vector of the 1st, 2nd, ... mth order subtree pattern with node v as the root, Concat() represents the concatenation of ft 1 ,ft 2 ,...,ft m Therefore, the feature vector of the event attribute graph node with a dimension of m*dim is finally obtained. In the event attribute graph of the trajectory, the subtree pattern of a node refers to the local structural information of the node in the event attribute graph, that is, the contextual information of the event attribute value of the node. The PV-DBOW model can be used to learn the semantic feature vector of the event attribute graph node from the subtree pattern corpus. The semantic feature vector reflects the co-occurrence relationship of the subtree patterns of the event attribute graph nodes of different trajectories, which can be regarded as the semantic information of the node semantic feature vector. Therefore, the subtree pattern of each node in the event attribute graph can be obtained by the WL algorithm, thereby capturing the structural information of the business process and the dependency relationship between the event attribute values.
[0037] Step 2: Generate the feature vector of the trajectory
[0038] The event attribute values contained in the trajectory are sometimes repeated, and these repeated values correspond to the same node in the event attribute graph. However, it is worth noting that the same event attribute value in the trajectory has different meanings at different locations, which is crucial for analyzing and reconstructing business process behaviors. After extracting the structural information of the business process, the event attribute values in the trajectory are mapped to the feature vectors of the event attribute graph nodes, and the feature vector of the trajectory is constructed accordingly. For a given tr= <e 1 ,e 2 ,....,e n>, first construct the multi-event attribute graph of the trajectory, and then use the WL algorithm to extract subtree patterns of different heights. For each event attribute att in the trajectory tr, its corresponding subtree pattern set is expressed as Where m represents the maximum height of the subtree pattern, corresponding to the number of iterations of the WL algorithm. Each event e in the trajectory tr i The characteristic vector of the attribute att It can be obtained through the following two steps: First, initialize a mapping dictionary Map = (key, value), where key is the subtree pattern and value is the feature vector obtained by the PV-DBOW model. Then, for event e i , obtain the feature vector of the subtree pattern corresponding to its attribute att by searching the dictionary Map. Among them, 1≤i≤n, n=len(trace) represents the length of the trace tr. Assume fm:SP→R dim is a mapping function from a subtree pattern to its feature vector, where SP represents the set of all possible subtree patterns and dim is the dimension of the feature vector. Then the event e i The characteristic vector of the attribute att It can be obtained by formula (3): Among them, subtrptn(e i ) represents event e i The subtree pattern corresponding to the attribute att. The feature vector X of the event attribute att of the trajectory tr tr , can be obtained by concatenating the feature vectors of the attribute values of all events Assume maxtrlen is the length of the longest trace in the event log. If the length of the trace tr n < maxtrlen, use the value 0 in X tr The end is padded to length maxtrlen.
[0039] Step 3: Build a prediction model for the next event attribute value The present invention mainly uses the gated recurrent unit GRU, encoder-decoder structure and dot product attention mechanism in the recurrent neural network to build a prediction model for the attribute value of the next event, and uses the feature vector of the trajectory to train this model. The deep network structure of the prediction model is as follows Figure 5 As shown. Figure 5 In GRU 1 As an encoder, GRU 2 As a decoder. First, the track tr = <e 1 ,e 2 ,....,en >att i Feature vector from attribute perspective is input to the encoder GRU 1 Then, GRU 1 Capture the sequence information between events in the trajectory and output a set of feature tensors and the hidden state tensor Where n = len(trace), att i represents the i-th attribute in the attribute set AS, 1≤i≤|AS|, i∈Z. Figure 3 In GRU 2 GRU 1 Output feature tensor and the hidden state tensor Decoded into probability distribution. The higher the probability of occurrence of the event attribute value in the trajectory, the higher the probability that it is normal. The dot product attention mechanism is the key link between the encoder and the decoder. It is used to identify the degree of correlation between the attribute value of the event in the trajectory and the attribute value of the next event, and assign different attention weights to these attribute values to more effectively capture the long-distance dependency between events in the trajectory, thereby improving the accuracy of anomaly detection of business process instances in long trajectories. The following formulas (4) and (5) are used to calculate the query tensor and key tensors The dot product coefficient between The attention weights obtained by converting these dot product coefficients using the Softmax function in, and is a learnable parameter matrix, 1≤j≤n, j∈Z. It is GRU 1 The concatenation of the tensors output by the encoder at different attribute perspectives will and Converted to a d-dimensional tensor. In formula (5), the Softmax function is used to transform the event e in the trajectory j The attribute att i The dot product coefficient calculated at Normalize to obtain the corresponding attention weight Attention Weight Intuitively reflects the event attribute value encoding of the trajectory Importance,Attention Weight in Reconstructing Business Process Behavior The larger the value, the The more important the prediction of the current attribute value is. Then, in order to improve the ability of the next event attribute value prediction model to capture the dependencies between events in the long trajectory of the business process instance, the encoder GRU is used 1 Output attribute value tensor and its attention weight Weighted As shown in formula (6). In order to better reconstruct the event log trace, j The present invention uses the teacher forcing method widely used in the field of NLP. The teacher forcing method uses the real label to guide the training of the next event attribute value model, accelerates model convergence and improves prediction accuracy. For the activity attribute a of the event in the trajectory, the activity attribute value ev of the previous event is used. j-1,a Predict current events j The activity attribute value ev j,a The following formula (7) represents the probability distribution of the attribute value ev calculated using the word embedding algorithm (implemented by calling the torch.nn.embedding module in the present invention) j-1,a Encoding Formula (8) represents the use of attribute value encoding and attention encoding Calculate the next attribute value ev j,a The predicted output yes and The concatenation is input to the decoder GRU 2 Finally, the next event attribute value ev j,a The probability distribution of That is, event e in trajectory tr j The probability distribution of attribute a It can be calculated by formula (9), where Represents the parameter matrix of the fully connected layer. For the inactive attribute a′, since the current attribute value ev j,a' Depends on the name of the current activity j,a , also depends on the previous attribute value ev j-1,a' Therefore, the present invention uses the activity attribute value ev of the current event j,a , the attribute value ev of the attribute a′ of the previous event j-1,a' , to assist in predicting the probability distribution of the attribute It can be calculated by the following formulas (10), (11) and (12):
[0040] Step 4: Trajectory abnormality determination Trajectory anomaly determination requires anomaly scores of event attribute values in the trajectory. Calculating the anomaly score of the event attribute value of the business process requires the prediction result of the next event attribute value prediction model, that is, the probability distribution P of the event attribute. The present invention uses an anomaly scoring function to calculate the anomaly score of each event attribute value in the trajectory. For a certain event attribute value y in the trajectory tr, the anomaly scoring function f sc The definition of (P, y) is shown in formula (13). Among them, P y represents the probability assigned to the attribute value y by the prediction model. max(P) represents the maximum probability in the probability distribution P of the event attribute. max(P) and P y The difference between y and y can be used to evaluate the abnormality of the attribute value y. If the difference between the two is large, the prediction model believes that another attribute value should appear here instead of y. Therefore, y is an abnormal event attribute value. The anomaly scoring function assigns a higher anomaly score to uncommon events in the trajectory. Low-probability events in the trajectory may be caused by changes in the external environment or other special circumstances, not business process anomalies. Only when the anomaly score is greater than the anomaly threshold will it be considered an anomaly. Based on the anomaly score, the present invention uses an anomaly determination function f dt (s,τ) determines whether the trajectory tr is abnormal, as shown in formula (14). Among them, s is the abnormal score of all attribute values in the trajectory tr, and τ is the abnormal score threshold determined by the LP-Meanheuristic method. If a trajectory tr contains at least one anomaly score greater than the anomaly threshold, the trajectory is considered anomaly. The present invention uses an abnormal location function f al (s,τ) finds the location of the abnormal event attribute value in the trajectory tr, as shown in formula (15). The abnormal score s and the abnormal score threshold τ of all event attribute values in the trajectory tr are input into the abnormal location function, which can return the set FS of the index i of the abnormal attribute value. FS=f al (s,τ)={0≤i <sum|s[i]>τ} (15) Among them, sum is the number of anomaly scores, and the index i can be used to locate the abnormal attribute value of the event in the trajectory.
[0041] The embodiments of the present invention have achieved some positive effects during the development or use process, and indeed have great advantages over the prior art. The following content is described in conjunction with data, charts, etc. of the test process.
[0042] (1) Experimental setup
[0043] The present invention uses the dataset of Lahann et al., which contains synthetic event logs and real event logs. The synthetic event logs are Huge and Large event log datasets created by event log generation software PLG2 based on two specified business process models. The real event logs are BPIC 2012, BPIC 2013, BPIC 2015 and BPIC2017 datasets.
[0044] The specific information of the dataset is shown in Table 2. The Logs column indicates the number of event logs in the dataset. Table 2 Event log dataset
[0045] (2) Evaluation indicators
[0046] Due to the imbalance in the number of normal and abnormal trajectories in the event log dataset, the performance of the F1 score evaluation method widely used in business process anomaly detection research is improved. The F1 score is the harmonic mean of the precision and recall. The F1 score calculation method is as shown in formula (18), where TP: represents the number of samples correctly predicted by the model as positive. FP: represents the number of samples that the model incorrectly predicts as positive. FN: represents the number of samples that the model incorrectly predicts as negative.
[0047] The larger the F1 score value, the better the overall performance of the anomaly detection experiment of a certain method.
[0048] (3) Experimental results and analysis
[0049] Determine the number of iterations of the WL algorithm
[0050] This section evaluates the impact of the number of iterations of the WL algorithm on the performance of business process anomaly detection. This experiment examines the impact of the number of iterations iter on the performance of business process anomaly detection on event logs Huge_1, Large_1, BPIC2012_1, BPIC2013_1, and BPIC2015_1. Huge_1, Large_1, BPIC2012_1, BPIC2013_1, and BPIC2015_1 represent the first event logs in Huge, Large, BPIC2012, BPIC2013, and BPIC2015, respectively.
[0051] During the experiment, iter selected {1,2,3,4,5}, set the model training round number of epochs of the EAGE method of the present invention to 25, batch size batch_size = 64, learning rate learning_rate = 0.0006, and the learning rate decay factor was set to 0.85. During training, 80% of the trajectories in the event log were divided into training sets and 20% were divided into test sets. Using the F1 score as the evaluation index, 10 experiments were carried out on the EAGE method of the present invention on each of the above event logs, and the average F1 score of the 10 experiments was taken as the final F1 score of the event log. Figure 6 It shows how the F1 score of EAGE changes with the number of iterations iter of the WL algorithm.
[0052] from Figure 6 It can be found that on the datasets BPIC2012_1, Huge_1 and Large_1, as iter increases from 1 to 3, the performance of anomaly detection improves, and when iter increases from 3 to 5, the change is not large. This is because in the initial iterations, the node features of the event attribute graph are gradually enriched, and the structural information of the event attribute graph is gradually complete. However, when a certain number of iterations is reached, the node label information tends to be saturated, and further iterations will not significantly increase new useful information, and the performance tends to be stable. On the datasets BPIC2013_1 and BPIC2015_1, as iter increases from 1 to 5, the performance of anomaly detection first improves and then decreases. The main reason for this is that the event attribute graphs corresponding to these two datasets have more node labels. As iter increases, the types of generated node labels increase, and the number of common labels between event attribute graphs decreases, which will reduce the semantic relevance between the label sequences of event attribute graphs in the corpus. In the following experiments, the number of iterations of the WL algorithm on the dataset BPIC2015 is set to 4, and the number of iterations of the WL algorithm on other datasets is set to 3.
[0053] Effectiveness Evaluation of Anomaly Detection Methods in Business Processes
[0054] In order to verify whether the method of the present invention can effectively detect business process anomalies, this experiment is carried out on two synthetic event log data sets Huge and Large, and four real event log data sets BPIC 2012, BPIC 2013, BPIC 2015 and BPIC2017. The method EAGE of the present invention is compared with methods such as OC-SVM, Naive, Naive+, Sampling, t-STIDE+, DAE, Likelihood+, BINet, DAPNN and GAMA. Event logs are divided into groups, as shown in Table 2. For example, BPIC2013 has 3 event logs, where BPIC2013_3 represents the third event log in BPIC2013. Experiments are carried out on each event log data set in Table 2, and the experimental results of each data set use the average F1 score of the event logs in the data set. For example, the F1 score of the data set BPIC2013 is the average of BPIC2013_1, BPIC2013_2 and BPIC2013_3. During the experiment, the number of epochs, batch size, learning rate, learning rate decay factor, and the division of the training set and test set of the EAGE method are consistent with the experimental settings in the experiment of determining the number of iterations of the WL algorithm. The EAGE method is used to conduct 10 experiments on each event log in the above event log dataset, and the average F1 score of the 10 experiments is taken as the final F1 score of the event log. The experimental results are shown in Figure 7 shown.
[0055] from Figure 7 It can be seen that the method of the present invention achieves the best results on all event logs. and OC_SVM cannot detect anomalies in event attribute values, and method OC-SVM is not specifically designed for business process anomaly detection, so the performance of anomaly detection on data sets containing control flow anomalies and data flow anomalies is relatively low. Methods DAE, BINet, and DAPNN are all reconstruction-based methods, but ignore the structural information of the business process model. GAMA takes into account the structural information of the business process model, but the encoder designed by this method does not consider the order of the nodes, so the anomaly detection performance is lower than that of the method proposed in the present invention. From the above experimental results, it can be seen that the business process anomaly detection method based on event attribute graph embedding proposed in the present invention can more effectively detect anomalies in business processes. Therefore, in the anomaly detection of business processes, it is necessary to consider the structural information of the business process.
[0056] from Figure 7It can also be found that there is a clear performance gap between synthetic event logs (Huge, Large) and real event logs (BPIC 2012, BPIC 2013, BPIC 2015, and BPIC 2017). There may be two explanations for this gap: first, the algorithm only requires the identification of human anomalies, and it is unclear whether the unknown anomalies exist in the original event logs. Second, the synthetic event logs may cover business processes with simple characteristics, while the business processes in BPIC are more complex and therefore more difficult to understand and process. These factors together lead to the performance difference between synthetic event logs and BPIC event logs.
[0057] Effects of different encoding methods on anomaly detection performance
[0058] In order to test whether the method of combining the WL node relabeling algorithm and the PV-DBOW model to obtain the business process structure information can improve the performance of business process anomaly detection, an experiment was conducted to replace the graph node embedding module of EAGE with one-hot encoding. In this experiment, the One-HotNet anomaly detection method was obtained by replacing the graph node embedding module of EAGE with one-hot encoding. Experiments were carried out on four real event log datasets (BPIC 2012, BPIC 2013, BPIC 2015 and BPIC 2017) and two synthetic event log datasets (Huge and Large). Experiments were carried out on each event log in the above event log datasets using EAGE and One-HotNet methods. The division of training and test sets, experimental parameters, and calculation methods of evaluation indicators were consistent with the experimental settings in the effectiveness evaluation experiment of business process anomaly detection methods. Table 3 Comparison of graph node embedding encoding and one-hot encoding (F1 score)
[0059] It can be found from Table 3 that the EAGE method achieves higher F1 scores for both real event logs and synthetic event logs. In the BPIC13 and BPIC15 datasets, the relatively small number of trajectories leads to a small performance improvement on these two groups of datasets. This may be because the number of trajectories in the dataset is not enough to fully capture and represent the complexity and variability of the business process. Due to the small amount of data, the algorithm may not be able to fully learn the characteristics and patterns of the data, thereby limiting the performance improvement. In addition, due to the limited amount of data, there may be large randomness and noise, which will also affect the performance of the algorithm. Therefore, in the case of a small amount of data, the algorithm may not be able to give full play to its advantages, resulting in a small performance improvement. It can be concluded that the EAGE method of the present invention significantly improves the performance of business process anomaly detection compared to the One-HotNet method, which shows that the structural information of the business process is crucial for anomaly detection, especially in the datasets with long trajectories (BPIC12 and BPIC17), the improvement of the F1 score is particularly obvious.
[0060] Time efficiency analysis of EAGE method
[0061] In order to verify the time efficiency of the EAGE method proposed in the present invention, this experiment was conducted on two synthetic event log data sets Huge and Large, and four real event log data sets BPIC 2012, BPIC 2013, BPIC 2015 and BPIC 2017. The core of the experiment is to compare the time required for the method of the present invention to establish the next event attribute value prediction model with the current open source and best GAMA method, that is, the time for training a round of prediction models. The average training time of the EAGE method and the GAMA method on each data set is used as the final training time of the two methods on the data set. The experimental results are shown in Table 4, where the unit of the data is seconds. Table 4 Time efficiency comparison between EAGE and GAMA
[0062] The results in Table 4 show that the time used by the EAGE method to build a prediction model exceeds that of the GAMA method, with an average increase of about 1.08 times. This difference is mainly attributed to the fundamental difference in the design concepts of the two methods. The encoder of the GAMA method focuses on extracting structural information between event attributes and ignores the sequence information of events in the trajectory. Although this design simplifies the model, it may sacrifice the capture of deep-level features of the event sequence. In contrast, the EAGE method presents a more comprehensive and in-depth perspective in model construction. It not only mines the structural information of event attributes in the trajectory, but also considers the sequence characteristics of events in the trajectory. Although the EAGE method increases the architecture and computational complexity of the prediction model and prolongs the training time. Combined with the effectiveness evaluation experimental results of the business process anomaly detection method, although the training time of the EAGE method's prediction model has increased, it has achieved the best anomaly detection performance on all test data sets, described by the F1 score. It can be concluded that compared with the GMA method, the extra time investment of the EAGE method is worthwhile, and it significantly improves the accuracy and reliability of business process anomaly detection.
[0063] Ablation experiment
[0064] In the method EAGE of the present invention, the dot product attention mechanism weightedly integrates the information of neighboring nodes by calculating the similarity between nodes, thereby improving the performance of business process anomaly detection. However, in order to verify the role of dot product attention in the model, this experiment designed an ablation experiment to evaluate its specific impact on model performance by removing the dot product attention mechanism. Experiments were conducted on two synthetic event log datasets (Huge and Large) and four real event log datasets (BPIC 2012, BPIC2013, BPIC 2015, and BPIC 2017). Figure 8 As shown in the figure, EAGE and EAGE w / o DPA are the methods obtained after removing the dot product attention module. Experiments were conducted on each event log in the above dataset using the EAGE and EAGE w / o DPA methods. The division of the training set and the test set, the experimental parameters, and the calculation method of the evaluation index are consistent with the experimental settings in the effectiveness evaluation experiment of the business process anomaly detection method.
[0065] The performance of EAGE w / o DPA model has a small drop in Huge and Large datasets, but a large drop in BPIC2012, BPIC2013, BPIC2015, and BPIC2017 datasets. This is because the trajectory length of the event logs in these two datasets is shorter than that of the first four datasets. Dot product attention can effectively capture long-distance dependencies by calculating the degree of correlation between the attribute value of an event in the trajectory and the attribute value of the next event, and assigning different attention weights to these attribute values, and performs better in long sequence tasks. Through the above experiments, it can be concluded that the dot product attention mechanism is more effective in capturing long-distance dependencies in long trajectories, and the performance of EAGE w / oDPA drops significantly on datasets with longer trajectory lengths.
[0066] Experimental results show that the EAGE method can effectively improve the accuracy of business process anomaly detection on multiple public data sets. Through comparative experiments and ablation experiments, the present invention verifies the effectiveness of the dot product attention mechanism in the proposed method, as well as the advantages of extracting features by combining the WL algorithm and the PV-DBOW model. The method proposed in the present invention is of great significance to improving the detection accuracy of business process anomalies and maintaining the smooth operation of process perception information systems.
[0067] The above description is only a specific implementation mode of the present invention, but the protection scope of the present invention is not limited thereto. Any modifications, equivalent substitutions and improvements made by any technician familiar with the technical field within the technical scope disclosed by the present invention and within the spirit and principle of the present invention should be covered by the protection scope of the present invention.
Claims
1. The present invention proposes a business process anomaly detection method based on event attribute graph, characterized in that: include: Step 1: First, according to the dependency relationship between the attribute values of the same-name attributes of events in the trajectory, the trajectory is represented as multiple event attribute graphs; then, the Weisfeiler-Lehman algorithm is used to establish a subtree pattern corpus of the event attribute graph; finally, the PV-DBOW model is used to learn the semantic feature vectors of the event attribute graph nodes from the subtree pattern corpus; Step 2: Based on step 1, the event attribute values in the trajectory are mapped into feature vectors of event attribute graph nodes, and the feature vector of the trajectory is constructed accordingly; Step 3: Use the gated recurrent unit GRU, encoder-decoder structure and dot product attention mechanism in the recurrent neural network to build a prediction model for the attribute value of the next event, and use the feature vector of the trajectory to train this model; Step 4: Use the prediction result of the next event attribute value prediction model, that is, the probability distribution P of the event attribute to calculate the abnormal score of the event attribute value of the business process.
2. A method for detecting anomalies in a business process based on an event attribute graph according to claim 1, characterized in that: An event attribute graph defining the event log trace; The event attribute graph of the event attribute att in the trajectory tr is a two-tuple V att is a set of attribute values of event attributes att in trajectory tr, indicating that G att Node collection; Represents G att A set of edges, each edge corresponds to an attribute-value pair with a dependency relationship of attribute att; Assume that the set AS is the discrete attribute set of all events in the event log, and the trajectory tr = <e1,…,e n >, the steps for constructing the event attribute graph are as follows: (1) Obtain the node set of the event attribute graph and initialize the event attribute graph G using the attribute value set of the event attribute att∈AS of the trajectory tr att The node set V att , if event e j If there is no attribute att, fill it with 0; (2) Generate attribute-value pairs of event attributes and traverse each event e in the trajectory tr i , the event e i and e i+1 The attribute value of each attribute att Form a pair of attribute value pairs; (3) Obtain the edge set of the event attribute graph for each attribute value pair In the event attribute graph G att Generate a line from point to edge.
3. The method for detecting anomalies in a business process based on an event attribute graph according to claim 1, characterized in that: The prediction model of the next event attribute value; The present invention mainly uses the gated recurrent unit GRU, encoder-decoder structure and dot product attention mechanism in the recurrent neural network to construct a prediction model for the attribute value of the next event, and uses the feature vector of the trajectory to train the model; the deep network structure of the prediction model is shown in FIG5 , in which GRU1 is used as an encoder and GRU2 is used as a decoder; first, the trajectory tr= <e1,e2,....,e n >att i Feature vector from attribute perspective is input into the encoder GRU1; then, GRU1 captures the sequence information between events in the trajectory and outputs a set of feature tensors and the hidden state tensor Where n = len(trace), att i represents the i-th attribute in the attribute set AS, 1≤i≤|AS|, i∈Z; In Figure 3, GRU2 outputs the feature tensor of GRU1 and the hidden state tensor Decoded into probability distribution, the higher the probability of occurrence of event attribute values in the trajectory, the higher the probability of normality; the dot product attention mechanism is the key link between the encoder and the decoder. It is used to identify the correlation between the attribute value of an event in the trajectory and the attribute value of the next event, and assign different attention weights to these attribute values to more effectively capture the long-distance dependency between events in the trajectory, thereby improving the anomaly detection accuracy of business process instances in long trajectories. The following formulas (4) and (5) are used to calculate the query tensor respectively: and key tensors The dot product coefficient between The attention weights obtained by converting these dot product coefficients using the Softmax function in, and is a learnable parameter matrix, 1≤j≤n, j∈Z, is the concatenation of the tensors output by the GRU1 encoder at different attribute perspectives, which and Converted into a d-dimensional tensor; In formula (5), the Softmax function is used to calculate the event e in the trajectory j The attribute att i The dot product coefficient calculated at Normalize to obtain the corresponding attention weight Attention Weight Intuitively reflects the event attribute value encoding of the trajectory Importance,Attention Weight in Reconstructing Business Process Behavior The larger the value, the The more important the prediction of the current attribute value is; Then, in order to improve the ability of the next event attribute value prediction model to capture the dependencies between events in the long trajectory of the business process instance, the attribute value tensor output by encoder GRU1 is used. and its attention weight Weighted As shown in formula (6); In order to better reconstruct the event log trace, j The present invention uses the teacher forcing method widely used in the field of NLP. The teacher forcing method uses the real label to guide the training of the next event attribute value model, accelerates model convergence and improves prediction accuracy; for the activity attribute a of the event in the trajectory, the activity attribute value ev of the previous event is used. j-1,a Predict current events j The activity attribute value ev j,a The following formula (7) represents the use of the word embedding algorithm (implemented by calling the torch.nn.embedding module in the present invention) to calculate the attribute value ev j-1,a Encoding Formula (8) represents the use of attribute value encoding and attention encoding Calculate the next attribute value ev j,a The predicted output yes and The concatenation is input into the decoder GRU2; finally, the next event attribute value ev j,a The probability distribution of That is, event e in trajectory tr j The probability distribution of attribute a It can be calculated by formula (9), where Represents the parameter matrix of the fully connected layer For the inactive attribute a′, since the current attribute value ev j,a′ Depends on the name of the current activity j,a , also depends on the previous attribute value ev j-1,a′ Therefore, the present invention uses the activity attribute value ev of the current event j,a , the attribute value ev of the attribute a′ of the previous event j-1,a' , to assist in predicting the probability distribution of the attribute It can be calculated by the following formulas (10), (11) and (12):
4. The method for detecting anomalies in a business process based on an event attribute graph according to claim 1, characterized in that: The method for determining trajectory abnormality; Trajectory anomaly determination requires anomaly scores of event attribute values in the trajectory; calculating the anomaly score of event attribute values in the business process requires the prediction result of the next event attribute value prediction model, that is, the probability distribution P of the event attribute; the present invention uses an anomaly scoring function to calculate the anomaly score of each event attribute value in the trajectory. For a certain event attribute value y in the trajectory tr, the anomaly scoring function f sc The definition of (P, y) is shown in formula (13): Among them, P y represents the probability assigned to the attribute value y by the prediction model, max(P) represents the maximum probability in the probability distribution P of the event attribute, max(P) and P y The difference between y and y can be used to evaluate the abnormality of the attribute value y. If the difference between the two is large, the prediction model believes that another attribute value should appear here, so y is an abnormal event attribute value. The anomaly scoring function assigns a higher anomaly score to uncommon events in the trajectory. Low-probability events in the trajectory may be caused by changes in the external environment or other special circumstances, not business process anomalies. Only when the anomaly score is greater than the anomaly threshold will it be considered an anomaly. Based on the anomaly score, the present invention uses an anomaly determination function f dt (s,τ) determines whether the trajectory tr is abnormal, as shown in formula (14), where s is the abnormality score of all attribute values in the trajectory tr, and τ is the abnormality score threshold determined by the LP-Meanheuristic method; If the trajectory tr contains at least one anomaly score greater than the anomaly threshold, the trajectory is judged to be abnormal; The present invention uses an abnormal location function f al (s, τ) finds the location of the abnormal event attribute value in the trajectory tr, as shown in formula (15); the abnormal score s and the abnormal score threshold τ of all event attribute values in the trajectory tr are input into the abnormal location function, which can return the set FS of the index i of the abnormal attribute value; FS=f al (s,τ)={0≤i <sum|s[i]>τ} (15) Among them, sum is the number of anomaly scores, and the index i can be used to locate the abnormal attribute value of the event in the trajectory.
5. A method for detecting anomalies in a business process based on event attribute graph embedding according to any one of claims 1 to 4, characterized in that: The business process anomaly detection method based on event attribute graph embedding includes four modules: First, in the module of extracting business process structure information, each trajectory in the event log is represented as multiple event attribute graphs, and the semantic feature vectors of the event attribute graph nodes are learned using the WL algorithm and the PV-DBOW model; then, in the module of generating trajectory feature vectors, the feature vector of the trajectory is constructed using the semantic feature vectors of the event attribute graph nodes according to the mapping relationship between the subtree pattern of the event attribute values in the trajectory and the event attribute graph nodes; then, in the module of constructing the prediction model for the next event attribute value, a prediction model for the next event attribute value is constructed, and the prediction model is trained using the feature vector of the trajectory; finally, in the module of determining the anomaly score and trajectory anomaly, the next event attribute value prediction model is used to obtain the probability distribution of the next event attribute of the business process instance, and the anomaly score of the next event attribute value is calculated; if the anomaly score exceeds the set anomaly score threshold, the event attribute value is determined to be abnormal, otherwise, the event attribute value is determined to be normal.
Citation Information
Cited By
Fire fighting system-oriented event rule mining method and device
CN120258123A