Business process log case attribute automatic identification method and system
By applying supervised machine learning methods in business process event logs, automatically identifying case attributes, the problems of inefficiency of manual marking and difficulty in ensuring accuracy are solved, and efficient and accurate process analysis and optimization are achieved.
Patent Information
- Application Number
- CN202510074548.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, manual marking of business process event logs is inefficient, prone to errors, and difficult to quickly adapt to business changes, resulting in inconsistent process analysis results and difficult to ensure the accuracy of process analysis.
A method for automatic identification of case attributes of business process logs is proposed, and a supervised machine learning method, especially the gradient-enhancing decision tree algorithm is used to build a binary classifier to automatically identify case attributes in the event log, and to achieve accurate attribute recognition through feature extraction and candidate attribute combination evaluation.
This method can greatly reduce human errors, improve data accuracy and reliability, improve process mining efficiency, support rapid data processing and analysis, help enterprises respond to market changes in a timely manner, and improve business flexibility and competitiveness.
Smart Images

Figure CN120045872A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of process mining, and particularly relates to a method and system for automatically identifying case attributes of business process logs. Background Art
[0002] Business process event logs are the basis for analyzing and optimizing business processes in process mining. They contain detailed records of an organization or business process, usually including information such as case numbers, activities, resources, and timestamps. These information are crucial for understanding the execution of the process, identifying bottlenecks, and optimizing resource allocation. In order to mine valuable process models from these event logs, accurate tag identification is crucial. Traditionally, people usually rely on manual marking to identify and classify the fields in these event logs. However, with the increase in the complexity of business processes, the manual marking method gradually exposes a series of deficiencies.
[0003] First of all, the manual marking method is inefficient and error-prone. As the number of enterprise activities increases, the scale of the generated event logs often grows exponentially. Manually marking each column is not only time-consuming and laborious, but also in the face of a large amount of data, the accuracy of manual marking is difficult to guarantee. The interference of human factors may lead to inconsistent marking, which in turn affects the subsequent process analysis results. In contrast, automated algorithms can process a large amount of data in a short time, ensuring the accuracy and consistency of marking, thereby improving the efficiency of process mining. Secondly, the business processes of enterprises often evolve continuously with changes in market demand and internal management. The flexibility of manual marking is poor and it is difficult to quickly adapt to new business scenarios. By designing an automatic recognition algorithm, the marking rules can be flexibly adjusted according to different business requirements and event characteristics, and quickly adapt to changes. This scalability enables enterprises to maintain a keen insight into processes in a dynamic environment. In addition, automated marking can also reduce the influence of human biases and subjective judgments, providing a richer basis for subsequent data analysis and mining. Therefore, it is particularly important to design an algorithm to automatically identify the case attributes in event logs. This algorithm can provide strong support for the process mining of enterprises, helping enterprises to achieve continuous optimization and innovation in a highly competitive market environment. Summary of the Invention
[0004] In order to solve the above problems, the present invention proposes a method and system for automatically identifying case attributes of business process logs. For business process event logs, it can automatically identify the case attributes of accurate event logs, reduce the incidence of human error marking and the dependence on users' professional knowledge, improve the accuracy and reliability of data, and thus achieve process optimization and efficiency improvement.
[0005] The technical solution of the present invention is as follows:
[0006] A method for automatically identifying case attributes of business process logs includes the following steps:
[0007] Step 1: Obtain business process event logs, extract features from the event logs, and analyze the features of each column.
[0008] Step 2: Use supervised machine learning methods to construct binary classifiers for each key attribute, output the probabilities that each column belongs to the four key attributes respectively, and select the top N attributes with the highest probability for each column to form the candidate attribute set for each column.
[0009] Step 3: According to the number of attributes included in the candidate attribute set of each column, divide the candidate attribute set of each column into single candidate attributes or candidate attribute combinations.
[0010] Step 4: For the candidate attribute combinations, use model discovery algorithms to discover process models, evaluate the performance of each candidate attribute combination through cross-validation, and calculate the scores of each candidate attribute combination for each column.
[0011] Step 5: According to the scores of each candidate attribute combination, select the candidate attribute combination with the highest score as the final key attribute output.
[0012] Further, the specific process in Step 1 is as follows:
[0013] Step 1.1: The obtained business process event logs consist of several columns, and each column contains several values composed of letters, numbers, and special characters.
[0014] Step 1.2: Extract five features from the values of each column in the event logs, including letter ratio, number ratio, special character ratio, unique value ratio, and average string length; let the column set A of the event logs = {A 1 , A 2 ,..., A k}, where the kth column A k contains n string values: A k = {s 1 , s 2 ,..., s n}, each string is composed of multiple characters, and the ith string is denoted as: s i = {c 1 , c 2 ,..., c m}, m is the total number of characters; among them, each character is a letter, a number, or a special character.
[0015] The letter ratio is a measure of the proportion of English letters in the string, and the calculation process is as follows:
[0016] First, define the indicator function of the letter ratio:
[0017]
[0018] Among them, c j is the j-th character; Π letter (·) is the indicating function of the letter ratio; the 26 letters include both uppercase and lowercase;
[0019] Then, calculate the letter ratio:
[0020]
[0021] Among them, Letter ratio; |S i | is the number of characters in s i ;
[0022] The digit ratio is a measure of the proportion of digits in the string, and the calculation process is as follows:
[0023] First, define the indicating function of the digit ratio:
[0024]
[0025] Among them, Π digit (·) is the indicating function of the digit ratio;
[0026] Then, calculate the digit ratio:
[0027]
[0028] Among them, D(·) is the digit ratio;
[0029] The special character ratio is a measure of the proportion of non-letter and non-digit characters in the string, including,.!:@-#%, and the calculation process is as follows:
[0030] First, define the indicating function of the special character ratio:
[0031]
[0032] Among them, ∏ special (·) is the indicating function of the special character ratio;
[0033] Then, calculate the special character ratio:
[0034]
[0035] Among them, S(·) is the special character ratio;
[0036] The unique value ratio is the proportion of the number of different string values to the total number, and the calculation formula is:
[0037]
[0038] Among them, U(·) is the unique value ratio; unique(A k ) represents the number of different values in column A k If U(A k ) is close to 1, it indicates that the values in this column are not repeated with other columns;
[0039] The calculation process of the average string length is as follows:
[0040]
[0041] Among them, M(·) is the average string length;
[0042] Step 1.3. Analyze the five features of each column and preliminarily determine the candidate key attribute types of the column; specifically:
[0043] The key attributes include case number, timestamp, activity, and resource;
[0044] If the current column contains both letters and numbers and has a high unique value ratio, that is, U(A k ) ≥ 0.8, then it is determined that the candidate key attribute type of the current column is the case number;
[0045] If the digital ratio and special character ratio of the current column are the highest compared to other columns and conform to the time format, then it is determined that the candidate key attribute type of the current column is the timestamp;
[0046] If the current column has a high letter ratio, that is the average string length is longer than that of columns with a high letter ratio, then it is determined that the candidate key attribute type of the current column is the activity;
[0047] If the current column has a high letter ratio and a high unique value ratio, that is then it is determined that the candidate key attribute type of the current column is the resource.
[0048] Further, in step 2, the key attributes include case number, activity, timestamp, and resource; the case number is an identifier used to uniquely identify a specific process instance, and each case represents a complete business process, including all activities from start to end; the activity refers to the specific operations or tasks performed in the business process, and each activity represents a step, usually associated with business operations, decisions, or events; the timestamp is an identifier that records the time when the activity occurs, usually represented in the form of date and time, and the timestamp is used to analyze the execution order and duration of activities; the resource refers to the personnel, systems, or devices that participate in or are responsible for the activity during its execution, and resources include human resources, technical resources, or other support resources; attributes other than the above four key attributes are non-key attributes, and are directly marked as others after the key attributes are determined;
[0049] The supervised machine learning method specifically uses the gradient boosting decision tree algorithm model. During the training process of the gradient boosting decision tree algorithm model, the columns in the known event log are labeled as corresponding tag categories, and the gradient boosting decision tree algorithm model is used to classify the column features. Four binary classifiers are established to respectively identify the four key attributes of case number, activity, timestamp, and resource. The five features of the letter ratio, digit ratio, special character ratio, unique value ratio, and average string length of each column in the event log are input into the gradient boosting decision tree algorithm model, and each column is labeled as a target category or a non-target category;
[0050] For each binary classifier, multiple gradient boosting decision tree algorithm models are trained respectively. The loss function is optimized using the gradient boosting method. The model builds a new decision tree in each iteration. The new decision tree learns the samples that the previous all decision trees failed to predict correctly to reduce the error. First, the model is initialized as a constant term, representing the initial prediction value; for the t-th iteration, calculate the residual:
[0051]
[0052] where, is the residual of the -th sample in the t-th iteration; is the true label of the -th sample; is the predicted value of the -th sample in the (t - 1)-th iteration;
[0053] Train a new decision tree with the residual as the target, and update the predicted value of the gradient boosting decision tree algorithm model:
[0054]
[0055] where, is the The predicted value of the t-th round of iteration for a sample; v is the learning rate; h t (·) is the new decision tree obtained in the t-th round of iteration; For the sample;
[0056] The final gradient boosting decision tree algorithm model will output the probabilities that each column belongs to case number, activity, timestamp, or resource. The probability calculation formula is:
[0057]
[0058] where P(·) is the probability; is the predicted value of the final gradient boosting decision tree algorithm model for the sample ;
[0059] According to the output probabilities, select the top N attributes to form the candidate attribute set for each column. The specific value range of N is 1 - 5, and the specific value is adjusted according to the actual event log characteristics.
[0060] Furthermore, in step 3, if the number N of attributes included in the candidate attribute set of the current column is 1, then determine that the candidate attribute set of the current column is a single candidate attribute. In this case, this 1 attribute in the candidate attribute set is the key attribute, and this key attribute is the attribute name of the current column;
[0061] If the number N of attributes included in the candidate attribute set of the current column is greater than 1, then determine that the candidate attribute set of the current column is a candidate attribute combination. In this case, it is necessary to continue to make judgments and evaluations to determine the best attribute combination.
[0062] Furthermore, in step 4, the model discovery algorithm adopted is specifically the inductive mining algorithm. This algorithm adopts the divide-and-conquer idea, decomposes the problem of discovering the process model of an event log into the problem of discovering the sub-processes of multiple sub-logs obtained by splitting the event log. First, initialize the process model as empty; recursively divide the event log. Each division is based on the discrimination rules in the event log, including activity name and resource type; apply the inductive mining algorithm to each sub-log to discover the sub-process model, and gradually construct an accurate and concise process model; finally, merge the sub-process models into a complete process model to form a process description of the entire event log;
[0063] Evaluate the quality of the model mined by the inductive mining algorithm for each candidate attribute combination by combining cross - validation with the F - measure value, a model evaluation metric in the field of process mining. Calculate the score of the F - measure value of the model. First, randomly divide the data set into two subsets, and each subset maintains the data distribution consistency. For each candidate attribute combination, perform cross - validation. In each iteration, select one subset as the test set and the other subset as the training set. Use the training set to train the model and use the test set to evaluate the performance of the process model, and record the F - measure value. Swap the training set and the test set for re - evaluation.
[0064] The F - measure value is the harmonic mean of the fitness and the precision. The fitness quantifies the ability of the process model to regenerate the recorded traces in the event log, and the precision quantifies the ability of the process model to generate only the recorded traces in the event log. The formula for calculating the F - measure value is:
[0065]
[0066] where F - measure(·) is the F - measure value; fitness(·) is the fitness; precision(·) is the precision; L is the event log; M is the process model mined from the event log.
[0067] Finally, take the average F - measure value as the score of the final required candidate attribute combination.
[0068] A business process log case attribute automatic recognition system adopts the above - mentioned business process log case attribute automatic recognition method. The input of this system is the business process event log, and the output is the event log with recognized case attributes. This system includes the following modules:
[0069] An event log acquisition and feature extraction module, which is used to acquire the business process event log, extract features from the event log, and analyze the features of each column.
[0070] A module for narrowing the range of candidate key attributes, which is used to build a binary classifier for each key attribute by using supervised machine learning methods, output the probabilities that each column belongs to the four key attributes respectively, and select the N attributes with the highest probability for each column to form the candidate attribute set for each column.
[0071] A candidate attribute set discrimination module, which is used to judge whether the candidate attribute set is a single candidate attribute or a candidate attribute combination.
[0072] The process model discovery and evaluation module uses the model discovery algorithm to discover process models for candidate attribute combinations, evaluates the performance of each candidate attribute combination through cross-validation, and calculates the scores of each candidate attribute combination for each column.
[0073] The best candidate combination selection module is used to select the candidate attribute combination with the highest score as the final key attribute output.
[0074] The beneficial technical effects brought by the present invention are as follows.
[0075] For business process event logs, the present invention proposes a method that can automatically identify the case attributes of business process event logs, which can greatly shorten the time for data preparation, enabling users to enter the data analysis and process discovery stages faster, thereby improving work efficiency. At the same time, aiming at the problem of easy errors in the manual annotation process, the automatic identification of case attributes in the method of the present invention can reduce the incidence of these human errors and improve the accuracy and reliability of data. In addition, in a rapidly changing business environment, real-time analysis and decision-making are crucial. The automatic identification of case attributes in the method of the present invention can support rapid data processing and analysis, helping enterprises respond to market changes in a timely manner and enhancing the flexibility and competitiveness of their businesses. The ability of the method of the present invention to automatically identify case attributes enables process mining technology to be applied to more fields, such as healthcare, finance, manufacturing, etc., promoting cross-industry digital transformation.
[0076] The present invention can automatically extract multi-dimensional features from event logs, covering key information such as letter ratio, digit ratio, special character ratio, unique value ratio, and average string length. This process requires no manual intervention, greatly improving the efficiency and accuracy of feature extraction. Traditional manual marking methods require professionals to check log fields one by one, which is not only time-consuming and laborious but also prone to errors due to subjective judgment. The automated feature extraction of the present invention ensures the consistency and objectivity of feature acquisition.
[0077] By constructing a supervised machine learning model based on the gradient boosting decision tree algorithm, the present invention realizes the accurate classification of log attributes. During the model training process, using a small amount of labeled sample data, it automatically learns the rules and patterns of different attribute features without the need for manual setting of complex rules. Compared with manual marking that relies on expert experience and knowledge, the intelligent model of the present invention can continuously optimize and adapt to new data changes, and has stronger generalization ability and robustness. At the same time, the probability value output by the model provides a confidence reference for attribute recognition, further enhancing the reliability of the recognition results.
[0078] The present invention innovatively introduces a dynamic candidate attribute set partitioning mechanism, which intelligently partitions candidate attributes into single forms and combined forms according to attribute probability distributions and correlations. This dynamic partitioning strategy effectively addresses the diversity and complexity of log attributes and avoids the recognition limitations caused by fixed classification in traditional manual marking. Combined with an intelligent optimization mechanism, based on the comprehensive scoring of machine learning models, the best attribute combination is automatically selected, greatly improving the accuracy of attribute recognition and reducing the workload of manual screening and decision-making.
[0079] When evaluating the algorithm, the present invention does not simply apply the classic evaluation method, the F-measure value, in the field of process mining. Instead, it combines cross-validation with F-measure value evaluation. By dividing the dataset into subsets for iterative training and testing, it fully explores the performance of the model under different data distributions, effectively reducing the accidental errors caused by single testing, and at the same time avoiding the problems of overfitting and underfitting of the model. This makes the evaluation results more stable and reliable, and can more truly reflect the performance of the model on the overall dataset, ensuring the robustness of the model under different data distributions. Brief Description of the Drawings
[0080] Figure 1 It is a flowchart of the method for automatically identifying the attributes of business process log cases of the present invention.
[0081] Figure 2 It is an architecture diagram of the system for automatically identifying the attributes of business process log cases of the present invention. Detailed Embodiments
[0082] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments:
[0083] The flowchart of the present invention is as Figure 1 shown. Starting from the business process event log, first, the feature information of each column is extracted. Secondly, the range of candidate key attributes is narrowed down to judge single candidate attributes. If there are multiple candidate attributes, process model discovery and evaluation are carried out. Finally, the best candidate combination is selected according to the evaluation results, and the event log after identifying the case attributes is output. The present invention also designs a system for automatically identifying the attributes of business process log cases according to the above process. The main functional modules of this system include: an event log acquisition and feature extraction module, a candidate key attribute range narrowing module, a candidate attribute set discrimination module, a process model discovery and evaluation module, and a best candidate combination selection module.
[0084] A method for automatically identifying the attributes of business process log cases includes the following steps:
[0085] Step 1: Obtain the business process event log, extract features from the event log, and analyze the features of each column. The specific process is as follows:
[0086] Step 1.1, the obtained business process event log consists of several columns, each column contains several values consisting of letters, numbers, and special characters;
[0087] Step 1.2: Extract five features from the values of each column of the event log, including letter ratio, number ratio, special character ratio, unique value ratio and average string length. Suppose the column set A of the event log is A={A 1 ,A 2 ,...,A k}, where the kth column A k Contains n string values: A k ={s 1 ,s 2 ,...,s n}, each string consists of multiple characters, the i-th string is recorded as: s i ={c 1 ,c 2 ,...,c m}, m is the total number of characters; each character may be a letter, a number or a special character;
[0088] The letter ratio is a measure of the proportion of English letters in a string. The calculation process is:
[0089] First, define the indicator function of letter ratio as shown in formula (1):
[0090]
[0091] Among them, c j is the jth character; letter (·) is the indicator function of letter ratio; 26 letters include uppercase and lowercase;
[0092] Then, the letter ratio is calculated as shown in formula (2):
[0093]
[0094] in, Letter ratio; |s i | is i The number of characters;
[0095] The digital ratio is a measure of the proportion of numbers (0-9) in a string. The calculation process is:
[0096] First, define the indicator function of the digital ratio as shown in formula (3):
[0097]
[0098] Among them, digit(·) is an indicator function for the digital ratio;
[0099] Then, the digital ratio is calculated as shown in formula (4):
[0100]
[0101] where D(·) is the digital ratio;
[0102] The special character ratio measures the proportion of non-alphabetic and non-numeric characters in a string. For example,.,!:@-#%, and the calculation process is as follows:
[0103] First, define the indicator function of the special character ratio as shown in formula (5):
[0104]
[0105] where Π special (·) is the indicator function of the special character ratio;
[0106] Then, the special character ratio is calculated as shown in formula (6):
[0107]
[0108] where S(·) is the special character ratio;
[0109] Measure the ratio of unique values in a column, that is, the ratio of the number of different string values to the total number. The calculation process is as shown in formula (7):
[0110]
[0111] where U(·) is the ratio of unique values; unique(A k ) represents the number of different values in column A k If U(A k ) is close to 1, it means that the values in this column are almost non-repetitive with other columns;
[0112] The average string length measures the average length of strings in a column. The calculation process is as shown in formula (8):
[0113]
[0114] where M(·) is the average string length;
[0115] Step 1.3. Analyze the five characteristics of each column and initially determine the candidate key attribute types of the column; specifically:
[0116] The key attributes include case number, timestamp, activity, and resource;
[0117] If the current column contains both letters and numbers and has a high proportion of unique values, i.e., U(A k ) ≥ 0.8, then it is determined that the candidate key attribute type of the current column is the case number;
[0118] If the proportion of numbers and special characters in the current column is the highest compared to other columns and conforms to the time format, then it is determined that the candidate key attribute type of the current column is the timestamp;
[0119] If the current column has a high proportion of letters, i.e., and the average string length is longer than the average string lengths of other columns with a high proportion of letters, then it is determined that the candidate key attribute type of the current column is the activity;
[0120] If the current column has a high proportion of letters and a high proportion of unique values, i.e., then it is determined that the candidate key attribute type of the current column is the resource.
[0121] The method of the present invention is applicable to various business process scenarios, that is, for any given business process event log, the method of the present invention can automatically identify case attributes. For example, in the embodiment of the present invention, there is a patient treatment process event log of a hospital as shown in Table 1. In Table 1, the following columns are included: the first column is the patient number (Patient), the second column is the visit date (VisitDate), the third column is the treatment activity (Treatment Activity), the fourth column is the doctor name (Doctor Name), and the fifth column is the treatment result (Treatment Result).
[0122] Table 1 Example of Event Log
[0123]
[0124]
[0125] Table 2 Feature Extraction Results
[0126]
[0127] Extract features for each column to train a classifier to identify key attributes. After extracting features from the data in Table 1, the results shown in Table 2 can be obtained. Through preliminary analysis, it can be seen that the first column contains both letters and numbers and has a high proportion of unique values, and may be the case number; the second column contains numbers and special characters and conforms to the time format, and may be the timestamp; the third column has a high proportion of letters, and the average string length is longer than the average string lengths of other columns with a high proportion of letters, and may be the activity; the fourth column and the fifth column have the same proportion of letters and unique values, and further judgment is required.
[0128] Step 2: According to the characteristics of each column obtained in Step 1, use the supervised machine learning method to construct a binary classifier for each key attribute, output the probabilities that each column belongs to the four key attributes respectively, and select the top N attributes with the highest probability for each column to form the candidate attribute set for each column.
[0129] The key attributes include case number, activity, timestamp, and resource; the case number is an identifier used to uniquely identify a specific process instance, and each case represents a complete business process, including all activities from start to end; the activity refers to the specific operations or tasks executed in the business process, and each activity represents an identifiable step, usually associated with business operations, decisions, or events; the timestamp refers to the identifier that records the time when the activity occurs, usually represented in the form of date and time, and the timestamp can be used to analyze the execution order and duration of activities; the resource refers to the personnel, systems, or devices that participate in or are responsible for the activity during the execution of the activity, and the resource can be human resources, technical resources, or other support resources; attributes other than the above four key attributes are non-key attributes and can be directly marked as others after judging the key attributes.
[0130] The supervised machine learning method specifically uses the gradient boosting decision tree algorithm model. During the training process of the gradient boosting decision tree algorithm model, the columns in the known event log are marked as the corresponding label categories, and the gradient boosting decision tree algorithm model is used to classify the column features. Four binary classifiers are established to respectively identify the four key attributes of case number, activity, timestamp, and resource. The five features of the letter ratio, number ratio, special character ratio, unique value ratio, and average string length of each column in the event log are input into the gradient boosting decision tree algorithm model, and each column is marked as a target category or a non-target category.
[0131] For each binary classifier, multiple gradient boosting decision tree algorithm models are trained respectively. The gradient boosting method is used to optimize the loss function. The model builds a new decision tree in each iteration. The new decision tree learns the samples that the previous all decision trees failed to correctly predict to reduce the error. First, initialize the model as a constant term, representing the initial prediction value; for the t-th iteration, calculate the residual as shown in formula (9):
[0132]
[0133] where is the residual of the -th sample in the t-th iteration; is the true label of the -th sample; is the predicted value of the -th sample in the (t - 1)-th iteration;
[0134] Train a new decision tree with the residual as the target, and update the predicted value of the gradient boosting decision tree algorithm model as shown in formula (10):
[0135]
[0136] where is the predicted value of the -th sample in the t-th iteration; v is the learning rate, usually taking the value of 0.1; h t (·) is the new decision tree obtained in the t-th iteration; is the -th sample;
[0137] The final gradient boosting decision tree algorithm model will output the probabilities that each column belongs to the case number, activity, timestamp, or resource. The probability calculation formula is as shown in formula (11):
[0138]
[0139] where P(·) is the probability; is the predicted value of the final gradient boosting decision tree algorithm model for the -th sample ;
[0140] According to the output probabilities, select the top N attributes to form the candidate attribute set for each column. The specific value range of N is 1 - 5, and the specific value can be adjusted according to the actual event log characteristics.
[0141] In the embodiments of the present invention, four binary classifiers are constructed using the gradient boosting decision tree algorithm model. The five features of each column are input into the gradient boosting decision tree algorithm model to predict the probabilities that each column belongs to the case number, timestamp, resource, and activity respectively. The prediction results are shown in Table 3.
[0142] Table 3 Classification Results of Classifiers
[0143]
[0144] Step 3: According to the number of attributes included in the candidate attribute set of each column, divide the candidate attribute set of each column obtained in Step 2 into a single candidate attribute or a candidate attribute combination.
[0145] If there is only 1 attribute in the candidate attribute set of the current column (i.e., when the number of included attributes N is 1), it is determined that the candidate attribute set of the current column is a single candidate attribute. In this case, this 1 attribute in the candidate attribute set is the key attribute, and this key attribute is the attribute name of the current column; for example, if the candidate attribute set of a certain column only contains the "case number" attribute, then this column is identified as the case number attribute;
[0146] If there is at least one attribute in the candidate attribute set of the current column (that is, when the number N of included attributes is greater than 1, and the candidate attribute set contains two or more attributes), then the candidate attribute set of the current column is determined as a candidate attribute combination. In this case, it is necessary to continue judgment and evaluation to determine the best attribute combination. The single candidate attribute means that when the probability of a certain column belonging to a certain key attribute is the largest, it can be judged as a single candidate attribute; the candidate attribute combination means that when the probabilities of multiple columns belonging to a certain key attribute are all the largest, these columns are judged as multiple candidate attributes, and further judgment is required to distinguish them.
[0147] In the embodiment of the present invention, from the partitioning result, it can be seen that N = 1, and Patient, Visit Date, and TreatmentActivity are all in the form of single candidate attributes. Patient corresponds to "case number", Visit Date corresponds to "timestamp", and Treatment Activity corresponds to "activity"; since the "resource" probability results of Doctor Name and Treatment Result are the same, N = 2, so they are in the form of candidate attribute combinations, and the candidate columns in the candidate attribute set are all the candidate columns of the "resource" key attribute.
[0148] Step 4: Use the model discovery algorithm for the candidate attribute combinations obtained in Step 3 to discover the process model, evaluate the performance of each candidate attribute combination by means of cross-validation, and calculate the scores of each column for each candidate attribute combination.
[0149] The model discovery algorithm adopted is specifically the inductive mining algorithm. This algorithm adopts the idea of divide and conquer, decomposes the problem of discovering the process model of an event log into the problem of discovering the sub-processes of multiple sub-logs obtained by splitting the event log. First, initialize the process model as empty; recursively partition the event log, and each partition is based on the discrimination rules in the event log, such as activity name, resource type, etc.; apply the inductive mining algorithm to each sub-log to discover the sub-process model, and gradually build an accurate and concise process model; finally, merge the sub-process models into a complete process model to form a process description of the entire event log.
[0150] Evaluate the quality of the model mined by the inductive mining algorithm for each candidate attribute combination by combining cross-validation with the F-measure value, a model evaluation metric in the field of process mining; calculate the score of the F-measure value. First, randomly divide the dataset into two subsets, and try to keep the data distribution consistent for each subset; for each candidate attribute combination, perform cross-validation. In each iteration, select one subset as the test set and the other subset as the training set. Use the training set to train the model and use the test set to evaluate the performance of the process model, and record the F-measure value; swap the training set and the test set for re-evaluation.
[0151] The F-measure value adopted by the present invention is the harmonic mean of the fitness and the accuracy. The fitness quantifies the ability of the process model to regenerate the recorded traces in the event log, and the accuracy quantifies the ability of the process model to generate only the recorded traces in the event log. The calculation of the F-measure value is shown in formula (12):
[0152]
[0153] where F-measure(·) is the F-measure value; fitness(·) is the fitness; precision(·) is the accuracy; L is the event log; M is the process model mined from the event log.
[0154] Finally, take the average F-measure value as the score of the final required candidate attribute combination.
[0155] Table 4 Scores of candidate attribute combinations
[0156]
[0157] In the embodiment of the present invention, after two-fold cross-validation, the scores of the candidate attribute combinations of Doctor Name and Treatment Result are shown in Table 4, where combination 1 identifies Doctor Name as "resource", and combination 2 identifies TreatmentResult as "resource".
[0158] Step 5: According to the scores of each candidate attribute combination obtained in step 4, select the candidate attribute combination with the highest score as the final key attribute for output.
[0159] Table 5 Final event log marking results
[0160]
[0161]
[0162] In the embodiments of the present invention, finally, combination 1 is selected, that is, "Resource" in the "Doctor Name" column. The results are shown in Table 5. It can be seen from the table that the four key attributes of the case number, timestamp, activity, and resource in the event log are accurately and automatically identified and marked. Compared with manual marking, the method of the present invention can avoid human errors. When marking large-scale event logs, the key attributes in the event log can be effectively marked faster and more accurately, which is convenient for subsequent other process mining tasks and analysis.
[0163] Based on the above method process, the present invention designs a system for automatically identifying case attributes of business process logs. When the business process event log is input into this system, an event log with identified case attributes will be obtained as output, as Figure 2 shown. The system includes the following modules:
[0164] An event log acquisition and feature extraction module, which is used to acquire the business process event log, extract features from the event log, and analyze the features of each column;
[0165] A module for narrowing the range of candidate key attributes, which is used to build a binary classifier for each key attribute by using supervised machine learning methods, output the probabilities that each column belongs to the four key attributes respectively, and select the N attributes with the highest probability for each column to form the candidate attribute set for each column;
[0166] A candidate attribute set discrimination module, which is used to determine whether the candidate attribute set is a single candidate attribute or a candidate attribute combination;
[0167] A process model discovery and evaluation module, which uses a model discovery algorithm for the candidate attribute combination to discover the process model, evaluates the performance of each candidate attribute combination by cross-validation, and calculates the scores of each column for each candidate attribute combination;
[0168] A module for selecting the best candidate combination, which is used to select the candidate attribute combination with the highest score as the final key attribute output.
[0169] The present invention also designs a storage medium storing a program, and when the program is executed by a processor, the above-mentioned method for automatically identifying case attributes of business process logs is implemented.
[0170] The present invention also designs a computing device, including a processor and a memory for storing the program executable by the processor. When the processor executes the program stored in the memory, the above-mentioned method for automatically identifying case attributes of business process logs is implemented. The computing device can be a desktop computer, a laptop computer, a smart phone, a PDA handheld terminal, a tablet computer, a programmable logic controller (PLC), or other terminal devices with processor functions.
[0171] The present invention provides a new solution to the problems of high error rate and low efficiency that are prone to occur in the case of manually annotating event log attributes in the traditional way. It can reduce the incidence of human errors, improve the accuracy and reliability of data, and at the same time support fast data processing and analysis, helping enterprises to respond to market changes in a timely manner. In addition, it can also enable process mining technology to be applied to more fields, such as medical, financial, manufacturing, etc., and has practical promotion value and is worthy of promotion.
[0172] Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Changes, modifications, additions or substitutions made by those skilled in the art within the scope of the essence of the present invention should also fall within the protection scope of the present invention.
Claims
1. A method for automatically identifying attributes of business process log cases, characterized in that: The steps include: Step 1: Obtain the business process event log, extract features from the event log, and analyze the features of each column; Step 2: Use supervised machine learning methods to build a binary classifier for each key attribute, output the probability that each column belongs to the four key attributes, and select the N attributes with the highest probability for each column to form a candidate attribute set for each column; Step 3: Divide the candidate attribute set of each column into single candidate attributes or candidate attribute combinations according to the number of attributes contained in the candidate attribute set of each column; Step 4: Combine candidate attributes and use the model discovery algorithm to discover the process model. Evaluate the performance of each candidate attribute combination by cross-validation and calculate the score of each candidate attribute combination in each column. Step 5: According to the score of each candidate attribute combination, select the candidate attribute combination with the highest score as the final key attribute output.
2. The method for automatically identifying business process log case attributes according to claim 1 is characterized in that: The specific process in step 1 is: Step 1.1, the obtained business process event log consists of several columns, each column contains several values consisting of letters, numbers, and special characters; Step 1.2: Extract five features from the values of each column of the event log, including letter ratio, number ratio, special character ratio, unique value ratio and average string length; suppose the column set A of the event log is A = {A1, A2, ..., A k }, where the kth column A k Contains n string values: A k ={s1, s2, ..., s n }, each string consists of multiple characters, the i-th string is recorded as: s i ={c1, c2, ..., c m }, m is the total number of characters; each character is a letter, number or special character; The letter ratio is a measure of the proportion of English letters in a string. The calculation process is: First, define the indicator function for letter proportions: Among them, c j is the jth character; letter (·) is the indicator function of letter ratio; 26 letters include uppercase and lowercase; Then, calculate the letter ratio: in, Letter ratio; |s i | is i The number of characters; The digital ratio is a measure of the proportion of numbers in a string. The calculation process is: First, define the indicator function for the digital ratio: Among them, digit (·) is the indicator function of the digital ratio; Then, calculate the digital ratio: Where D(·) is the digital ratio; The special character ratio is a measure of the proportion of non-alphabetic and non-numeric characters in a string, including .! :@-#%. The calculation process is: First, define the indicator function for the special character ratio: Among them, special (·) is the indicator function of the special character ratio; Then, calculate the special character ratio: Among them, S(·) is the special character ratio; The unique value ratio is the ratio of the number of different string values to the total number, and the calculation formula is: Among them, U(·) is the unique value ratio; unique(A k ) represents column A k The number of different values in , if U(A k ) is close to 1, indicating that the values in this column are not repeated in other columns; The calculation process of average string length is: Where M(·) is the average string length; Step 1.3: Analyze the five characteristics of each column and preliminarily determine the candidate key attribute types of the column; specifically: Key attributes include case number, timestamp, activity, and resource; If the current column contains letters and numbers, and the unique value ratio is high, that is, U(A k )≥0.8, then the candidate key attribute type of the current column is determined to be case number; If the current column has the highest digital ratio and special character ratio compared to other columns and conforms to the time format, the candidate key attribute type of the current column is determined to be timestamp; If the current column contains a high proportion of letters, that is, If the average string length is longer than the average string length of other columns with high letter ratios, the candidate key attribute type of the current column is determined to be active; If the current column contains a high proportion of letters and a high proportion of unique values, that is, It is determined that the candidate key attribute type of the current column is resource.
3. The method for automatically identifying business process log case attributes according to claim 2 is characterized in that: In step 2, the key attributes include case number, activity, timestamp, and resource; the case number is an identifier used to uniquely identify a specific process instance, and each case represents a complete business process, including all activities from the beginning to the end; the activity refers to a specific operation or task performed in a business process, and each activity represents a step, which is usually associated with a business operation, decision or event; the timestamp refers to an identifier that records the time when an activity occurs, usually expressed in the form of date and time, and the timestamp is used to analyze the execution sequence and duration of the activity; the resource refers to the person, system or equipment that participates in or is responsible for the activity when the activity is executed, and the resource includes human resources, technical resources or other supporting resources; the attributes other than the above four key attributes are non-key attributes, and are directly marked as others after the key attributes are determined; The supervised machine learning method specifically adopts the gradient boosting decision tree algorithm model. During the training process of the gradient boosting decision tree algorithm model, the columns in the known event logs are marked as corresponding label categories. The gradient boosting decision tree algorithm model is used to classify the column features. Four binary classifiers are established to respectively identify the four key attributes of case number, activity, timestamp, and resource. The five features of the letter ratio, number ratio, special character ratio, unique value ratio, and average string length of each column in the event log are input into the gradient boosting decision tree algorithm model. Each column is marked as a target category or a non-target category. For each binary classifier, multiple gradient boosting decision tree algorithm models are trained separately, and the loss function is optimized using the gradient boosting method. The model establishes a new decision tree in each iteration. The new decision tree learns the samples that all previous decision trees failed to correctly predict in order to reduce the error. First, the model is initialized as a constant term, which represents the initial prediction value; for the tth iteration, the residual is calculated: in, For the The residual of the tth iteration of samples; For the The true labels of samples; For the The predicted value of the t-1th iteration of the sample; Train a new decision tree, using the residual as the target, and update the prediction value of the gradient boosting decision tree algorithm model: in, For the The predicted value of the tth iteration of the sample; v is the learning rate; h is the t (·) is the new decision tree obtained in the tth round of iteration; For the samples; The final gradient boosting decision tree algorithm model outputs the probability of each column belonging to a case number, activity, timestamp, or resource. The probability calculation formula is: Where P(·) is the probability; For the final gradient boosting decision tree algorithm model Samples The predicted value of According to the output probability, the highest N attributes are selected to form the candidate attribute set of each column. The specific value range of N is 1-5, and the specific value is adjusted according to the actual event log characteristics.
4. The method for automatically identifying business process log case attributes according to claim 3 is characterized in that: In step 3, if the number N of attributes included in the candidate attribute set of the current column is 1, the candidate attribute set of the current column is determined to be a single candidate attribute. In this case, the attribute in the candidate attribute set is the key attribute, and the key attribute is the attribute name of the current column. If the number N of attributes included in the candidate attribute set of the current column is greater than 1, the candidate attribute set of the current column is determined to be a candidate attribute combination. In this case, further judgment and evaluation are required to determine the best attribute combination.
5. The method for automatically identifying business process log case attributes according to claim 4 is characterized in that: In step 4, the model discovery algorithm used is specifically an inductive mining algorithm, which adopts the idea of divide and conquer, and decomposes the problem of discovering a process model of an event log into the problem of discovering sub-processes of multiple sub-logs obtained by splitting the event log. First, the process model is initialized to be empty; the event log is recursively split, and each split is based on the distinction rules in the event log, including activity name and resource type; Apply inductive mining algorithms to each sub-log to discover sub-process models and gradually build an accurate and concise process model; finally, merge the sub-process models into a complete process model to form a process description of the entire event log; The quality of the model mined by each candidate attribute combination using the inductive mining algorithm is evaluated by combining cross-validation with the model evaluation index F-measure in the process mining field. The F-measure score of the model is calculated by first randomly dividing the data set into two subsets, and each subset maintains the distribution consistency of the data. For each candidate attribute combination, cross-validation is performed. In each iteration, one of the subsets is selected as the test set and the other subset is selected as the training set. The model is trained using the training set, and the performance of the process model is evaluated using the test set, and the F-measure value is recorded. The training set and the test set are exchanged for re-evaluation. The F-measure value is the harmonic mean of fit and accuracy. The fit quantifies the ability of the process model to regenerate the traces recorded in the event log, and the accuracy quantifies the ability of the process model to only generate the traces recorded in the event log. The F-measure value calculation formula is: Where, F-measure(·) is the F-measure value; fitness(·) is the degree of fit; precision(·) is the accuracy; L is the event log; M is the process model mined from the event log; Finally, the average F-measure value is taken as the score of the final required candidate attribute combination.
6. A business process log case attribute automatic identification system, characterized in that: The method for automatically identifying case attributes of a business process log according to any one of claims 1 to 5 is adopted, wherein the input of the system is a business process event log, and the output is an event log with identified case attributes; the system comprises the following modules: The event log acquisition and feature extraction module is used to acquire business process event logs, extract features from the event logs, and analyze the features of each column; The module for narrowing down the range of candidate key attributes is used to build a binary classifier for each key attribute using the supervised machine learning method, output the probability that each column belongs to the four key attributes, and select the N attributes with the highest probability for each column to form the candidate attribute set for each column; A candidate attribute set discrimination module is used to determine whether a candidate attribute set is a single candidate attribute or a combination of candidate attributes; The process model discovery and evaluation module uses the model discovery algorithm to discover the process model for the candidate attribute combination, evaluates the performance of each candidate attribute combination through cross-validation, and calculates the score of each candidate attribute combination in each column; The best candidate combination selection module is used to select the candidate attribute combination with the highest score as the final key attribute output.
Citation Information
Patent Citations
Decision tree model training method, and method and apparatus for determining data attributes in OCR result
CN107273883A
Decision tree model-based data processing method and related equipment
CN114219596A
Extracting and labeling custom information from log messages
US20180285432A1
Field operations framework
WO2024254075A2
Cited By
System log multi-dimensional dynamic filtering and clearing control method and device
CN121579314A