An activity sequence sampling process mining method and system for medical data

Through event embedding model and dynamic weighted trajectory importance sampling method, the problems of model complexity and key events ignored in traditional process mining technology are solved, and efficient and concise process model generation is achieved, which is suitable for large-scale medical data processing.

CN120032911BActive Publication Date: 2025-07-08QINGDAO UNIV OF TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510519781.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-07-08
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

When traditional process mining technology processes large-scale medical data logs, it faces the problems of high-frequency repeat sequences that lead to the complexity of the model, the critical paths are masked, the activities of low-frequency key event are ignored, and the model efficiency and quality are difficult to balance.

Method used

Through the event embedding model, semantically similar event activities are mapped to low-dimensional space and clustered to generate unified semantic identifiers. Combined with the dynamically weighted trajectory importance sampling method, key low-frequency events and their direct follow-up relationships are preferred to generate high-quality process models.

Benefits of technology

Significantly compress log size, reduce model complexity, retain key low-frequency events, generate process models with high fit and low complexity, and improve processing efficiency and model readability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032911B_ABST
    Figure CN120032911B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for mining activity sequence sampling process of medical data, belonging to the field of information technology. The method steps are as follows: obtaining the original event log and performing data cleaning; clustering semantically similar event activities into the same cluster based on the event embedding model and the DBSCAN algorithm, and generating a unique semantic identifier for each cluster; generating a redundant-free event log; sampling the redundant-free event log based on the comprehensive importance score of the trace to obtain a sample event log; and inputting the sample event log into an inductive mining algorithm to generate a corresponding process model. The system includes a data acquisition and cleaning module, a clustering module, a redundancy removal module, a sampling module, and a process model generation module. The present invention greatly improves the processing efficiency while ensuring the representativeness of the log, is applicable to the large-scale log processing requirements in the medical field, and provides reliable support for business process optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of information technology, and in particular relates to an activity sequence sampling process mining method and system applied to medical data. Background Art

[0002] Traditional process mining technology mainly relies on algorithms such as Alpha Miner and Heuristic Miner to directly build process models through event logs. However, these methods face significant challenges when processing large-scale logs. For example, high-frequency repetitive sequences (such as "vital signs monitoring") appear in large numbers in medical data logs, resulting in a sharp increase in the number of model nodes and obscuring the critical path. In addition, among the existing log sampling methods, the LogRank method and the set coverage sampling method both reduce the log size by selecting subsets, while the LogRank++ method distinguishes event contexts based on label refinement. These methods all have certain limitations. They are usually unable to effectively process semantically similar sequences, and low-frequency but critical behaviors (such as "emergency surgery") may be missed during the sampling process, resulting in overfitting or underfitting of the model. The specific defects are as follows: First, redundant events will lead to model complexity; for example, repeated sequences (such as periodic detection) will increase the number of model nodes, mask the critical path, and reduce the readability and practicality of the model; second, the sampling method loses behavior; random or heuristic sampling ignores low-frequency critical event activities (such as "emergency surgery"), resulting in the model being unable to accurately reflect the entire business process; third, it is difficult to balance efficiency and quality; traditional distributed solutions (such as MapReduce) rely on hardware expansion and are not optimized from the data preprocessing level, making it difficult to take into account both processing efficiency and model quality. Summary of the invention

[0003] In order to solve the above problems, the present invention proposes an activity sequence sampling process mining method and system applied to medical data. Through the event embedding model, semantically similar event activities are mapped to a low-dimensional space and clustered to generate a unified semantic identifier, which significantly compresses the log size. Based on the dynamic weighted trajectory importance sampling method, key low-frequency event activities and their direct follow-up relationships are preferentially retained to ensure the behavioral completeness of the sample log. Finally, the process model is generated through the Inductive Miner algorithm.

[0004] The technical solution of the present invention is as follows:

[0005] An activity sequence sampling process mining method applied to medical data includes the following steps:

[0006] Step 1: Get the original event log and clean the data;

[0007] Step 2: Cluster semantically similar event activities into the same cluster based on the event embedding model and the DBSCAN algorithm, and generate a unique semantic identifier for each cluster;

[0008] Step 3: Replace the semantically similar event activities in the event log with the semantic identifier corresponding to the cluster to generate a redundant-free event log;

[0009] Step 4: Sample the redundant-free event log based on the comprehensive importance score of the trace to obtain a sample event log;

[0010] Step 5: Input the sample event log into the inductive mining algorithm to generate the corresponding process model, and evaluate the quality and complexity of the process model.

[0011] Further, the specific process of Step 1 is as follows:

[0012] Step 1.1: Export the standard event log from the hospital information system of the medical institution as the original event log. The format of the event log is a structured file conforming to the XES specification, including the event activity name, timestamp, case ID, and the participant who executes the activity;

[0013] The event log is a set containing several traces, and a trace is a sequence composed of several event activities;

[0014] Define the original event log as:

[0015] (1);

[0016] where, is the original event log; is the total number of traces; is the th trace, specifically:

[0017] (2);

[0018] where, is the th event activity in the th trace;

[0019] Step 1.2: Perform data cleaning on the original event log, remove the traces that do not contain complete lifecycle information, and standardize the labels of the event activities to unify the labels of the event activities.

[0020] Further, the specific process of Step 2 is as follows:

[0021] Step 2.1: Train the event embedding model; for each trace in the cleaned event log, use a length of The sliding window method forms context pairs All the context pairs constitute the training sample set which is used to train the event embedding model; where represents the target event activity represents the context event activity co-occurring with the target event activity within the sliding window;

[0022] During training, the set parameters include: vector dimension, negative sampling number, learning rate;

[0023] Set the objective function for model training as:

[0024] (3);

[0025] where is the event embedding model; represents the vector inner product operation; is the negative sampling candidate event activity; is the complete set of event activities;

[0026] When the number of training iterations reaches the preset number, the training ends;

[0027] Step 2.2. Use the trained event embedding model to vectorize the event activities to obtain an embedding sequence;

[0028] Step 2.3. Apply the DBSCAN algorithm to cluster the embedding sequence; The DBSCAN algorithm is a density-based clustering algorithm that will automatically divide the density-connected embedding sequences into the same cluster by calculating the density reachability between the embedding sequences, thereby identifying semantically similar event activities; During the clustering process, distributed computing is used to slice the embedding sequence to 6 Spark nodes for parallel computing; The final clustering result contains several clusters, and each cluster contains several semantically similar event activities;

[0029] The DBSCAN algorithm parameters include: neighborhood radius, minimum number of samples;

[0030] Step 2.4. Preset the first threshold and the second threshold, count the number of members of each cluster and compare it with the first threshold, calculate the minimum Euclidean distance between each cluster and other clusters in the embedding space, and compare it with the second threshold; If the th cluster has a number of members less than the first threshold and its minimum Euclidean distance from other clusters in the embedding space is greater than the second threshold, then the th cluster is an isolated noise cluster; Assign a unique semantic identifier to each cluster other than the noise cluster, in the format of: ; is the semantic identifier corresponding to the th cluster.

[0031] Furthermore, the specific process of step 2.2 is as follows:

[0032] Step 2.2.1, construct the semantic representation space of activities based on the trained event embedding model:

[0033] (4);

[0034] where, is the semantic representation space of the event activity ; represents the embedded vector representation of the event activity ; represents the dimension of the embedded vector;

[0035] Step 2.2.2, extract subsequences of a fixed length from each trajectory in a sliding window manner, denoted as the embedded sequence: ;

[0036] (5);

[0037] where, is the embedded sequence, representing the representation sequence of the subsequences extracted by the sliding window in the embedding space; represents the th event activity in the embedded sequence corresponding to the embedded vector representation.

[0038] Furthermore, the specific process of step 3 is as follows:

[0039] Replace the event activities with similar semantics in the event log with the semantic identifiers corresponding to the clusters, in the form of:

[0040] (6);

[0041] Thus, a redundancy-removed event log with a simplified structure is formed:

[0042] (7);

[0043] where, is the update operation; is the redundancy-removed event log; , are respectively the th and th event activities in the th trajectory; represents the th trajectory, the an event activity, and the event activity is an event activity with similar semantics; for the th trajectory after clustering; indicating the th event activity in the th trajectory after clustering, and the event activity is the semantic identifier corresponding to the cluster.

[0044] Furthermore, the specific process of step 4 is as follows:

[0045] Step 4.1: Predetermine a third threshold. If the occurrence frequency of an event activity is less than the third threshold and it is confirmed by medical experts or a domain knowledge base to have key business significance, then mark the event activity as a key low-frequency event activity, mark the trajectory where the key low-frequency event activity is located as a low-frequency trajectory, and mark the remaining trajectories as normal trajectories;

[0046] Step 4.2: Calculate the activity importance and structural importance of each trajectory, and calculate the comprehensive importance score by weighted calculation;

[0047] The calculation formula for activity importance is:

[0048] (8);

[0049] where, is the activity importance of the th trajectory; is the coverage rate of the event activity in the redundant event log;

[0050] The calculation formula for structural importance is:

[0051] (9);

[0052] where, is the structural importance of the th trajectory; , are respectively the th and the th event activities in the th trajectory; represents the direct following frequency of the event activity pair in the redundant event log;

[0053] The comprehensive importance score of the trajectory is calculated by linear weighting:

[0054] (10);

[0055] where, is the Comprehensive importance score of a trajectory; is the weight;

[0056] When calculating the comprehensive importance score of low-frequency trajectories, set the weight coefficient , and for ordinary trajectories, set ;

[0057] Step 4.2. Sort all trajectories according to the comprehensive importance score, and select the top trajectories with the highest comprehensive importance scores as the sample event log:

[0058] (11);

[0059] Among them, is the sample event log; is all trajectories the top trajectories with the highest comprehensive importance scores; is the number of target trajectories under the preset sampling rate; is the sampling rate, and the sampling rate supports dynamic adjustment by the user according to computing resources, with the adjustment range being 20% - 30%.

[0060] Furthermore, in the said Step 5, the process model includes a Petri net, and the Petri net contains places and transitions;

[0061] Evaluate the quality of the process model through a fitness index, and the calculation formula of the fitness index is:

[0062] (12);

[0063] Among them, is the fitness index; is the process model; represents length calculation; represents the trajectory in the replay path, and the calculation formula is:

[0064] (13);

[0065] Among them, is the transition sequence in the process model , is the th transition; is the set composed of all transitions in the process model ; is the th event activity and the a transition matching function; is the total number of event activities;

[0066] By calculating the edge-node ratio of the process model to evaluate the complexity of the process model, the formula is:

[0067] (14);

[0068] where is the edge-node ratio of the process model ; is the edge of the process model corresponding to the number of transitions; is the node of the process model corresponding to the number of places;

[0069] If > 3, it is determined that the process model is overfitted and it is necessary to return to step 1 for reprocessing;

[0070] During online application, the sample event log is input into the inductive mining algorithm to generate the corresponding process model, which completes the process mining task.

[0071] An activity sequence sampling process mining system applied to medical data uses the activity sequence sampling process mining method applied to medical data as described above for process mining. The system includes a data collection and cleaning module, a clustering module, a redundancy removal module, a sampling module, and a process model generation module. The data collection and cleaning module is connected to the hospital information system of the medical institution to obtain real-time data and clean the obtained standard event log. The clustering module clusters semantically similar event activities into the same cluster based on the event embedding model and the DBSCAN algorithm, and generates a unique semantic identifier for each cluster. The redundancy removal module is used to replace semantically similar event activities in the event log with the semantic identifier corresponding to the cluster to generate a redundancy-removed event log. The sampling module samples the redundancy-removed event log based on the comprehensive importance score of the trace to obtain a sample event log. The process model generation module is used to input the sample event log into the inductive mining algorithm to generate the corresponding process model, and evaluate the quality and complexity of the process model.

[0072] Beneficial technical effects brought by the present invention: Through the event embedding model, the present invention maps semantically similar repetitive sequences into a low-dimensional space and clusters them to generate unified semantic identifiers, significantly compressing the log scale and reducing the model complexity. Based on the dynamic weighted trajectory importance sampling method, it preferentially retains low-frequency but critical event activities and their direct following relationships, ensuring the behavioral completeness of the sample log and avoiding the loss of critical paths. Finally, a high-quality process model is generated through Inductive Miner, which has both high fitting degree and low complexity, effectively supporting subsequent process mining and analysis. While ensuring the representativeness of the log, this method greatly improves the processing efficiency, is applicable to the large-scale log processing requirements in the medical field, and provides reliable support for business process optimization. Description of the Drawings

[0073] Figure 1 It is a flowchart of the activity sequence sampling process mining method applied to medical data in the present invention. Detailed Embodiments

[0074] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments:

[0075] As Figure 1 shown, an activity sequence sampling process mining method applied to medical data includes the following steps:

[0076] Step 1: Obtain the original event log and perform data cleaning; the specific process is as follows:

[0077] Step 1.1: Export the standard event log from the hospital information system (HIS) of the medical institution as the original event log. The format of the event log is a structured file conforming to the XES specification, including fields such as event activity name, timestamp, case ID, and participants performing the activity;

[0078] The event log is a set containing several traces, and a trace is a sequence composed of several event activities;

[0079] Define the original event log as:

[0080] (1);

[0081] Among them, is the original event log; is the total number of traces; is the th trace, and a trace represents the complete activity process of a business instance:

[0082] (2);

[0083] Among them, For the th event activity in the th trajectory;

[0084] The structure of the trajectory is based on the linear-ordered event modeling foundation, providing a unified representation for subsequent structure compression and sampling.

[0085] Step 1.2: Clean the data of the original event log, remove the trajectories that do not contain complete life cycle information, standardize the labels of event activities, and unify the labels of event activities.

[0086] Step 2: Cluster the semantically similar event activities into the same cluster based on the event embedding model (Skip-gram model) and DBSCAN algorithm, and generate a unique semantic identifier for each cluster to reduce the modeling complexity; the specific process is as follows:

[0087] Step 2.1: Train the event embedding model; for each trajectory (i.e., each activity sequence) in the cleaned event log, use a sliding window method with a length of to form context pairs , and all the context pairs constitute the training sample set , which is used to train the event embedding model. Among them, represents the target event activity, and represents the context event activities that co-occur with the target event activity within the sliding window.

[0088] During training, the set parameters include: vector dimension, negative sampling number, and learning rate. The vector dimension is used to balance the computational efficiency and semantic expression ability; the negative sampling number is used to reduce the high-frequency event bias; the learning rate is determined through cross-validation;

[0089] Set the objective function for model training as:

[0090] (3);

[0091] Among them, is the event embedding model, which is used to map events to a low-dimensional vector space to capture semantic similarity; represents the vector inner product operation; is the negative sampling candidate event activity; is the complete set of event activities;

[0092] When the number of training iterations reaches the preset epoch: 200 times, the training ends.

[0093] Step 2.2: Use the trained event embedding model to vectorize the event activities to obtain an embedding sequence; the specific process is as follows:

[0094] Step 2.2.1: Different from the traditional redundancy judgment that relies on frequency or location, the present invention constructs a semantic representation space for activities based on the trained event embedding model:

[0095] (4);

[0096] Among them, is the semantic representation space of the event activity ; represents the embedding vector representation of the event activity ; represents the dimension of the embedding vector;

[0097] Step 2.2.2: Each trajectory extracts subsequences of a fixed length in a sliding window manner, denoted as the embedding sequence: ;

[0098] (5);

[0099] Among them, is the embedding sequence, representing the representation sequence of the subsequences extracted by the sliding window in the embedding space; represents the -th event activity in the embedding sequence, corresponding to the embedding vector representation;

[0100] Step 2.3: Apply the DBSCAN algorithm to the embedding sequence for clustering. The DBSCAN algorithm is a density-based clustering algorithm. By calculating the density reachability between the embedding sequences, it automatically divides the density-connected embedding sequences into the same cluster, thereby identifying semantically similar event activities. During the clustering process, distributed computing is adopted, and the embedding sequences are sliced to 6 Spark nodes for parallel computing. The final clustering result contains several clusters, and each cluster contains several semantically similar event activities.

[0101] The parameters of the DBSCAN algorithm include: neighborhood radius and minimum number of samples.

[0102] After clustering, semantic analysis is performed to evaluate the clustering effect. Specifically, by calculating the average Euclidean distance of the embedding sequences within each cluster, the degree of semantic similarity is confirmed to ensure that the events within the formed clusters have a high degree of semantic consistency.

[0103] Step 2.4: Preset the first threshold and the second threshold, count the number of members in each cluster and compare it with the first threshold, calculate the minimum Euclidean distance between each cluster and other clusters in the embedding space, and compare it with the second threshold; if the -th cluster If the number of members is less than the first threshold and its minimum Euclidean distance from other clusters in the embedding space is greater than the second threshold, then the th cluster is an isolated noise cluster. The noise cluster does not have semantic generality and is directly removed without entering step 3 for replacement. Assign a unified semantic identifier to each cluster other than the noise cluster, in the format of: ; is the semantic identifier corresponding to the th cluster.

[0104] Step 3: Replace the semantically similar event activities in the event log with the semantic identifiers corresponding to the clusters to generate a redundant-reduced event log;

[0105] Replace the semantically similar event activities in the event log with the semantic identifiers corresponding to the clusters, in the form of:

[0106] (6);

[0107] Thus, a redundant-reduced event log with a simplified structure is formed:

[0108] (7);

[0109] wherein, is an update operation; is a redundant-reduced event log; , are respectively the th and the th event activities in the th trajectory; represents the th event activity in the th trajectory, and this event activity is a semantically similar event activity; is the th trajectory after clustering; represents the th event activity in the th trajectory after clustering, and this event activity is the semantic identifier corresponding to the cluster;

[0110] This representation retains the control structure and semantic context of the original process, while reducing the number of duplicate nodes, providing a more readable basis for subsequent analyses such as graph modeling and visualization.

[0111] Step 4: Sample the redundant-reduced event log based on the comprehensive importance score of the trajectories to obtain a sample event log; the specific process is as follows:

[0112] Step 4.1: First, identify critical low-frequency event activities according to a pre-set third threshold. If the occurrence frequency of an event activity is less than the third threshold and it is confirmed by medical experts or a domain knowledge base to have critical business significance (such as involving patient safety or major medical decisions), then it is marked as a critical low-frequency event activity. The trajectory where the critical low-frequency event activity is located is marked as a low-frequency trajectory, and a higher weight is assigned to the low-frequency trajectory by increasing the weight value during the subsequent trajectory scoring process. The remaining trajectories are marked as ordinary trajectories. The value assigns a higher weight to the low-frequency trajectory. The remaining trajectories are marked as ordinary trajectories.

[0113] Step 4.2: Calculate the activity importance and structural importance of each trajectory, and calculate the comprehensive importance score by weighted calculation; this process aims to select a representative subset of trajectories, taking into account the main process and critical low-frequency behaviors, and avoiding the information loss problem caused by traditional sampling strategies (such as random sampling or frequency priority). The specific process is as follows:

[0114] Step 4.2.1: The activity importance is the coverage of all event activities in the trajectory in the de-duplicated event log, which is used to measure the contribution of each event activity in the trajectory to the overall process, and is defined as:

[0115] (8);

[0116] Among them, is the activity importance of the th trajectory; is the coverage rate of the event activity in the de-duplicated event log.

[0117] Step 4.2.2: The structural importance is the coverage of the direct following relationship in the trajectory in the de-duplicated event log, considering the ability of the direct following relationship in the trajectory to depict the overall process structure:

[0118] (9);

[0119] Among them, is the structural importance of the th trajectory; , are the th, th event activities in the th trajectory respectively; represents the direct following frequency of the event activity pair in the de-duplicated event log, that is, the number of times this activity sequence appears in the de-duplicated event log;

[0120] Step 4.2.3: Use linear weighting to calculate the comprehensive importance score of the trajectory:

[0121] (10);

[0122] Among them, is the comprehensive importance score of the th trajectory; is the weight;

[0123] When calculating the comprehensive importance score of low-frequency trajectories, set the weight coefficient to emphasize the role of activity importance in the comprehensive evaluation. For ordinary trajectories, set to ensure the trade-off evaluation between event activities and structures;

[0124] Step 4.3. Sort all trajectories according to the comprehensive importance score, and select the top trajectories with the highest comprehensive importance scores as the sample event log:

[0125] (11);

[0126] Among them, is the sample event log; is all trajectories the top trajectories with the highest comprehensive importance scores; is the number of target trajectories under the preset sampling rate; is the sampling rate, with a default value of 25%, and supports dynamic adjustment by users according to computing resources, with an adjustment range of 20% - 30%.

[0127] Step 5. Input the sample event log into the inductive mining algorithm to generate the corresponding process model, and evaluate the quality and complexity of the process model.

[0128] The process model includes a Petri net, and the Petri net contains places and transitions;

[0129] For the quality evaluation of the process model, verify the behavioral consistency between the sample event log and the original event log through the fitness index. The calculation formula of the fitness index is:

[0130] (12);

[0131] Among them, is the fitness index; is the process model; represents length calculation; represents the replay path of the trajectory in , and the calculation formula is:

[0132] (13);

[0133] Among them, is the transition sequence in the process model, is the th transition; is the set composed of all transitions in the process model; is the th event activity and the th transition matching function. If they match, it is 1; if not, it is 0. The matching relationship is determined according to the Token Replay algorithm; is the total number of event activities;

[0134] The complexity of the process model is evaluated by calculating the edge-node ratio of the process model , and the formula is:

[0135] (14);

[0136] Among them, is the edge-node ratio of the process model ; is the edge of the process model corresponding to the number of transitions; is the node

[0137] If > 3, it is determined that the process model is overfitted and it is necessary to return to step 1 for reprocessing;

[0138] During online application, the sample event log is input into the inductive mining algorithm to generate the corresponding process model, thus completing the process mining task.

[0139] Steps 2 and 4 are implemented through the Apache Spark distributed framework, and the clustering and sampling processes are executed in parallel according to the log shards.

[0140] An activity sequence sampling process mining system for medical data, including a data collection and cleaning module, a clustering module, a redundancy removal module, a sampling module, and a process model generation module; the data collection and cleaning module is connected to the hospital information system of a medical institution to obtain real-time data, and clean the obtained standard event log; the clustering module clusters semantically similar event activities into the same cluster based on an event embedding model and the DBSCAN algorithm, and generates a unique semantic identifier for each cluster; the redundancy removal module is used to replace semantically similar event activities in the event log with the semantic identifier corresponding to the cluster, generating a redundancy-removed event log; the sampling module samples the redundancy-removed event log based on the comprehensive importance score of the trace, obtaining a sample event log; the process model generation module is used to input the sample event log into an inductive mining algorithm to generate a corresponding process model, and evaluate the quality and complexity of the process model.

[0141] To prove the feasibility and superiority of the present invention, the following embodiments are given.

[0142] This embodiment is described by taking the medical business process optimization scenario as an example. This embodiment is based on the public dataset BPIC2012-W and simulates the diagnosis and treatment process of emergency patients in a hospital.

[0143] The hardware configuration of this embodiment is as follows:

[0144] Cluster environment: Apache Spark cluster (6 nodes, 64 cores / 256GB memory per node).

[0145] Storage: HDFS distributed storage system, and the log is sharded into 6 blocks (about 21,000 traces per block).

[0146] The software tools used in this embodiment are as follows:

[0147] Process mining tool: ProM 6.11 (integrated with custom plugins, supporting clustering, sampling, and visualization).

[0148] Embedding model training: Python 3.9 + Gensim library.

[0149] Distributed computing: Spark MLlib, Spark SQL (trace analysis)

[0150] The format of the public dataset BPIC2012-W is the XES standard event log, that is, BPIC2012-W.xes. Taking the public dataset BPIC2012-W as the original event log, this original event log contains a total of 13,087 traces, 262,200 event activities, the event activity types are divided into 36 types, and the trace length ranges from 12 steps to 306 steps. An example of a trace is: ;

[0151] Perform data cleaning on the original event log; first, eliminate invalid traces: remove traces that do not contain event activities of "discharge" or "termination of treatment" (a total of 1203 traces). Then, standardize the event activity labels: rename "Blood Test_A" and "Blood Test_B" to "Blood Test_Type 1" and "Blood Test_Type 2" respectively.

[0152] Preset the high-frequency threshold to 1000. In this embodiment, the number of occurrences of the event activity "Vital Sign Monitoring" and the event activity "Blood Test" are 8000 and 6500 respectively, both of which are greater than the high-frequency threshold. Then, determine that "Vital Sign Monitoring" and "Blood Test" are high-frequency events.

[0153] Extract all traces (i.e., all activity sequences) from the cleaned public dataset. For each activity sequence, construct a sliding window with a length of to form context pairs, and finally obtain a training sample set containing 120 million context pairs for training the event embedding model. When training, the set parameters include: the vector dimension is 128, the negative sampling number is 15, and the learning rate is 0.025.

[0154] After the training is completed, use the trained event embedding model to perform vector representation on the event activities, obtain the embedding sequence, and visually display the embedding sequence. In the embodiment of the present invention, the Euclidean distance between "Blood Test_Type 1" and "Blood Test_Type 2" in the embedding space is less than 0.3. Therefore, it is determined that the two are semantically similar events.

[0155] Use the DBSCAN algorithm to cluster the semantically similar event activities. The parameter settings of the DBSCAN algorithm include: the neighborhood radius is 0.7, which is determined by analyzing the distribution in the embedding space; the minimum number of samples is 5 to avoid over-segmentation. The input data of the DBSCAN algorithm is the embedding sequence of the extracted high-frequency events, with a total of 28 activities. During the clustering process, distributed computing is adopted, and the embedding vectors are sliced into 6 Spark nodes for parallel calculation of density reachability. By clustering, semantically similar event activities are merged to reduce log redundancy.

[0156] Example of clustering results:

[0157] Cluster 1: Event activities related to vital sign monitoring;

[0158] Containing event activities:

[0159] Vital Sign Monitoring_1 (frequency: 4200 times);

[0160] Vital Sign Monitoring_2 (frequency: 3800 times);

[0161] Vital signs monitoring_3 (frequency: 3500 times);

[0162] Semantic analysis: The average Euclidean distance of these event activities in the embedding space is 0.25, indicating that the semantics are highly similar (both represent regular monitoring of patient vital signs), proving that the clustering is accurate.

[0163] Semantic identifier assignment: ;in, represents cluster 1; is the semantic identifier corresponding to cluster 1, which is used to simplify the process modeling representation;

[0164] Cluster 2: blood test type event activities;

[0165] Events included:

[0166] Blood test_type 1 (frequency: 2700 times);

[0167] Blood test_type 2 (frequency: 2300 times);

[0168] Blood test_type 3 (frequency: 1500 times);

[0169] Semantic analysis: The cosine similarity of the embedding vector angle is 0.92, indicating that different detection types belong to the same semantic category (if the difference in the detection items is only the parameter adjustment), proving that the clustering is accurate.

[0170] Semantic identifier assignment: ;in, represents cluster 2; is the semantic identifier corresponding to cluster 2, which is used to simplify the process modeling representation;

[0171] Cluster 3: imaging examination event activity;

[0172] Events included:

[0173] CT scan_routine (frequency: 1200 times);

[0174] MRI scan_enhanced (frequency: 980 times);

[0175] X-ray examination_chest (frequency: 850 times);

[0176] Semantic analysis: Dense subregions (radius 0.4) are formed in the embedding space, indicating that the contextual associations of different imaging examinations are strong (e.g., they are all used for diagnostic support), proving that the clustering is accurate.

[0177] Semantic identifier assignment: ;in, represents cluster 3; It is the semantic identifier corresponding to Cluster 3, used to simplify the process modeling representation;

[0178] Cluster 4: Drug treatment event activities;

[0179] Including event activities:

[0180] Drug injection_antibiotics (frequency: 3000 times);

[0181] Oral drug_analgesics (frequency: 2600 times);

[0182] Intravenous infusion_nutritional support (frequency: 2200 times);

[0183] Semantic analysis: The event vectors highly overlap in the "treatment method" dimension (principal component analysis shows that the variance contribution rate > 85%), proving accurate clustering.

[0184] Semantic identifier assignment: ; Among them, is Cluster 4; is the semantic identifier corresponding to Cluster 4, used to simplify the process modeling representation;

[0185] Cluster 5: Nursing operation event activities;

[0186] Including event activities:

[0187] Nursing record_daily (frequency: 5500 times);

[0188] Nursing assessment_night (frequency: 4800 times);

[0189] Bed adjustment_assistance (frequency: 3700 times);

[0190] Semantic analysis: Through t-SNE visualization, these event activities form independent clusters in the two-dimensional projection, proving accurate clustering.

[0191] Semantic identifier assignment: ; Among them, is Cluster 5; is the semantic identifier corresponding to Cluster 5, used to simplify the process modeling representation;

[0192] Cluster 6:

[0193] Including event activities:

[0194] Temporary equipment detection_1 (frequency: 2 times);

[0195] Disinfection record_special (frequency: 1 time);

[0196] The number of members in Cluster 6 is 2, which is less than the first threshold of 3, and the minimum Euclidean distance from other clusters in the embedding space is 1.45, which is greater than the second threshold of 1.2. Therefore, Cluster 6 is a noise cluster and is directly removed;

[0197] Generate a redundant event log according to the process in Step 3, aiming to compress the log scale and retain key semantics.

[0198] Table 1 shows a specific comparison example of a certain trajectory conversion. 36 original events are compressed into 5 semantic identifiers (ID_Vital, ID_Blood, ID_Medication, ID_Imaging, ID_Nursing) and 11 non-repetitive sequences (such as "emergency surgery").

[0199] Semantic consistency: Event activities within the same cluster belong to the same type of operation in actual business. Specifically: All variants of "vital sign monitoring" are classified into ID_Vital; all variants of "blood test" are classified into ID_Blood; all variants of "imaging examination" are classified into ID_Imaging; all variants of "medication treatment" are classified into ID_Medication; all variants of "nursing operation" are classified into ID_Nursing.

[0200] Interpretability: The identifier naming directly reflects the business semantics (such as ID_Blood represents all operations related to blood tests).

[0201] Table 1 Specific comparison example of a certain trajectory conversion

[0202] 。

[0203] Perform trajectory importance sampling according to the process in Step 4, aiming to select representative trajectories to ensure coverage of key behaviors. First, identify key low-frequency event activities; the third threshold is preset to 31. In this embodiment, the frequency of the event activity "emergency surgery" = 32 times, and the frequency of "transfer to ICU" = 45 times, both of which are greater than the low-frequency threshold. Therefore, it is determined that "emergency surgery" and "transfer to ICU" are key low-frequency event activities. Then set the weight of low-frequency trajectories , to emphasize the role of activity importance in comprehensive evaluation. For ordinary trajectories, set , calculate the comprehensive importance score of each trajectory according to formula (10); then sort the trajectories in descending order of the comprehensive importance score, and select the top 25% of the trajectories (3272) as the sample log . For comparison with traditional methods, key path retention is performed here. After calculation, the retention rate of the trajectories containing "emergency surgery" in the present invention is 98%, while that of the traditional method LogRank is only 72%. Therefore, the effect of the present invention is better.

[0204] The quality of the model is evaluated and verified according to the process in step 5. First, the process model is generated using the InductiveMiner algorithm, which is implemented based on the ProM plug-in. The process model includes Petri nets and critical paths; the Petri net contains 45 places and 85 transitions; the critical path contains several, such as "initial diagnosis → emergency surgery → ICU transfer". Then, according to formula (12), the fit is calculated to be 0.94. The fit of 0.94 means that the sample model can reproduce 94% of the original behavior, exceeding the threshold of 0.9. Finally, the complexity is evaluated. The edge of the process model is the number of transitions, 85, and the node of the process model is the number of places, 45. According to formula (14), the edge-node ratio is approximately equal to 1.89, and the edge-node ratio is 1.89<3, indicating that the model is simple and not overfitted.

[0205] The method of the present invention was compared with the traditional method (LogRank) and the experimental results are shown in Table 2.

[0206] Table 2 Comparison results between the method of the present invention and the traditional method

[0207] .

[0208] As can be seen from Table 2, the number of activities in the present invention is reduced by 39%, the model complexity is reduced by 45%, and the de-redundancy effect is significant; the present invention retains key behaviors: the retention rate of low-frequency key event activities (such as "emergency surgery") is increased by 26%. The present invention improves computing efficiency, and Spark parallelization reduces the clustering and sampling time to 1 / 3 of the traditional method. Among them, ↓68% means that the process discovery time of the present invention method is reduced by 68% compared with the original event log.

[0209] The present invention can also perform intelligent manufacturing log optimization;

[0210] Scenario: Compress device detection event activities (such as "Sensor calibration_1" "Sensor calibration_2").

[0211] Dynamic weight adjustment: adjusted according to the equipment failure rate , giving priority to covering fault-related paths.

[0212] In this embodiment, the event embedding model is used to map semantically similar event activities (such as "Blood Test_Type 1" and "Blood Test_Type 2") into a low-dimensional space and perform clustering to generate a unified semantic identifier (such as "ID_Blood"), significantly compressing the log scale; based on the dynamic weighted trajectory importance sampling method, key low-frequency event activities (such as "Emergency Surgery") and their direct following relationships are preferentially retained to ensure the behavioral completeness of the sample event log; finally, the process model is generated through the Inductive Miner algorithm, with a fitness of 0.94 and an edge-node ratio of 1.89, verifying the high quality and simplicity of the model. Experiments show that the method of the present invention reduces the process discovery time by 68% on the BPIC2012 log, and the key path retention rate is increased to 98%, which is applicable to the processing of large-scale logs in fields such as healthcare.

[0213] Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Changes, modifications, additions, or substitutions made by those skilled in the art within the scope of the essence of the present invention should also fall within the protection scope of the present invention.

Claims

1. A method for mining activity sequence sampling process of medical data, characterized in that, It includes the following steps: Step 1: Obtain the original event log and perform data cleaning; Step 2: Based on the event embedding model and the DBSCAN algorithm, cluster event activities with similar semantics into the same cluster, and generate a unique semantic identifier for each cluster; Step 3: Replace event activities with similar semantics in the event log with the semantic identifier corresponding to the cluster to generate a redundant-free event log; Step 4: Sample the redundant-free event log based on the comprehensive importance score of the trace to obtain a sample event log; Step 5: Input the sample event log into an inductive mining algorithm to generate a corresponding process model, and evaluate the quality and complexity of the process model; The specific process of Step 1 is as follows: Step 1.1: Export the standard event log from the hospital information system of the medical institution as the original event log. The format of the event log is a structured file conforming to the XES specification, including the event activity name, timestamp, case ID, and the participant who executes the activity; The event log is a set containing several traces, and a trace is a sequence composed of several event activities; Define the original event log as: (1); Among them, is the original event log; is the total number of trajectories; is the th trajectory, specifically: (2); Among them, is the th event activity in the th trajectory; Step 1.2: Perform data cleaning on the original event log, remove traces that do not contain complete lifecycle information, standardize the labels of event activities, and unify the labels of event activities; The specific process of Step 2 is as follows: Step 2.

1. Train an event embedding model; for each trace in the cleaned event log, use a sliding window method with a length of to form context pairs , and all the context pairs constitute a training sample set for training the event embedding model; where represents the target event activity, and represents the context event activity that co-occurs with the target event activity within the sliding window. During training, the set parameters include: vector dimension, negative sampling number, learning rate; Set the objective function for model training as follows: (3); Among them, is an event embedding model; represents a vector inner product operation; is a negative sampling candidate event activity; is the complete set of event activities; When the number of training iterations reaches the preset number, the training ends; Step 2.2: Use the trained event embedding model to vectorize the event activities to obtain an embedding sequence; Step 2.3: Apply the DBSCAN algorithm to cluster the embedding sequence; The DBSCAN algorithm is a density-based clustering algorithm. By calculating the density reachability between the embedding sequences, it will automatically divide the density-connected embedding sequences into the same cluster, thereby identifying event activities with similar semantics; During the clustering process, distributed computing is adopted, and the embedding sequence is sliced into 6 Spark nodes for parallel computing; The final clustering result contains several clusters, and each cluster contains several event activities with similar semantics; The DBSCAN algorithm parameters include: neighborhood radius, minimum sample number; Step 2.4: Preset a first threshold and a second threshold, count the number of members in each cluster and compare it with the first threshold, calculate the minimum Euclidean distance between each cluster and other clusters in the embedding space, and compare it with the second threshold; if the number of members in the cluster is less than the first threshold and its minimum Euclidean distance from other clusters in the embedding space is greater than the second threshold, then the cluster is an isolated noise cluster; assign a unique semantic identifier to each cluster other than the noise cluster, in the format of: ; is the semantic identifier corresponding to the th cluster; The specific process of Step 2.2 is as follows: Step 2.2.1: Construct a semantic representation space for activities based on the trained event embedding model; (4); Among them, is the semantic representation space of the event activity ; represents the embedded vector representation of the event activity ; represents the dimension of the embedded vector Step 2.2.2: Each trajectory extracts subsequences of a fixed length in a sliding window manner, denoted as embedding sequences: ​ (5); Among them, is the embedding sequence, representing the representation sequence of the subsequences extracted by the sliding window in the embedding space; represents the th event activity in the embedding sequence, corresponding to the embedding vector representation; The specific process of Step 3 is as follows: Replace event activities with similar semantics in the event log with the semantic identifier corresponding to the cluster, in the form of: (6); Thus, a redundant-free event log with a simplified structure is formed: (7); Among them, is an update operation; is to remove redundant event logs; , are respectively the th and the th event activities in the th trajectory; represents the th event activity in the th trajectory, and this event activity is a semantically similar event activity; is the th trajectory after clustering; represents the th event activity in the th trajectory after clustering, and this event activity is the semantic identifier corresponding to the cluster; The specific process of Step 4 is as follows: Step 4.1: Preset a third threshold. If the occurrence frequency of an event activity is less than the third threshold and it is confirmed by medical experts or the domain knowledge base to have key business significance, then mark this event activity as a key low-frequency event activity, mark the trace where the key low-frequency event activity is located as a low-frequency trace, and mark the remaining traces as ordinary traces; Step 4.2: Calculate the activity importance and structural importance of each trace, and calculate the comprehensive importance score by weighting; The calculation formula for activity importance is: (8); Among them, is the activity importance of the th trajectory; is the coverage rate of the event activity in the redundant event log; The calculation formula for structural importance is: (9); Among them, is the structural importance of the th trajectory; , are respectively the th and th events in the th trajectory; represents the direct following frequency of the event on in the redundant event log. The comprehensive importance score of the trajectory is calculated by linear weighting: (10); Among them, is the comprehensive importance score of the th trajectory; is the weight; When calculating the comprehensive importance score of low-frequency trajectories, set the weight coefficient , for ordinary trajectories, set ; Step 4.

2. Sort all the trajectories according to the comprehensive importance score, and select the top trajectories with the highest comprehensive importance scores as the sample event log: (11); Among them, is the sample event log; is all trajectories with the top trajectories having the highest comprehensive importance scores; is the number of target trajectories at the preset sampling rate; is the sampling rate, and the sampling rate supports dynamic adjustment by the user according to computing resources, and the adjustment range is 20% - 30%.

2. The method for mining the activity sequence sampling process applied to medical data according to claim 1, wherein In the said step 5, the process model includes a Petri net, and the Petri net contains places and transitions; The quality of the process model is evaluated through a goodness-of-fit index, and the calculation formula of the goodness-of-fit index is: (12); Among them, is the goodness-of-fit index; is the process model; represents length calculation; represents the trajectory in the replay path, and the calculation formula is: (13); Among them, is the transition sequence in the process model, is the th transition; is the set composed of all transitions in the process model; is the th event activity and the th transition matching function; is the total number of event activities; By calculating the edge-node ratio of the process model to evaluate the complexity of the process model, the formula is: (14); Among them, is the edge-node ratio of the process model ; is the edge of the process model , corresponding to the number of transitions; is the node of the process model , corresponding to the number of places; If > 3, it is determined that the process model is overfitted and it is necessary to return to step 1 for reprocessing; During online application, the sample event log is input into the inductive mining algorithm to generate the corresponding process model, thus completing the process mining task.

3. An activity sequence sampling process mining system applied to medical data, characterized in that, The process mining is carried out by using the activity sequence sampling process mining method applied to medical data as described in any one of claims 1-2. The system includes a data collection and cleaning module, a clustering module, a redundancy removal module, a sampling module, and a process model generation module; The data collection and cleaning module is connected to the hospital information system of the medical institution to obtain real-time data, and cleans the obtained standard event log; the clustering module clusters the semantically similar event activities into the same cluster based on the event embedding model and the DBSCAN algorithm, and generates a unique semantic identifier for each cluster; the redundancy removal module is used to replace the semantically similar event activities in the event log with the semantic identifier corresponding to the cluster to generate a redundancy-removed event log; the sampling module samples the redundancy-removed event log based on the comprehensive importance score of the trajectory to obtain a sample event log; the process model generation module is used to input the sample event log into the inductive mining algorithm to generate the corresponding process model, and evaluate the quality and complexity of the process model.

Citation Information

Patent Citations

  • Automatic software process modeling method and system based on process mining

    CN115374595A

  • Case feature tag-based medical path precise classification model generation method and system

    CN115910359A

  • Supervised learning-based data identification method and system

    CN119537899A

  • Business risk assessment method and device, computer equipment and storage medium

    CN119721689A