Visual Modeling Method and System Based on Process Mining
Through the visual modeling method based on process mining, the existing process modeling methods have been solved, and the problem of poor flexibility, poor interpretation and lack of visual interface are achieved, and an intuitive, flexible and efficient process modeling process is achieved.
Patent Information
- Application Number
- CN202510405458.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-04-02
AI Technical Summary
The existing process modeling methods have problems such as poor flexibility, poor model interpretability, complex training process and lack of visual interface.
Using a visual modeling method based on process mining, we collect process data from log files, clean and preprocess data, provide a graphical interface for visual modeling, and allow users to create and modify data processing models through drag and drop operations.
It realizes an intuitive, easy-to-understand, flexible and efficient process modeling process, improving user experience and work efficiency.
Smart Images

Figure CN119917086B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of process modeling, and particularly relates to a visualization modeling method and system based on process mining. Background Art
[0002] In traditional process modeling, the method relying on manual input is not only time-consuming and laborious, but also prone to errors. Although existing process mining tools can extract process models from log data, they often lack an intuitive visualization interface, making it difficult for non-professional users to understand and modify the models. In addition, although some visualization tools provide a graphical interface, they fail to be closely integrated with process mining technology and cannot achieve an automated modeling process. Existing solutions have deficiencies in terms of user-friendliness, flexibility, automation level, and processing speed, specifically manifested as follows:
[0003] 1. Rule-based automated modeling: Although it has a high degree of automation and fast processing speed, rule definition and maintenance are complex and lack flexibility.
[0004] 2. Machine learning-assisted modeling: It can handle complex data, but requires a large amount of training data, has poor model interpretability, and a complex training process.
[0005] 3. Traditional process mining tools: They are technically mature and have a fast processing speed, but lack a visualization interface and require a high technical level.
[0006] Based on the above problems, it is very important to design an intuitive, flexible, and automated visualization modeling method and system based on process mining. Summary of the Invention
[0007] The present invention aims to overcome the problems in the existing process modeling methods, such as poor flexibility, poor model interpretability, complex training process, and lack of a visualization interface, and provides an intuitive, flexible, and automated visualization modeling method and system based on process mining.
[0008] To achieve the above invention purpose, the present invention adopts the following technical solutions:
[0009] A visualization modeling method based on process mining includes the following steps;
[0010] S1, collecting process data from a log file;
[0011] S2, cleaning and preprocessing the collected process data;
[0012] S3, using the process data processed in step S2 for visualization modeling; the visualization modeling provides a graphical interface, and users can create and modify a data processing model through drag-and-drop and editing operations;
[0013] S4. The user exports the constructed data processing model as a file for subsequent use or sharing.
[0014] Step S3 includes the following steps:
[0015] S31. Select a data source or reference an existing model: The user selects the required data source or references an existing model saved previously through a graphical interface to provide a basic data set.
[0016] S32. Data integration: Based on the data source or existing model selected in step S31, a data node is generated, and the user performs subsequent operations based on the data node to form a data set in series; the subsequent operations include field management, table association, and table merging.
[0017] S33. Operation node: Using the data set integrated in step S32, the user performs specific data processing tasks in the operation node step to obtain a customized result table, which will be used for final data analysis and decision support; the data processing tasks include data filtering, sorting, and aggregation.
[0018] Preferably, in step S2, cleaning the collected process data includes missing value processing, outlier processing, duplicate value processing, and error data processing.
[0019] The missing value processing refers to deleting, filling, or predicting missing values.
[0020] The outlier processing refers to identifying and processing outliers; the processing methods include deleting, modifying, and marking.
[0021] The duplicate value processing refers to deleting duplicate records.
[0022] The error data processing refers to correcting errors in the data set.
[0023] Preferably, in step S2, preprocessing the collected process data includes data integration, data transformation, data normalization, data dimensionality reduction, and data augmentation.
[0024] The data integration refers to merging multiple data sources.
[0025] The data transformation refers to performing logarithmic transformation, normalization, and discretization operations.
[0026] The data normalization refers to unifying the data format and unit.
[0027] The data dimensionality reduction refers to reducing the data dimension through aggregation, reduction, and compression methods.
[0028] The data augmentation refers to enriching the data set by adding additional information.
[0029] Preferably, in step S32, the field management refers to that the user manages the fields in the data nodes, including addition, deletion, and renaming operations.
[0030] Preferably, in step S32, the table association refers to associating multiple data nodes to form a new data set, specifically including merging two tables through a common field.
[0031] Preferably, in step S32, the table merging refers to merging multiple data nodes into a single data set; the merging methods include vertical merging or horizontal merging.
[0032] Preferably, in step S32, the data set has the following characteristics:
[0033] Comprehensiveness, specifically, the comprehensiveness refers to merging data sets from different sources through field management, table association, and table merging operations to cover all relevant information required for user analysis;
[0034] Consistency, specifically, the consistency refers to ensuring the unity of field names, formats, and units through field management.
[0035] The present invention also provides a visualization modeling system based on process mining, including:
[0036] An acquisition module for acquiring process data from a log file;
[0037] A preprocessing module for cleaning and preprocessing the acquired process data;
[0038] A visualization modeling module for performing visualization modeling using the process data processed by the preprocessing module; the visualization modeling provides a graphical interface, and the user can create and modify a data processing model through dragging and editing operations;
[0039] An export module for exporting the constructed data processing model as a file.
[0040] Compared with the prior art, the beneficial effects of the present invention are: (1) Intuitive and easy to understand: Through the graphical interface, the user can more intuitively understand each step of data processing; (2) Flexible and efficient: The present invention enables the user to easily adjust the data processing flow with low trial-and-error costs; (3) Strong reusability: The constructed model of the present invention can be reused multiple times to improve work efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 It is a flowchart of a visualization modeling method based on process mining in the present invention;
[0042] Figure 2A schematic diagram of an interface for selecting a data source or referencing an existing model process in the present invention;
[0043] Figure 3 A schematic diagram of an interface for the data integration process in the present invention;
[0044] Figure 4 A schematic diagram of an interface for an operation node in the present invention;
[0045] Figure 5 A schematic diagram of an interface for saving and exporting a data processing model in the present invention;
[0046] Figure 6 A schematic diagram of a modeling process in the practical application of the visualization modeling method based on process mining provided by an embodiment of the present invention. Detailed implementation manners
[0047] To more clearly illustrate the embodiments of the present invention, the specific implementation manners of the present invention will be described below with reference to the accompanying drawings. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings, and other implementation manners can also be obtained.
[0048] As Figure 1 shown, the visualization modeling method based on process mining of the present invention includes the following steps;
[0049] 1. Collect process data from a log file;
[0050] 2. Clean and preprocess the collected process data to ensure the quality and consistency of the data;
[0051] 3. Use the process data processed in step 2 to perform visualization modeling; the visualization modeling provides a graphical interface, and the user can create and modify a data processing model through drag-and-drop and editing operations;
[0052] 4. The user exports the constructed data processing model as a file for subsequent use or sharing.
[0053] Further, in step 2, cleaning the collected process data includes handling missing values, handling outliers, handling duplicate values, and handling error data;
[0054] Handling missing values means deleting, filling, or predicting missing values to ensure the integrity of the data;
[0055] Handling outliers means identifying and handling outliers to avoid adverse effects on the analysis results;
[0056] Among them, the handling methods include but are not limited to:
[0057] Deletion: Delete abnormal data records that seriously deviate from the normal range and have no business value;
[0058] Modification: Reasonably replace or correct outliers according to business rules or algorithms (such as median interpolation, neighboring value filling, etc.) to ensure data integrity and accuracy;
[0059] Marking: Add special marks to outliers that may have analytical value for subsequent analysis;
[0060] Duplicate value handling refers to deleting duplicate records to improve data uniqueness and accuracy;
[0061] Error data handling refers to correcting errors in the dataset to enhance data reliability.
[0062] Preprocessing the collected process data includes data integration, data transformation, data normalization, data dimensionality reduction, and data augmentation;
[0063] Data integration refers to merging multiple data sources to solve entity recognition and attribute redundancy problems;
[0064] Data transformation refers to performing logarithmic transformation, normalization, and discretization operations to eliminate data inconsistencies;
[0065] Data normalization refers to unifying data formats and units to ensure data consistency;
[0066] Data dimensionality reduction refers to reducing the data dimension through aggregation, reduction, and compression methods to improve analysis efficiency;
[0067] Data augmentation refers to enriching the dataset by adding additional information, such as adding geographical location information through an external database.
[0068] Furthermore, in the visualization modeling stage, the specific steps are as follows:
[0069] 3-1. Select a data source or reference an existing model: As Figure 2 shown, the user first selects the required data source through the graphical user interface or references an existing model saved previously, which provides the basic dataset for subsequent data integration.
[0070] 3-2. Data integration: As Figure 3As shown, based on the data source or existing model selected in step 3-1, a data node will be generated. Users can perform subsequent operations based on this node, including field management, table association, and table merging, etc., thus forming a comprehensive and consistent dataset in series, laying a solid foundation for further analysis and modeling. The "comprehensiveness" of the dataset is reflected in that by performing field management, table association, and table merging operations, datasets from different sources are combined to cover all relevant information required for user analysis. For example, table association can connect multiple tables through common fields to ensure the integrity of data content; table merging can expand the dimensions and depth of the dataset through vertical or horizontal merging. The "consistency" of the dataset is reflected in the standardization and unity of the data. By field management, the consistency of field names, formats, and units is ensured, avoiding the occurrence of duplicate and redundant data, thereby improving data quality and ensuring the accuracy and usability of the dataset.
[0071] 3-3. Operation Node: As Figure 4 shown, using the dataset integrated in step 3-2, users can perform specific data processing tasks in the operation node, such as data filtering, sorting, aggregation, etc., to obtain a customized result table, which will be used for final data analysis and decision support.
[0072] Furthermore, as Figure 5 shown, for step 4, users can export the constructed data processing model as a file, such as in JSON or XML format, for convenient subsequent use or sharing.
[0073] In addition, the present invention also provides a visualization modeling system based on process mining, including:
[0074] A collection module for collecting process data from log files;
[0075] A preprocessing module for cleaning and preprocessing the collected process data;
[0076] A visualization modeling module for performing visualization modeling using the process data processed by the preprocessing module; the visualization modeling provides a graphical interface, and users can perform creation and modification of the data processing model through dragging and editing operations;
[0077] An export module for exporting the constructed data processing model as a file.
[0078] The following combines specific example scenarios, with the Figure 6 process shown, to detail the specific implementation steps of the present invention in practical applications:
[0079] Suppose there is an e-commerce dataset containing an order table and a user table. Visualization modeling is used to analyze users' purchase behaviors.
[0080] Step 1: Select a data source or reference an existing model (corresponding to the data node 1 in Figure 6 )
[0081] In this step, the order table and user table in the e-commerce dataset are selected as the data sources. The order table contains detailed information about orders, such as order ID, user ID, order amount, purchased items, etc.; the user table contains basic information about users, such as user ID, name, age, gender, region, etc.
[0082] Step 2: Data integration (corresponding to the data node 2 in Figure 6 )
[0083] In the data integration step, the order table and user table are associated through the "user ID" field to generate a new dataset. The specific content of this new dataset includes the following:
[0084] Order details: including order ID, order amount, details of purchased items, order date, etc. for each order.
[0085] User basic information: including name, age, gender, region, etc. of the user associated with the order.
[0086] Association information: showing the relationship between each user and their orders, such as the types, frequencies, and consumption habits of the items purchased by the user.
[0087] Through this step, a comprehensive data view can be obtained, which directly associates the personal characteristics of users with their purchase behaviors, providing a rich information basis for further analysis.
[0088] Step 3: Operation node (corresponding to the data node 3 in Figure 6 )
[0089] In the operation node step, the generated new dataset is further processed. For example, aggregation operations can be added to calculate the total consumption amount, average consumption amount, purchase times, etc. for each user. In addition, users can be grouped according to their age and region to analyze the differences in purchase behaviors among users of different age groups and regions.
[0090] Step 4: Set as the result table (corresponding to the final model in Figure 6 )
[0091] In the set as the result table step, the merged dataset is set as the result table. The specific content of this result table includes the following:
[0092] User purchase behavior statistics: including statistical data such as the total consumption amount, average consumption amount, purchase times, etc. for each user.
[0093] User grouping analysis: Group users according to their age and region, and display the purchase behavior characteristics of different groups.
[0094] Commodity sales analysis: Analyze which commodities are the most popular, which commodities have the highest sales volume, and the relationship between commodity sales and user characteristics.
[0095] Trend prediction: Based on historical purchase data, predict future sales trends and changes in user purchase behavior.
[0096] This result table provides in-depth analysis of user purchase behavior for e-commerce enterprises, helping enterprises optimize marketing strategies and improve sales efficiency. For example, by analyzing user purchase behavior, enterprises can find that users in certain regions are more inclined to purchase high-value commodities, so they can target the promotion of high-end products in these regions.
[0097] Through this process-based processing, the present invention not only improves the efficiency and accuracy of modeling, but also enhances the reusability of the model and the shareability of the results. Each step is a logical continuation of the previous step, ensuring the coherence and systematicness of the entire modeling process.
[0098] The above description only details the preferred embodiments and principles of the present invention. For those of ordinary skill in the art, according to the idea provided by the present invention, there will be changes in the specific implementation manners, and these changes should also be regarded as the protection scope of the present invention.
Claims
1. A visual modeling method based on process mining, characterized in that: The method comprises the following steps: S1, collects process data from log files; S2, cleaning and preprocessing the collected process data; S3, using the process data processed in step S2 to perform visual modeling; the visual modeling provides a graphical interface, and the user can create and modify the data processing model through dragging and editing operations; S4, the user exports the constructed data processing model as a file for subsequent use or sharing; Step S3 includes the following steps: S31, select data source or reference existing model: the user selects the required data source or references the previously saved existing model through the graphical interface to provide the basic data set; S32, data integration: according to the data source or existing model selected in step S31, a data node is generated, and the user performs subsequent operations based on the data node to form a data set in series; the subsequent operations include field management, table association and table merging; S33, operation node: using the data set integrated in step S32, the user performs a specific data processing task in the operation node step to obtain a customized result table, which will be used for final data analysis and decision support; the data processing task includes data filtering, sorting and aggregation; In step S32, the field management refers to the user managing the fields in the data node, including adding, deleting and renaming operations; In step S32, the table association refers to associating multiple data nodes to form a new data set, specifically including merging two tables through a common field; In step S32, the table merging refers to merging multiple data nodes into a single data set; the merging method includes vertical merging or horizontal merging; In step S32, the data set has the following characteristics: Comprehensiveness: Specifically, comprehensiveness refers to merging data sets from different sources through field management, table association, and table merging operations to cover all relevant information required for user analysis; Consistency, specifically refers to ensuring the uniformity of field names, formats, and units through field management.
2. The visual modeling method based on process mining according to claim 1 is characterized in that: In step S2, the collected process data is cleaned including missing value processing, abnormal value processing, duplicate value processing and erroneous data processing; The missing value processing refers to deleting, filling or predicting missing values; The outlier processing refers to identifying and processing outliers; the processing methods include deletion, modification and marking; The duplicate value processing refers to deleting duplicate records; The error data processing refers to correcting errors in the data set.
3. The visual modeling method based on process mining according to claim 2 is characterized in that: In step S2, the collected process data is preprocessed including data integration, data conversion, data normalization, data dimension reduction and data enhancement; The data integration refers to merging multiple data sources; The data conversion refers to performing logarithmic transformation, normalization and discretization operations; The data standardization refers to the unification of data formats and units; The data dimensionality reduction refers to reducing the data dimension by aggregation, reduction and compression methods; The data enhancement refers to adding additional information to enrich the dataset.
4. A visual modeling system based on process mining, used to implement the visual modeling method based on process mining according to any one of claims 1 to 3, characterized in that: The process mining-based visual modeling system includes: The collection module is used to collect process data from log files; The preprocessing module is used to clean and preprocess the collected process data; A visual modeling module is used to perform visual modeling using the process data processed by the preprocessing module; the visual modeling provides a graphical interface, and the user can create and modify the data processing model through dragging and editing operations; The export module is used to export the constructed data processing model as a file.
Citation Information
Patent Citations
Big data mining tool and method based on dragging process
CN110909039A
Process optimization method, electronic equipment, storage medium and program product
CN119558640A