An end-to-end business process bottleneck prediction and simulation method based on process mining

CN122596360APending Publication Date: 2026-08-18GUANGDONG POWER GRID CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611052700.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-15
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

[0003]然而,当前针对业务流程瓶颈识别的方法存在一些深层次的局限性

Benefits of technology

本发明公开了一种基于流程挖掘的端到端业务流程瓶颈预测与仿真方法,针对业务流程中隐含环节导致的时间偏差和瓶颈问题,通过从事件日志中提取时间戳记录和关键路径数据,解析相邻节点间隔时长与服务时间分布,计算时间戳偏差值,识别未记录的人工审核等静默活动占比,进而修正节点服务时间,标记瓶颈候选节点。本发明通过综合风险指数和动态阈值调整,结合支持向量机分类高风险瓶颈节点,并利用逻辑回归算法预测瓶颈覆盖率,最终通过仿真平台验证优化效果。本发明最核心的创新在于将静默活动耗时剥离与动态阈值调整相结合,精准定位业务流程中的隐性瓶颈,提升了流程优化的准确性和预测能力,为企业流程管理提供了科学依据。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122596360A_ABST
    Figure CN122596360A_ABST
Patent Text Reader

Abstract

The application provides an end-to-end business process bottleneck prediction and simulation method based on process mining, comprising: identifying the time proportion of implicit links such as manual review in the path segment which are not recorded by the system through analyzing the timestamp deviation value set, determining the silent activity proportion distribution and the service time expansion range of adjacent nodes; identifying the nodes in the adjacent nodes whose service time expansion range exceeds the preset expansion threshold, obtaining the corrected node service time set by stripping the silent activity time consumption component of each node in the event log; according to the adjusted bottleneck judgment threshold, classifying the time expansion characteristics of each node in the bottleneck candidate node set by using the support vector machine algorithm, and identifying high-risk silent bottleneck nodes; simulating the high-risk silent bottleneck nodes, obtaining path simulation data in the simulation platform, calculating the coverage rate improvement range by using the logistic regression algorithm, and predicting the bottleneck coverage rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information technology, and in particular to an end-to-end business process bottleneck prediction and simulation method based on process mining. Background Technology

[0002] In the field of enterprise operations management, researching how to optimize business processes to improve efficiency and reduce costs is of paramount importance. The analysis and improvement of business processes directly impact a company's core competitiveness. In particular, identifying bottlenecks in complex processes and implementing targeted measures can significantly improve overall operational effectiveness. This area of ​​research is not only a hot topic in academia but also a pressing need in business practice.

[0003] However, current methods for identifying bottlenecks in business processes have some deep-seated limitations. These methods often rely too heavily on recorded data, neglecting steps in the process that are not captured by the system, leading to discrepancies between the analysis results and actual operations. In particular, unrecorded steps, such as manual review or offline communication, although consuming significant time in actual operation, are often overlooked due to a lack of data support, thus affecting the comprehensiveness and accuracy of bottleneck identification. These unrecorded steps often cause the time of adjacent recorded steps to be erroneously amplified, creating a bias in time distribution. For example, in an order processing flow, the system records the time for order submission and order confirmation, but the intermediate manual review step is not recorded. Its time is incorrectly attributed to the order confirmation step, making the confirmation step appear excessively time-consuming and misjudged as a bottleneck, while the truly time-consuming steps are completely ignored.

[0004] This time distribution bias not only masks the real problem but can also mislead the direction of optimization. Therefore, accurately identifying the impact of unrecorded steps in the process on the time distribution and correcting the resulting misjudgments of bottlenecks has become a key issue in business process optimization. Summary of the Invention

[0005] This invention provides a method for predicting and simulating business process bottlenecks based on process mining, including: Read the timestamp records of each node and the path segment data on the critical path from the event log. Extract the interval between adjacent nodes and the service time distribution of each node through the log parsing unit. Perform a difference operation on the theoretical completion timestamp within the interval between adjacent nodes and the service time distribution to obtain a set of timestamp deviation values. By analyzing the set of timestamp deviation values, the time proportion of hidden business processes that are not recorded within the path segment is identified, the distribution of silent activities is determined, and the service time inflation of adjacent nodes is extracted. If the service time inflation rate of a node among the adjacent nodes exceeds a preset inflation threshold, the silent activity time component of each node in the event log is removed, and the corrected node service time set is extracted. Extract the expansion limit range of each node on the critical path from the modified node service time set, evaluate the time consumption weight of each path segment in combination with the silent activity proportion distribution, determine the nodes whose expansion limit range and time consumption weight both exceed the corresponding average, and aggregate to generate a bottleneck candidate node set. Analyze the bottleneck candidate node set, obtain the fluctuation range of the silent activity ratio of the path segment where each candidate node is located, perform a weighted fusion operation based on the fluctuation range of the ratio and the time consumption weight to generate a comprehensive risk index, extract the distribution range of the comprehensive risk index and perform dynamic scaling on the preset judgment benchmark threshold to generate an adjustment judgment threshold; Based on the adjusted judgment threshold, the time dilation features of each node in the bottleneck candidate node set are classified to identify high-risk bottleneck nodes. Simulation processing is performed on the high-risk bottleneck nodes to obtain path simulation data, and the overall bottleneck coverage is predicted based on the coverage improvement rate calculated by logistic regression.

[0006] Furthermore, the process involves reading timestamp records of each node and path segment data on the critical path from the event log, extracting the interval between adjacent nodes and the service time distribution of each node through the log parsing unit, and performing a difference calculation on the service time distribution and the theoretical completion timestamps within the interval between adjacent nodes to obtain a set of timestamp deviation values, including: Identify the start and end nodes of each path segment on the critical path; Extract the node service time based on the end timestamp minus the start timestamp of the process instance; Based on the node service time of the multi-process instance, the mean and variance of service time are calculated, and the service time distribution is established. The average service time between adjacent nodes is summed to generate a theoretical completion timestamp; Calculate the difference between the theoretical completion timestamp and the actual completion timestamp to generate a set of timestamp deviation values.

[0007] Furthermore, the step of identifying the time proportion of unrecorded hidden business processes within the path segment by analyzing the timestamp deviation value set, determining the distribution of silent activity proportions, and extracting the service time inflation of adjacent nodes includes: The intervals where the deviation values ​​are densely clustered in the set of timestamp deviation values ​​are extracted as the regions where hidden links exist; The percentage of silent activity time is generated based on the ratio of the cumulative duration of deviation values ​​within the area where the hidden link exists to the total duration of the path segment. Perform a sequence permutation on the percentage of silent activity time in multiple groups to construct a distribution of silent activity percentage; Calculate the difference between the original service time and the baseline service time to extract the service time inflation. The ratio of the service time inflation to the baseline service time is calculated to generate the service time inflation magnitude.

[0008] Furthermore, for the nodes whose service time inflation exceeds a preset inflation threshold among the adjacent nodes, the silent activity time component of each node in the event log is stripped, and the corrected node service time set is extracted, including: Nodes whose service time expansion exceeds a preset expansion threshold are selected from the adjacent nodes as nodes that have exceeded the expansion limit. Obtain the ratio of the original service time to the silent activity time associated with the node that exceeds the expansion limit; The silent activity time component is obtained by multiplying the silent activity time percentage by the original service time. The stripped service time is generated by subtracting the silent activity time component from the original service time; The service time after stripping is combined with the original service time of nodes that have not exceeded the limit to form the corrected set of node service times.

[0009] Furthermore, the step of extracting the expansion limit range of each node on the critical path from the corrected node service time set, evaluating the time consumption weight of each path segment in conjunction with the silent activity proportion distribution, determining nodes whose expansion limit range and time consumption weight both exceed the corresponding average, and aggregating to generate a bottleneck candidate node set includes: Identify the expansion amplitude of each node and subtract it from the preset expansion threshold to obtain the expansion excess range; The arithmetic mean of the expansion excess range is calculated to obtain the mean value of the expansion excess range; The time consumption weight is calculated by multiplying the percentage of time spent in silent activities by the percentage of total time spent in silent activities. The arithmetic mean of the time consumption weights is calculated to obtain the mean of the time consumption weights; A bottleneck candidate node set is established by combining nodes that meet the criteria of having an expansion limit greater than the average of the expansion limit range and a time consumption weight greater than the average of the time consumption weight.

[0010] Furthermore, the analysis of the bottleneck candidate node set obtains the fluctuation range of the silent activity proportion of the path segment where each candidate node is located. Based on the fluctuation range of the proportion and the time consumption weight, a weighted fusion operation is performed to generate a comprehensive risk index. The distribution range of the comprehensive risk index is extracted and dynamically scaled against a preset judgment benchmark threshold to generate an adjusted judgment threshold, including: Extract the maximum and minimum percentages of silent activity time among multiple process instances within the same path segment; The difference between the extracted maximum and minimum values ​​is used to generate the percentage fluctuation range of each path segment; After assigning independent preset weight values ​​to the aforementioned percentage fluctuation range and the time consumption weight, a summation process is performed to generate a comprehensive risk index. The comprehensive risk index of all path segments is retrieved to obtain the global maximum and global minimum values; The difference between the global maximum value and the global minimum value is used to calculate the span of the distribution interval; The numerical scaling factor is obtained by dividing the distribution interval span by a preset judgment benchmark threshold. In response to a determination result where the span of the distribution interval is greater than a preset determination benchmark threshold, an adjustment determination threshold is obtained by multiplying the preset determination benchmark threshold by the numerical scaling factor.

[0011] Furthermore, the step of classifying the time dilation features of each node in the bottleneck candidate node set according to the adjusted judgment threshold to identify high-risk bottleneck nodes includes: Obtain the service time inflation rate, inflation limit range, and comprehensive risk index associated with each candidate node; The feature vectors of each candidate node are constructed by combining the three feature values ​​obtained. In response to the condition that the comprehensive risk index is greater than the adjustment judgment threshold, the feature vector is labeled with a high-risk category to construct labeled samples; For the labeled samples input into the classification model, perform maximum margin boundary training to extract the classifier; The classifier performs pattern matching on the feature vector of the node to be detected to read the corresponding category data and identify high-risk bottleneck nodes.

[0012] Furthermore, the step of performing simulation processing on the high-risk bottleneck nodes to obtain path simulation data, and predicting the overall bottleneck coverage based on logistic regression to calculate the coverage improvement, includes: A simulation testing process was established for the aforementioned high-risk bottleneck nodes; The test service time corresponding to each node is read by simulating the node sequence of the critical path and the corrected node service time set through the simulation test process. Set up multiple sets of silent activity time percentage parameters with numerical differences to drive the test process and output independent path simulation data; Read the simulation service time of each node and calculate the difference between it and the original corrected service time to establish the service time change; Read the total time of the simulated path and calculate the difference between it and the total time of the original path to establish the change in path time; Combine the service time change and the path time change to construct coverage feature association data; The coverage feature-related data is input into a logistic regression analyzer to perform maximum likelihood calculation and extract the independent coverage probability value of each node; The difference between the independent coverage probability value and the original baseline coverage probability value is calculated to determine the coverage improvement. The overall bottleneck coverage rate is generated by summing up the coverage improvement of all nodes.

[0013] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects: This invention discloses an end-to-end business process bottleneck prediction and simulation method based on process mining. Addressing time deviations and bottlenecks caused by hidden steps in business processes, it extracts timestamp records and critical path data from event logs, analyzes the interval duration between adjacent nodes and service time distribution, calculates timestamp deviation values, identifies the proportion of unrecorded silent activities such as manual reviews, and then corrects node service times to mark candidate bottleneck nodes. This invention combines a comprehensive risk index and dynamic threshold adjustment with support vector machine classification of high-risk bottleneck nodes and uses logistic regression to predict bottleneck coverage. Finally, the optimization effect is verified through a simulation platform. The core innovation of this invention lies in combining silent activity time stripping with dynamic threshold adjustment to accurately locate hidden bottlenecks in business processes, improving the accuracy and predictive ability of process optimization, and providing a scientific basis for enterprise process management. Attached Figure Description

[0014] Fig. 1 This is a flowchart of an end-to-end business process bottleneck prediction and simulation method based on process mining, according to the present invention.

[0015] Fig. 2 This is a schematic diagram of an end-to-end business process bottleneck prediction and simulation method based on process mining according to the present invention.

[0016] Fig. 3 This is another schematic diagram of an end-to-end business process bottleneck prediction and simulation method based on process mining according to the present invention. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0018] like Figs. 1-3 This embodiment of an end-to-end business process bottleneck prediction and simulation method based on process mining may specifically include: S101. Obtain the timestamp records of each node and the path segment data on the critical path from the event log. Extract the interval duration between adjacent nodes and the service time distribution of each node through the log parsing unit. Perform a difference calculation between the service time distribution and the theoretical completion timestamp within the interval duration of adjacent nodes to obtain a set of timestamp deviation values.

[0019] The start and end timestamp records of each node are read from the event log. The log parsing unit arranges these timestamp records according to the execution order of the nodes, identifying the start and end nodes of each path segment on the critical path. For each process instance, the service time of each node within the path segment is calculated. The service time is equal to the end timestamp minus the start timestamp of that node. The log parsing unit statistically aggregates the service times of each node in multiple process instances, calculating the mean service time μ and the variance service time σ², where μ is the arithmetic mean of the service times of that node across all instances, and σ² is the corresponding variance. A service time distribution for each node is generated based on the mean and variance service times. Based on the mean service time of each node in the service time distribution, the mean service time between adjacent nodes is accumulated to obtain the theoretical completion timestamp within the interval between adjacent nodes. The difference between the theoretical completion timestamp and the actual completion timestamp recorded in the event log is calculated to obtain the set of timestamp deviation values ​​for each path segment.

[0020] In one implementation, the event log reading process includes scanning each business process execution record one by one. Each record contains a node identifier, execution time, and process instance number. The log parsing unit groups the records according to the process instance number, and within each group, sorts the timestamp records of each node in ascending order according to the execution time, thereby reconstructing the complete execution trajectory of a single process instance.

[0021] Specifically, the identification of the critical path is based on the dependencies between nodes in the process definition. The log parsing unit traverses the execution trajectory, marks the node sequence on the critical path, and defines two adjacent nodes as a path segment. For each path segment, the completion timestamp of the starting node and the start timestamp of the ending node are extracted, and the difference between the two is the original interval duration of the path segment.

[0022] In one embodiment, the statistical aggregation process of service time is performed on the execution records of the same node across multiple process instances. The log parsing unit collects the service time values ​​of the node across all process instances, calculates the arithmetic mean of these values ​​as the service time mean, and simultaneously calculates the sum of the squares of the differences between each value and the mean, dividing by the number of instances to obtain the service time variance. The service time mean reflects the typical time consumption level of the node, while the service time variance reflects the dispersion of time consumption among different instances; together, they constitute the service time distribution of the node.

[0023] It should be noted that the theoretical completion timestamp is calculated using a node-by-node accumulation method. Starting with the completion timestamp of the initial node of the path segment, the average service time of each subsequent node is added sequentially to obtain the theoretical completion time of each node under ideal conditions. This theoretical completion timestamp represents the time when the node should complete under conditions of no additional delay.

[0024] For example, the theoretical completion timestamp is compared one by one with the actual completion timestamps recorded in the event log. The theoretical completion timestamp is subtracted from the actual completion timestamp to obtain the timestamp deviation value of a single node. The above calculation is repeated for all path segments on the critical path, and the deviation values ​​of each path segment are summarized to form a timestamp deviation value set. This set reflects the distribution of the extent to which the actual execution time in each path segment exceeds the theoretical expectation.

[0025] S102. By analyzing the set of timestamp deviation values, identify the time proportion of hidden links such as manual review that are not recorded by the system within the path segment, and determine the distribution of silent activities and the service time expansion of adjacent nodes.

[0026] Path segments with deviation values ​​exceeding a preset deviation threshold are selected from the timestamp deviation value set. Deviation values ​​within each path segment are grouped by numerical range, and densely clustered intervals are identified as areas with hidden links. The silent activity time percentage for each path segment is obtained by comparing the cumulative duration of deviation values ​​within these hidden link areas to the total duration of the path segment. Based on the numerical differences in the silent activity time percentages across path segments, the silent activity time percentages are arranged sequentially to obtain a silent activity percentage distribution. For path segments with percentage values ​​exceeding a preset percentage threshold, the average service time of adjacent nodes within that path segment across multiple process instances is extracted as the original service time. The baseline service time of adjacent nodes is obtained from historical normal process records. The difference between the original service time and the baseline service time is calculated to obtain the service time inflation of adjacent nodes. The ratio of the service time inflation to the baseline service time is then calculated to obtain the service time inflation magnitude of each adjacent node.

[0027] In one implementation, the filtering process for the timestamp deviation value set is based on a preset deviation threshold, which is determined according to the historical operation data of the business process. A typical value is 1.5 times the average timestamp deviation value of each path segment in the historical normal process.

[0028] For example, if the historical average deviation of a normal process is 2 minutes, then the preset deviation threshold is set to 3 minutes, and the typical value of the preset percentage threshold is 0.2. That is, a path segment is judged as having a high percentage when the silent activity time accounts for more than 20% of the total duration of the path segment. The specific value can be adjusted according to the prevalence of manual intervention activities in the business process. For each path segment, all deviation values ​​within it are traversed. If a deviation value exceeds the preset deviation threshold, the path segment is marked as a path segment to be analyzed, and the deviation values ​​within the path segment to be analyzed are divided into several intervals according to their numerical values.

[0029] Specifically, the identification of areas with hidden links is based on the distribution density of deviation values ​​within each interval. The number of deviation values ​​within each interval is counted. If the proportion of deviation values ​​in a certain interval to the total number of deviation values ​​in that path segment exceeds a preset density threshold, that interval is determined to be a densely clustered interval. A typical preset density threshold is 0.3, meaning that an interval is considered densely clustered when the number of deviation values ​​in a certain interval accounts for more than 30% of the total number of deviation values. This threshold can be adjusted according to the concentration of deviation value distribution. These densely clustered intervals reflect the concentrated temporal distribution of unrecorded manual review or offline communication activities, and are thus marked as areas where hidden links exist.

[0030] It should be noted that the calculation of the silent activity time percentage uses the ratio of cumulative duration to total duration. For densely clustered intervals marked as areas containing hidden links, the durations corresponding to all deviation values ​​within that interval are accumulated to obtain the cumulative duration of the hidden links. This cumulative duration is then divided by the total duration of that path segment from the start node to the end node. The resulting ratio is the silent activity time percentage of that path segment. The silent activity time percentages of each path segment are arranged according to the order of the path segments on the critical path to form the silent activity time percentage distribution. The silent activity time percentage is calculated using formula P. silent =T hidden / T total Calculate, where T hidden T is the cumulative value of the duration corresponding to all deviation values ​​within the region where the hidden link exists. total P is the total duration of the path segment from the starting node to the ending node. silent The value range is [0,1].

[0031] In one embodiment, the baseline service time is obtained from historical normal process records, which are filtered from process instances that have not experienced abnormal delays. The abnormal delay identification process involves calculating the average μ and standard deviation σ of the service time t of each adjacent node in the historical process instance. If all nodes satisfy |t-μ| / σ, then... <T z, Among them, T z The anomaly detection threshold is typically set to 1.5, corresponding to approximately 86.6% confidence interval of a normal distribution. It can be adjusted within the range of 1.0 to 2.0 based on the sensitivity of the business process to anomalies, indicating no abnormal delay has occurred. The average service time of adjacent nodes is extracted from the historical normal process records and used as the baseline service time B for that adjacent node. The original service time A is calculated from multiple process instances to be analyzed, taking the arithmetic mean of the service times of adjacent nodes. Further, the service time inflation D is obtained by subtracting B from A; this inflation reflects the absolute increase in service time of adjacent nodes due to implicit links. The service time inflation magnitude R is obtained by dividing D by B; this magnitude, expressed as a percentage, characterizes the degree of inflation of adjacent node service time relative to the normal state, facilitating horizontal comparisons between nodes with different service time magnitudes.

[0032] S103. Identify nodes whose service time inflation exceeds a preset inflation threshold among adjacent nodes, and obtain the corrected node service time set by stripping the silent activity time component of each node into the event log.

[0033] From the service time inflation rate, adjacent nodes whose values ​​exceed a preset inflation threshold are selected and marked as inflation-overloaded nodes. Adjacent nodes whose values ​​do not exceed the preset inflation threshold are marked as normal nodes. For each inflation-overloaded node, the original service time recorded in the event log and the proportion of silent activity time are extracted. The silent activity time component of each inflation-overloaded node is obtained by multiplying the silent activity time component by the original service time. The silent activity time component is then subtracted from the original service time to obtain the stripped service time of each inflation-overloaded node. The stripped service time and the original service time of the normal nodes are summarized according to the order of the nodes on the critical path to obtain the corrected node service time set.

[0034] In one implementation, the preset inflation threshold is determined based on the historical operational characteristics of the business process. It is formed by statistically analyzing the upper limit of the service time inflation of adjacent nodes in a normal process instance, and then adding a certain tolerance range to this upper limit. The tolerance range is typically 10% to 20% of the upper limit.

[0035] For example, if the upper limit of the normal process instance service time expansion is R max=0.30, then the preset inflation threshold T inflate =R max ×(1+0.15)≈0.345), the tolerance ratio can be adjusted according to the actual fluctuation characteristics of the business process. For adjacent nodes whose service time expansion exceeds the preset expansion threshold, they are marked as expansion over-limit nodes; for adjacent nodes that do not exceed the preset expansion threshold, they are marked as normal nodes.

[0036] Specifically, the process of identifying nodes exceeding the expansion limit involves traversing the service time expansion range of all adjacent nodes on the critical path and comparing each one with a preset expansion threshold. If the expansion range of an adjacent node is greater than the preset expansion threshold, the node is determined to have interference from silent activity time consumption and is included in the set of nodes exceeding the expansion limit. If the expansion range is less than or equal to the preset expansion threshold, the node's service time is determined to be within the normal range and is included in the set of normal nodes.

[0037] It should be noted that the silent activity time component is calculated by multiplying the silent activity time percentage by the original service time. For each node exceeding the expansion limit, its original service time is extracted from the event log, and then multiplied by the silent activity time percentage of the path segment containing that node. The resulting product is the silent activity time component that the node was incorrectly classified into due to unrecorded hidden steps. The silent activity time component is calculated according to formula C. silent =T raw ×P silent Calculate, where T raw P represents the original service time of this over-expansion node. silent The percentage of silent activity time in the path segment containing this node; service time T after stripping. adj =T raw -C silent =T raw ×(1-P silent ).

[0038] In one embodiment, the service time after stripping is obtained by subtraction, subtracting the silent activity time component from the original service time of the over-expansion node, and the difference is the true service time of the node after removing the influence of hidden links.

[0039] For example, if the original service time of an order confirmation node is a certain number of time units, and the proportion of silent activity time in its path segment is a certain percentage, then the silent activity time component is the product of the two. The service time after removing silent activity interference is the difference between the original service time and this product. Furthermore, the revised set of node service times is compiled according to the order of nodes on the critical path. The revised service times of each over-limit node are arranged sequentially with the original service times of each normal node, forming a revised service time sequence covering all adjacent nodes on the critical path. This sequence reflects the true time distribution of each node after removing silent activity interference.

[0040] S104. Extract the expansion limit range of each node on the critical path from the corrected node service time set, evaluate the time consumption weight of each path segment in combination with the silent activity proportion distribution, and mark the nodes whose expansion limit range and time consumption weight both exceed the corresponding average as bottleneck candidate nodes to determine the bottleneck candidate node set.

[0041] Extract the inflation amplitude values ​​corresponding to each node from the service time inflation amplitude. Perform a difference calculation between these inflation amplitude values ​​and a preset inflation threshold to obtain the inflation excess range for each node. Calculate the arithmetic mean of the inflation excess ranges for all nodes on the critical path to obtain the average inflation excess range. Based on the silent activity time proportion of each path segment in the silent activity proportion distribution, multiply the silent activity time proportion of each path segment by its proportion in the total critical path duration to obtain the time consumption weight of each path segment. Calculate the arithmetic mean of the time consumption weights for all path segments to obtain the average time consumption weight. If the inflation excess range of a node exceeds the average inflation excess range, and the time consumption weight of the path segment containing that node exceeds the average time consumption weight, then that node is marked as a bottleneck candidate node. Summarize all bottleneck candidate nodes to determine the bottleneck candidate node set.

[0042] In one implementation, the expansion limit range is calculated based on the difference between the service time expansion magnitude of each node and a preset expansion threshold. For each node on the critical path, its expansion magnitude is subtracted from the preset expansion threshold, and the resulting difference is the expansion limit range for that node. The expansion limit range is calculated according to formula E. over =RT inflate Calculate, where R is the service time inflation rate of the node, and T... inflate E is the preset inflation threshold; over >0 indicates exceeding the limit; the larger the value, the more severe the exceedance. E over Nodes with a value ≤0 will not participate in subsequent bottleneck candidate selection. A positive value for the expansion limit indicates that the expansion of the node exceeds the normal allowable range; the larger the value, the more severe the exceedance.

[0043] Specifically, the average expansion range is calculated using an arithmetic mean. The expansion ranges of all nodes on the critical path are summed, and then divided by the total number of nodes. The resulting quotient is the average expansion range. This average serves as a benchmark for determining whether a single node is a high-expansion node.

[0044] It should be noted that the time consumption weight reflects the contribution of each path segment to the overall process time consumption. For each path segment, the percentage of silent activity time for that path segment and the percentage of that path segment's duration in the total critical path duration are obtained. These two values ​​are then multiplied; the product is the time consumption weight for that path segment. Time consumption weight W seg According to formula W seg =P silent ×P dur Calculate, where P silent P represents the percentage of silent activity time for this path segment. dur The percentage of this path segment's duration in the total duration of the critical path (P) dur =T seg / T total path T seg T represents the path segment duration. total path (W is the total duration of the critical path). The physical meaning of multiplying the two is: the path segment with the greater contribution to path segment duration and the more significant the silent activity, the greater its comprehensive impact on the overall process time consumption. seg A higher value indicates a greater potential for optimization of that path segment.

[0045] In one embodiment, the average time consumption weight is also calculated using an arithmetic mean method, by summing the time consumption weights of all path segments and dividing by the total number of path segments.

[0046] For example, in an order processing workflow, if a path segment contains two adjacent nodes, order review and order confirmation, and this path segment has a high proportion of silent activity time and a large proportion of its duration in the overall process, then the time consumption weight of this path segment is high. Furthermore, the selection of bottleneck candidate nodes employs a dual-condition judgment method. Each node on the critical path is traversed; if the expansion limit of a node exceeds the average expansion limit, and the time consumption weight of the path segment containing that node also exceeds the average time consumption weight, then that node is marked as a bottleneck candidate node. All nodes meeting both conditions are aggregated to form a bottleneck candidate node set. Nodes in this set are characterized by high expansion levels and significant time consumption impact.

[0047] S105. By analyzing the bottleneck candidate node set, identify the fluctuation range of the silent activity proportion of each candidate node in the path segment, weight and fuse the fluctuation range of the proportion with the time consumption weight to generate a comprehensive risk index for each path segment, identify the distribution range of the comprehensive risk index, and dynamically scale the preset bottleneck judgment benchmark threshold to obtain the adjusted bottleneck judgment threshold.

[0048] The silent activity time percentage of each candidate node in the bottleneck candidate node set is extracted. The silent activity time percentages of multiple process instances within the same path segment are statistically analyzed to obtain the maximum and minimum values ​​of the silent activity time percentage for each path segment. The difference between the maximum and minimum values ​​is calculated to obtain the percentage fluctuation range for each path segment. Based on the percentage fluctuation range and the time consumption weight of each path segment, both are multiplied by their respective preset weight values ​​and then summed to obtain the comprehensive risk index for each path segment. The comprehensive risk indices of all path segments are sorted, and the maximum and minimum values ​​of the comprehensive risk indices are identified. The difference between the maximum and minimum values ​​is taken as the distribution interval span. The distribution interval span is divided by a preset bottleneck judgment benchmark threshold to obtain a scaling factor. If the distribution interval span exceeds the preset bottleneck judgment benchmark threshold, the preset bottleneck judgment benchmark threshold is multiplied by the scaling factor to amplify it; if the distribution interval span does not exceed the preset bottleneck judgment benchmark threshold, the preset bottleneck judgment benchmark threshold is divided by the scaling factor to shrink it, resulting in an adjusted bottleneck judgment threshold.

[0049] In one implementation, the percentage fluctuation range is extracted based on historical operational data of the path segments where each candidate node in the bottleneck candidate node set is located. For each path segment, the percentage of silent activity time for that path segment in multiple process instances is collected. All values ​​are iterated to determine the maximum and minimum values, and the difference between the two is the percentage fluctuation range for that path segment. The larger the percentage fluctuation range value, the more significant the difference in silent activity performance of that path segment in different process instances.

[0050] Specifically, the calculation of the percentage fluctuation range covers all path segments involved in the bottleneck candidate node set.

[0051] For example, in the procurement approval process, if a certain path segment contains two adjacent nodes, supplier qualification review and contract terms confirmation, and the proportion of silent activity time of this path segment fluctuates in each procurement process, then the extreme values ​​of the proportion of silent activity time of this path segment in all process instances are counted to obtain the range of fluctuation of the proportion of this path segment.

[0052] It should be noted that the comprehensive risk index is calculated using a weighted summation method. The weights for both the percentage fluctuation range and time consumption are multiplied by their respective preset weight values ​​and then summed. These preset weight values ​​are determined in advance based on the characteristics of the business process. The weight corresponding to the percentage fluctuation range reflects the relative importance of process stability in risk assessment, while the weight corresponding to the time consumption reflects the relative importance of time consumption in risk assessment. The weighted sum of these two values ​​is the comprehensive risk index for that path segment. A higher comprehensive risk index indicates a greater risk that the path segment will become a bottleneck.

[0053] In one embodiment, the preset weight values ​​are determined based on the historical bottleneck distribution of the business process. If historical data shows that bottlenecks mostly occur in path segments with large fluctuations in silent activity, the weight value corresponding to the percentage fluctuation range is set higher; if historical data shows that bottlenecks mostly occur in path segments with high time consumption weight, the weight value corresponding to the time consumption weight is set higher. The sum of the two weight values ​​is usually set to 1 (i.e., w1 + w2 = 1) to maintain the comparability of the numerical range of the comprehensive risk index; typical values ​​are w1 = 0.6 (percentage fluctuation range weight) and w2 = 0.4 (time consumption weight), which can be adjusted according to the relative contribution of silent activity fluctuations and time consumption to the occurrence of bottlenecks in historical data. The comprehensive risk index is calculated as RI = w1 × V. occ +w2×W seg V occ To represent the range of fluctuations in percentage, W seg Time consumption is weighted. Furthermore, the distribution interval span is determined based on the extreme value distribution of the comprehensive risk index for all path segments. The comprehensive risk indexes of all path segments are sorted according to their numerical values, and the maximum and minimum values ​​after sorting are identified; the difference between these two values ​​is the distribution interval span. The distribution interval span reflects the dispersion of the comprehensive risk index for each path segment; a larger span indicates a more significant risk difference between different path segments.

[0054] For example, the scaling factor is calculated as the ratio of the distribution range to a preset bottleneck determination benchmark threshold. The scaling factor is the quotient obtained by dividing the distribution range by the preset bottleneck determination benchmark threshold. A scaling factor greater than 1 indicates that the current comprehensive risk index distribution range exceeds the range covered by the benchmark threshold, while a scaling factor less than 1 indicates that the current comprehensive risk index distribution range is within the range covered by the benchmark threshold.

[0055] In one possible implementation, the bottleneck determination threshold is dynamically adjusted based on the relationship between the distribution interval span and the preset bottleneck determination benchmark threshold. If the distribution interval span exceeds the preset bottleneck determination benchmark threshold, it indicates that the distribution range of the comprehensive risk index is wide. In this case, the preset bottleneck determination benchmark threshold is multiplied by a scaling factor to amplify it, allowing the adjusted bottleneck determination threshold to adapt to a larger risk distribution range. If the distribution interval span does not exceed the preset bottleneck determination benchmark threshold, it indicates that the distribution range of the comprehensive risk index is narrow. In this case, the preset bottleneck determination benchmark threshold is divided by a scaling factor to shrink it, allowing the adjusted bottleneck determination threshold to more accurately identify bottleneck nodes within the risk concentration area. The preset bottleneck determination benchmark threshold T... base A typical value for T is 1.2 times the average of the comprehensive risk index. base =mean(RI)×1.2, where mean(RI) is the average of the comprehensive risk index for all path segments. The specific value can be adjusted according to the requirements of the business process for bottleneck identification accuracy. To prevent the threshold from exceeding a reasonable range after dynamic scaling, the adjusted bottleneck judgment threshold is subject to upper and lower limit constraints: the upper limit is the maximum value of the comprehensive risk index, and the lower limit is the minimum value of the comprehensive risk index. If the scaling result exceeds the upper or lower limit, it will be truncated to the corresponding boundary value.

[0056] Understandably, the dynamic scaling mechanism allows the bottleneck identification threshold to be adaptively adjusted based on the distribution characteristics of the actual comprehensive risk index. In business processes with significantly different risk distributions, the enlarged threshold avoids misidentifying too many nodes as bottlenecks; in business processes with relatively concentrated risk distributions, the reduced threshold can more accurately capture potential bottleneck nodes, improving the targeting of bottleneck identification.

[0057] S106. Based on the adjusted bottleneck determination threshold, the support vector machine algorithm is used to classify the time dilation features of each node in the bottleneck candidate node set and identify high-risk silent bottleneck nodes.

[0058] The time dilation features of each candidate node in the bottleneck candidate node set are extracted. These features include the service time dilation magnitude, the dilation exceedance range, and the comprehensive risk index of the path segment where the node is located. For each candidate node, these three feature values ​​are arranged and combined in a fixed order to construct a feature vector. The feature vectors are then classified into risk categories based on an adjusted bottleneck determination threshold. If the comprehensive risk index of the path segment where the candidate node is located exceeds the adjusted bottleneck determination threshold, the feature vector is labeled as high-risk; otherwise, it is labeled as low-risk, resulting in a set of labeled samples with risk category labels. A support vector machine (SVM) algorithm is used to train the labeled sample set, using the feature vectors of each candidate node as input and the risk category labels as output. The classification boundary between high-risk and low-risk categories is determined by finding the maximum margin hyperplane, resulting in a trained classifier. The feature vectors of each candidate node in the bottleneck candidate node set are input into the classifier. Based on the classifier's output category determination results, candidate nodes judged as high-risk are identified as high-risk silent bottleneck nodes.

[0059] In one implementation, the extraction of time dilation features covers all candidate nodes in the bottleneck candidate node set. For each candidate node, its service time dilation magnitude, dilation exceedance range, and comprehensive risk index of the path segment where the node is located are obtained from the aforementioned processing. These three feature values ​​are used as components of the node's time dilation features. The service time dilation magnitude reflects the growth rate of the node's service time relative to the baseline service time, the dilation exceedance range reflects the extent to which the node's dilation exceeds a preset threshold, and the comprehensive risk index reflects the overall risk level of the path segment where the node is located.

[0060] Specifically, the feature vectors are constructed using a fixed-order arrangement, placing the service time dilation magnitude in the first dimension, the dilation exceeding the limit in the second dimension, and the comprehensive risk index in the third dimension. Through this arrangement, each candidate node corresponds to a three-dimensional feature vector, which can characterize the time dilation state of that node in a multi-dimensional space.

[0061] It should be noted that the risk category labels are based on the adjusted bottleneck determination threshold. The feature vectors of all candidate nodes are traversed, and for each candidate node, the relationship between the comprehensive risk index of its path segment and the adjusted bottleneck determination threshold is checked. If the comprehensive risk index exceeds the adjusted bottleneck determination threshold, the feature vector corresponding to that candidate node is labeled as high-risk; if the comprehensive risk index does not exceed the adjusted bottleneck determination threshold, the feature vector corresponding to that candidate node is labeled as low-risk. All labeled feature vectors and their corresponding risk category labels together constitute the labeled sample set.

[0062] In one embodiment, the training process of the Support Vector Machine (SVM) algorithm uses a set of labeled samples as input data. The core objective of the SVM algorithm is to find a hyperplane in the feature space such that high-risk class samples and low-risk class samples are located on opposite sides of the hyperplane, and the sum of the distances from the two classes to the hyperplane is maximized. This hyperplane is called the maximum margin hyperplane, which establishes the most discriminative classification boundary between the two classes. The sample points closest to the hyperplane are called support vectors, and these support vectors determine the position and orientation of the hyperplane. Further, the parameters of the maximum margin hyperplane are determined through iterative optimization during the training process. In the three-dimensional feature space, the hyperplane is represented as a two-dimensional plane, and its position is determined by the normal vector and the bias. During training, the SVM algorithm adjusts the values ​​of the normal vector and the bias so that all high-risk class samples are located on one side of the hyperplane and at a distance greater than a preset margin, while all low-risk class samples are located on the other side of the hyperplane and at a distance greater than the preset margin. When the labeled sample set is linearly separable in the 3D feature space, the Support Vector Machine (SVM) algorithm uses a linear kernel to construct a maximum margin hyperplane. When the samples are linearly inseparable, a Radial Basis Function (RBF) kernel is used to map the features to a high-dimensional space to construct a classification hyperplane. The kernel parameter γ is selected from the candidate set {0.01, 0.1, 1.0} through cross-validation. The regularization parameter C typically takes a value of 1.0 and can be adjusted according to the actual sample distribution. When the above constraints are met and the margin value reaches its maximum, the training process converges, and the trained classifier is obtained.

[0063] For example, in the bottleneck identification scenario of the order processing flow, if the feature vector of a candidate node is a combination of high service time inflation, medium inflation limit, and high comprehensive risk index, the position of this feature vector in the three-dimensional feature space will fall within the high-risk region. The maximum margin hyperplane obtained by training with the support vector machine algorithm can clearly distinguish such feature vectors from low-risk feature vectors.

[0064] In one possible implementation, the classifier is applied one by one to each candidate node in the bottleneck candidate node set. When the bottleneck candidate node set is small (less than 30 nodes), to prevent overfitting, leave-one-out cross-validation is used during training to evaluate the classifier's performance. Simultaneously, the labeled sample set is divided into a training set and a validation set in a 7:3 ratio. The training set is used to fit the hyperplane parameters, and the validation set is used to evaluate the classification accuracy, ensuring the classifier has a certain generalization ability. The feature vectors of the candidate nodes are input into the trained classifier, which determines the classifier based on the position of the feature vector relative to the maximum margin hyperplane in the three-dimensional feature space. If the feature vector is located on the side of the hyperplane corresponding to the high-risk class, the classifier outputs a high-risk class determination; if the feature vector is located on the side of the hyperplane corresponding to the low-risk class, the classifier outputs a low-risk class determination.

[0065] Understandably, the identification of high-risk silent bottleneck nodes is based on the classifier's output. By iterating through the classification results of all candidate nodes, those nodes judged as high-risk by the classifier are marked as high-risk silent bottleneck nodes. These nodes are characterized by significant service time inflation, large inflation exceedance ranges, and high overall risk in their respective path segments, representing potential bottleneck locations that require close monitoring in the business process.

[0066] S107. Simulate high-risk silent bottleneck nodes, obtain path simulation data from the simulation platform, calculate the coverage improvement rate using logistic regression algorithm, and predict the bottleneck coverage rate.

[0067] A simulation process is constructed in a simulation platform for the high-risk silent bottleneck nodes. The service time of each node is configured according to the node order of the critical path. The service time uses the values ​​of the corresponding nodes in the corrected node service time set. Multiple sets of different silent activity time ratios are set for the high-risk silent bottleneck nodes for simulation runs to obtain path simulation data for each high-risk silent bottleneck node under different silent activity time ratios. The simulation service time and total simulation path time of each high-risk silent bottleneck node are extracted from the path simulation data. The difference between the simulation service time and the corrected node service time is calculated to obtain the service time change. The difference between the total simulation path time and the original total path time is calculated to obtain the path time change. The service time change and path time change are combined to construct coverage feature data. A logistic regression algorithm is used to train the coverage feature data, using the service time change and path time change as input features and whether the bottleneck node is identified as the output label. The maximum likelihood estimation method is used to fit the mapping relationship between the features and the identification results to obtain the trained coverage predictor. The coverage feature data of each high-risk silent bottleneck node is input into the coverage predictor to obtain the coverage probability value of each node. The coverage probability value is then compared with the original coverage probability benchmark value to obtain the coverage improvement rate. The coverage improvement rates of all high-risk silent bottleneck nodes are summarized to predict the overall bottleneck coverage rate.

[0068] In one implementation, the simulation platform is configured based on the structure of high-risk silent bottleneck nodes and their corresponding critical paths. For each node on the critical path, a node sequence for the simulation process is established according to the node execution order, and the service time parameter of each node is taken from the value of the corresponding node in the corrected node service time set. The simulation platform can simulate the execution process of the business process under different conditions, and record the simulation service time of each node and the total simulation time of the entire path.

[0069] Specifically, the simulation variables are set based on the percentage of silent activity time for high-risk silent bottleneck nodes. In the simulation platform, multiple sets of different silent activity time percentage values ​​are configured for each high-risk silent bottleneck node, each representing a hypothetical level of silent activity intensity. The simulated silent activity time percentage is evenly divided into 5 groups from 0% to the maximum measured percentage (5 groups are typical values; the number of groups can be adjusted according to simulation accuracy requirements). For example, when the maximum measured percentage is 50%, the 5 groups are 0%, 12.5%, 25%, 37.5%, and 50%, respectively. The simulation data is recorded after each group runs independently. By running each set of simulation configurations sequentially, path simulation data for each high-risk silent bottleneck node under different silent activity time percentage conditions is obtained. The path simulation data includes the simulation service time of the node under the current simulation conditions and the total simulation path time for the entire critical path.

[0070] It should be noted that the coverage feature data is constructed using a two-dimensional interpolation method. The first dimension is the change in service time, obtained by subtracting the corrected node service time from the simulated service time of high-risk silent bottleneck nodes. This change reflects the degree of deviation of node service time from the corrected baseline under simulated conditions. The second dimension is the change in path time, obtained by subtracting the original path time from the simulated total path time. This change reflects the degree of deviation of overall path time from the original baseline under simulated conditions. The changes in these two dimensions are combined to form the coverage feature data.

[0071] In one embodiment, the training process of the logistic regression algorithm uses coverage feature data as the input dataset. Logistic regression is a supervised learning method widely used in binary classification problems. Its core principle is to convert continuous feature values ​​into discrete class probabilities by constructing a mapping function between features and output. In this embodiment, the input features are two-dimensional vectors of service time variation and path time variation, and the output label is a binary marker indicating whether the bottleneck node is identified. Furthermore, the logistic regression algorithm uses maximum likelihood estimation for parameter estimation. The core idea of ​​maximum likelihood estimation is to find a set of parameter values ​​that maximize the probability of observing the current training data under those parameter conditions. During training, the logistic regression algorithm iteratively adjusts the weight and bias parameters in the mapping function until the likelihood function value of the training data converges to a stable state. After training, the coverage predictor can output the corresponding coverage probability value based on the input coverage feature data. This probability value represents the likelihood that the bottleneck node is identified and covered under the current feature conditions.

[0072] For example, in a bottleneck prediction scenario for a procurement approval process, if the service time change and path time change of a high-risk silent bottleneck node are both positive under simulation conditions, it indicates that the increase in the service time of this node in the simulation environment has led to an increase in the overall path time. This type of feature data is input into a trained coverage predictor, which outputs the corresponding coverage probability value based on the position of the feature value in the mapping function.

[0073] In one possible implementation, the calculation of the coverage improvement is based on a comparison between the coverage probability value and a coverage probability baseline value. The coverage probability baseline value is calculated using the formula baseline=covered. nodes / total nodes Calculate (values ​​in the range [0,1], no need to multiply by 100), where covered nodes The total number of bottleneck nodes identified by the algorithm under the condition of zero silent percentage. nodes The baseline is the total number of high-risk silent bottleneck nodes. For example, if 8 out of 10 nodes are identified, then the baseline is 0.80; the coverage improvement Δp = p i -baseline, Δp>0 indicates an improvement in coverage capability. This baseline value reflects the coverage capability of the bottleneck identification method for this node under the original conditions. Subtracting the coverage probability baseline value from the coverage probability value output by the coverage probability predictor, the difference is the improvement in coverage rate for this node. A positive improvement in coverage rate indicates an increase in the probability of the node being identified after simulation optimization, while a negative value indicates a decrease in coverage probability.

[0074] Understandably, the overall bottleneck coverage rate is predicted by weighting the coverage probability values ​​of each high-risk silent bottleneck node, calculated using the formula Coverage. total =Σ(w node i ×p i ) / Σ(w node i ), where p i w is the coverage probability value of the i-th high-risk silent bottleneck node (output by the coverage predictor, with a value range of [0,1]). node i The Coverage is the weight of this node (the value is determined below). total The overall probability that each silent bottleneck node is fully identified and covered after simulation optimization. Node weight w node i Take the time consumption weight W of the path segment where the node is located. seg (In line with S104) This means that bottleneck nodes with a greater impact from time consumption have a higher contribution weight in the overall coverage prediction.

[0075] If the technical solution of this application involves the collection, processing, or application of personal information, the relevant products have strictly complied with the requirements of the "Personal Information Protection Law of the People's Republic of China" and other laws and regulations before implementing any personal information processing activities, clearly and explicitly informing individuals of the rules for personal information processing and obtaining their independent and voluntary authorization and consent. Specifically, if the information involved is sensitive personal information, the product has not only obtained the individual's separate consent before processing, but this consent is also an explicit consent made on the basis of full knowledge. For example, in areas where personal information collection devices such as cameras are deployed, prominent and eye-catching signs have been set up to clearly inform users that entering the area is considered as consenting to the collection of their personal information; or, on the personal information processing interface (such as applications, web pages, etc.), through pop-ups, checkboxes, or active uploads, the user is required to actively authorize the process after clearly displaying key rules such as the identity of the personal information processor, the purpose of processing, the processing method, and the types of information involved.

[0076] The above are only some preferred embodiments of the present invention, but the present invention is not limited thereto, and many improvements and modifications can be made. Any improvements and modifications made based on the basic principles of the present invention should be considered to fall within the protection scope of the present invention.

Claims

1. A business process bottleneck prediction and simulation method based on process mining, characterized in that, include: Read the timestamp records of each node and the path segment data on the critical path from the event log. Extract the interval between adjacent nodes and the service time distribution of each node through the log parsing unit. Perform a difference operation on the theoretical completion timestamp within the interval between adjacent nodes and the service time distribution to obtain a set of timestamp deviation values. By analyzing the set of timestamp deviation values, the time proportion of hidden business processes that are not recorded within the path segment is identified, the distribution of silent activities is determined, and the service time inflation of adjacent nodes is extracted. If the service time inflation rate of a node among the adjacent nodes exceeds a preset inflation threshold, the silent activity time component of each node in the event log is removed, and the corrected node service time set is extracted. Extract the expansion limit range of each node on the critical path from the modified node service time set, evaluate the time consumption weight of each path segment in combination with the silent activity proportion distribution, determine the nodes whose expansion limit range and time consumption weight both exceed the corresponding average, and aggregate to generate a bottleneck candidate node set. Analyze the bottleneck candidate node set, obtain the fluctuation range of the silent activity ratio of the path segment where each candidate node is located, perform a weighted fusion operation based on the fluctuation range of the ratio and the time consumption weight to generate a comprehensive risk index, extract the distribution range of the comprehensive risk index and perform dynamic scaling on the preset judgment benchmark threshold to generate an adjustment judgment threshold; Based on the adjusted judgment threshold, the time dilation features of each node in the bottleneck candidate node set are classified to identify high-risk bottleneck nodes. Simulation processing is performed on the high-risk bottleneck nodes to obtain path simulation data, and the overall bottleneck coverage is predicted based on the coverage improvement rate calculated by logistic regression.

2. The method according to claim 1, characterized in that, The process involves reading timestamp records of each node and path segment data on the critical path from the event log, extracting the interval between adjacent nodes and the service time distribution of each node through the log parsing unit, and performing a difference calculation on the service time distribution and the theoretical completion timestamp within the interval between adjacent nodes to obtain a set of timestamp deviation values, including: Identify the start and end nodes of each path segment on the critical path; Extract the node service time based on the end timestamp minus the start timestamp of the process instance; Based on the node service time of the multi-process instance, the mean and variance of service time are calculated, and the service time distribution is established. The average service time between adjacent nodes is summed to generate a theoretical completion timestamp; Calculate the difference between the theoretical completion timestamp and the actual completion timestamp to generate a set of timestamp deviation values.

3. The method according to claim 1, characterized in that, The process of analyzing the timestamp deviation value set to identify the time proportion of unrecorded hidden business processes within the path segment, determining the distribution of silent activity proportions, and extracting the service time inflation of adjacent nodes includes: The intervals where the deviation values ​​are densely clustered in the set of timestamp deviation values ​​are extracted as the regions where hidden links exist; The percentage of silent activity time is generated based on the ratio of the cumulative duration of deviation values ​​within the area where the hidden link exists to the total duration of the path segment. Perform a sequence permutation on the percentage of silent activity time in multiple groups to construct a distribution of silent activity percentage; Calculate the difference between the original service time and the baseline service time to extract the service time inflation. The ratio of the service time inflation to the baseline service time is calculated to generate the service time inflation magnitude.

4. The method according to claim 1, characterized in that, For nodes whose service time inflation exceeds a preset inflation threshold among adjacent nodes, the silent activity time component of each node in the event log is stripped, and a corrected set of node service times is extracted, including: Nodes whose service time expansion exceeds a preset expansion threshold are selected from the adjacent nodes as nodes that have exceeded the expansion limit. Obtain the ratio of the original service time to the silent activity time associated with the node that exceeds the expansion limit; The silent activity time component is obtained by multiplying the silent activity time percentage by the original service time. The stripped service time is generated by subtracting the silent activity time component from the original service time; The service time after stripping is combined with the original service time of nodes that have not exceeded the limit to form the corrected set of node service times.

5. The method according to claim 1, characterized in that, The process involves extracting the expansion limits of each node on the critical path from the corrected node service time set, evaluating the time consumption weight of each path segment in conjunction with the silent activity proportion distribution, determining nodes whose expansion limits and time consumption weights both exceed their corresponding averages, and aggregating these nodes to generate a bottleneck candidate node set, including: Identify the expansion amplitude of each node and subtract it from the preset expansion threshold to obtain the expansion excess range; The arithmetic mean of the expansion excess range is calculated to obtain the mean value of the expansion excess range; The time consumption weight is calculated by multiplying the percentage of time spent in silent activities by the percentage of total time spent in silent activities. The arithmetic mean of the time consumption weights is calculated to obtain the mean of the time consumption weights; A bottleneck candidate node set is established by combining nodes that meet the criteria of having an expansion limit greater than the average of the expansion limit range and a time consumption weight greater than the average of the time consumption weight.

6. The method according to claim 1, characterized in that, The analysis of the bottleneck candidate node set obtains the fluctuation range of the silent activity proportion of the path segment where each candidate node is located. Based on the fluctuation range of the proportion and the time consumption weight, a weighted fusion operation is performed to generate a comprehensive risk index. The distribution range of the comprehensive risk index is extracted and dynamically scaled against a preset judgment benchmark threshold to generate an adjusted judgment threshold, including: Extract the maximum and minimum percentages of silent activity time among multiple process instances within the same path segment; The difference between the extracted maximum and minimum values ​​is used to generate the percentage fluctuation range of each path segment; After assigning independent preset weight values ​​to the aforementioned percentage fluctuation range and the time consumption weight, a summation process is performed to generate a comprehensive risk index. The comprehensive risk index of all path segments is retrieved to obtain the global maximum and global minimum values; The difference between the global maximum value and the global minimum value is used to calculate the span of the distribution interval; The numerical scaling factor is obtained by dividing the distribution interval span by a preset judgment benchmark threshold. In response to a determination result where the span of the distribution interval is greater than a preset determination benchmark threshold, an adjustment determination threshold is obtained by multiplying the preset determination benchmark threshold by the numerical scaling factor.

7. The method according to claim 1, characterized in that, The step of classifying the time dilation features of each node in the bottleneck candidate node set according to the adjusted judgment threshold to identify high-risk bottleneck nodes includes: Obtain the service time inflation rate, inflation limit range, and comprehensive risk index associated with each candidate node; The feature vectors of each candidate node are constructed by combining the three feature values ​​obtained. In response to the condition that the comprehensive risk index is greater than the adjustment judgment threshold, the feature vector is labeled with a high-risk category to construct labeled samples; For the labeled samples input into the classification model, perform maximum margin boundary training to extract the classifier; The classifier performs pattern matching on the feature vector of the node to be detected to read the corresponding category data and identify high-risk bottleneck nodes.

8. The method according to claim 1, characterized in that, The step of performing simulation processing on the high-risk bottleneck nodes to obtain path simulation data, and predicting the overall bottleneck coverage based on logistic regression to calculate the coverage improvement, includes: A simulation testing process was established for the aforementioned high-risk bottleneck nodes; The test service time corresponding to each node is read by simulating the node sequence of the critical path and the corrected node service time set through the simulation test process. Set up multiple sets of silent activity time percentage parameters with numerical differences to drive the test process and output independent path simulation data; Read the simulation service time of each node and calculate the difference between it and the original corrected service time to establish the service time change; Read the total time of the simulated path and calculate the difference between it and the total time of the original path to establish the change in path time; Combine the service time change and the path time change to construct coverage feature association data; The coverage feature-related data is input into a logistic regression analyzer to perform maximum likelihood calculation and extract the independent coverage probability value of each node; The difference between the independent coverage probability value and the original baseline coverage probability value is calculated to determine the coverage improvement. The overall bottleneck coverage rate is generated by summing up the coverage improvement of all nodes.