A project execution risk assessment system and method based on data mining

Through detailed modeling and multi-dimensional data analysis of key tasks and their relationships during project execution, multiple regression models and dynamic data update mechanisms are adopted to solve the problem that existing technology is difficult to accurately evaluate the cumulative risks of task chains under complex dependencies, and the accurate assessment and timely risk warning of cumulative risks of task chain dependencies in projects are achieved, reducing the risk of project delays or failures.

CN119151297BActive Publication Date: 2025-05-09YANGZHOU JICHI INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411277306.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-12
Publication Date
2025-05-09
Estimated Expiration
2044-09-12

AI Technical Summary

Technical Problem

Existing data mining-based project execution risk assessment techniques are difficult to accurately evaluate and predict the cumulative risks of task chains under complex dependencies, making it difficult for managers to take risk mitigation measures in a timely manner, increasing the risk of project delays or failures.

Method used

By modeling the critical tasks and their relationships during project execution, and conducting comprehensive analysis with multi-dimensional data, using multiple regression models and dynamic data update mechanisms, the risk assessment results are adjusted in real time, high-risk, medium-risk and low-risk task chains are identified, and corresponding risk assessment reports are generated.

Benefits of technology

It realizes an accurate assessment of the cumulative risks of task chain dependence in the project, provides timely risk warning information, helps managers make scientific decisions, reduces the risk of project delays or failures, and optimizes resource allocation and monitoring efforts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119151297B_ABST
    Figure CN119151297B_ABST
Patent Text Reader

Abstract

The present invention discloses a project execution risk assessment system and method based on data mining, which relates to the technical field of project execution risk assessment, and specifically includes the following steps: obtaining comprehensive data information related to project execution, and storing it after obtaining; analyzing the stored comprehensive data information, evaluating the cumulative risk of each task chain dependency during the project execution process, and classifying each task chain in the project according to the evaluation results, identifying high-risk, medium-risk and low-risk task chains, and generating a corresponding risk assessment report; according to the risk assessment report, it is recommended that managers take corresponding risk mitigation measures for task chains of different risk levels; updating the comprehensive data information in real time, and regularly recalculating the cumulative risk assessment coefficient, and dynamically adjusting the risk assessment results. The present invention solves the problem of cumulative risk assessment of complex dependent task chains, and realizes precise classification management and dynamic risk adjustment effects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of project execution risk assessment, and in particular to a project execution risk assessment system and method based on data mining. Background Art

[0002] Project execution refers to the process of implementing the tasks and activities in the project plan to achieve the expected project goals. Project execution involves multiple aspects such as resource allocation, time management, and task coordination, and is a crucial link in the project life cycle. However, during project execution, there are often many risks such as schedule delays, cost overruns, and substandard quality. If these risks are not effectively managed, they may lead to project failure. Therefore, it is particularly important to conduct risk assessment on project execution. Through risk assessment, potential problems that may affect the success of the project can be identified, and preventive measures or emergency plans can be taken to mitigate the negative impact of risks on the project. There are many ways to assess project execution risks. Traditional methods mainly include expert review, experience summary, and qualitative analysis. However, these methods often rely on people's subjective judgment and are difficult to ensure comprehensiveness and accuracy. With the development of technology, data-driven quantitative analysis methods have gradually become an important means of risk assessment, among which data mining technology is particularly widely used.

[0003] The existing project execution risk assessment technology based on data mining mainly analyzes a large amount of data accumulated during the project execution process to mine potential risk factors and predict the risk status in project execution. This technology usually includes multiple steps such as data collection, data preprocessing, feature extraction, model training and evaluation. First, the system collects relevant data on project execution, including project progress, resource allocation, cost control, problem records and other information. Then, through the data preprocessing step, these data are cleaned, sorted and standardized to ensure the consistency and reliability of the data. In the feature extraction stage, the system uses data mining algorithms (such as classification, clustering, regression, etc.) to extract key risk-related features from massive data. These features may include delay patterns of specific tasks, cost deviation trends, workloads of team members, etc. Finally, the system evaluates the current or future project execution through the trained risk assessment model and outputs a risk assessment report or risk warning prompt. Although the existing technology can improve the accuracy of risk assessment to a certain extent, there are still problems such as high algorithm complexity, strong data dependence, and difficulty in interpretation, which limit its wide application in actual project management.

[0004] The prior art has the following deficiencies:

[0005] In projects involving complex dependencies between multiple tasks, especially in areas such as supply chain management or complex system development, there is often a strong dependency between tasks during project execution. When a predecessor task is delayed, it will affect multiple subsequent tasks in a chain reaction, thereby triggering cumulative risks in the task chain, which in turn affects the progress and success rate of the entire project. However, when dealing with the cumulative risks brought about by such dependencies, existing project execution risk assessment technologies based on data mining are usually difficult to accurately assess and predict their chain reactions. This is mainly because existing technologies lack in-depth modeling and effective analysis of complex dependencies between tasks, resulting in the inability to timely identify and warn of these potential major project delays or failure risks. Therefore, managers often fail to take timely and effective risk mitigation measures, which increases the risk of large-scale delays or even project failures during project execution.

[0006] The above information disclosed in this Background section is only for enhancement of understanding of the background of the present disclosure and therefore it may contain information that does not constitute the prior art that is already known to one of ordinary skill in the art. Summary of the invention

[0007] The purpose of the present invention is to provide a project execution risk assessment system and method based on data mining to solve the problems in the above-mentioned background technology.

[0008] In order to achieve the above object, the present invention provides the following technical solution: a project execution risk assessment method based on data mining, specifically comprising the following steps:

[0009] At the beginning of the project, determine the key tasks of the project and their interrelationships, configure the data collection system, define the information collection rules, and determine the data information structure for subsequent analysis;

[0010] Acquire comprehensive data information related to project execution and store it after acquisition;

[0011] Analyze the stored comprehensive data information, evaluate the cumulative risk of dependencies of various task chains during project execution, and classify the various task chains in the project according to the evaluation results, identify high-risk, medium-risk and low-risk task chains, and generate corresponding risk assessment reports;

[0012] Based on the risk assessment report, managers are advised to take appropriate risk mitigation measures for task chains with different risk levels;

[0013] During the project execution, comprehensive data information is updated in real time, and the cumulative risk assessment coefficient is recalculated regularly to dynamically adjust the risk assessment results.

[0014] Preferably, the comprehensive data information includes task execution time data, task dependency path data, task critical path data, task delay data, task chain length data and task chain concurrency data. The comprehensive data information is acquired from the project management system, the real-time monitoring platform and the task log through the configured automatic data acquisition module, and is stored in a data storage unit for subsequent analysis after acquisition.

[0015] Preferably, the stored comprehensive data information is analyzed to evaluate the cumulative risk of dependencies of various task chains during project execution, and the various task chains in the project are classified according to the evaluation results to identify high-risk, medium-risk and low-risk task chains, and generate a corresponding risk assessment report, which specifically includes the following steps:

[0016] Extracting task execution time data, task dependency path data, task critical path data, task delay data, task chain length data, and task chain concurrency data of each task chain from the data storage unit;

[0017] Preprocess the extracted data of each task chain;

[0018] Analyze the preprocessed data to generate task chain influence coefficients and cumulative delay indexes for each task chain;

[0019] A task chain cumulative risk assessment model is constructed based on the task chain impact coefficient and cumulative delay index of each generated task chain to generate a cumulative risk assessment coefficient for each task chain;

[0020] Determine the pre-set cumulative risk assessment coefficient threshold interval, and compare the generated cumulative risk assessment coefficient of each task chain with the threshold interval. According to the comparison results, evaluate the cumulative risk level of each task chain dependency during the project execution process, and classify the task chains in the project according to the evaluation results, identify high-risk, medium-risk and low-risk task chains, and generate a corresponding risk assessment report.

[0021] Preferably, the logic for obtaining the task chain influence coefficient and the cumulative delay index of each task chain is as follows:

[0022] Extract the actual execution time and planned execution time of each task in the task execution time data of each task chain, and mark the actual execution time and planned execution time of each task in each task chain as and Indicates the actual execution time of the nth task in the mth task chain, represents the planned execution time of the nth task in the mth task chain, m = 1, 2, 3, ..., k, n = 1, 2, 3, ..., g, k and g are both positive integers;

[0023] Compare the actual execution time of each task in each task chain with the planned execution time, and calculate the time deviation of each task. The specific calculation formula is as follows:

[0024]

[0025] In the formula, is the time deviation of the nth task in the mth task chain;

[0026] Extract the number of dependent tasks of each task in the task dependency path data of each task chain, and set the target number of dependent tasks of each task in the task dependency path data of each task chain as It represents the number of tasks that the nth task in the mth task chain directly depends on. The specific quantitative formula is as follows:

[0027]

[0028] In the formula, Indicates whether the nth task in the mth task chain depends on the pth task. If there is a dependency, Take 1, otherwise take 0, p = 1, 2, 3, ..., g, g is the total number of tasks in the task chain;

[0029] Extract the critical path data of each task chain and mark it as Indicates whether the nth task in the mth task chain is on the critical path. When the task is on the critical path, Take 1, otherwise take 0. At the same time, extract the resource consumption of each task and the total resource consumption of all tasks in the project and mark them as and R total , Indicates the resource consumption of the nth task in the mth task chain;

[0030] Calculate the comprehensive weight of each task in each task chain. The specific quantitative formula is as follows:

[0031]

[0032] In the formula, is the comprehensive weight of the nth task in the mth task chain;

[0033] Calculate the task chain influence coefficient of each task chain. The specific calculation formula is as follows:

[0034]

[0035] Where TCIC mis the task chain influence coefficient of the mth task chain;

[0036] Extract the task chain length data of each task chain and mark it as L m , L m It represents the number of path tasks from the starting task to the end task of the mth task chain. The specific quantitative formula is as follows:

[0037]

[0038] In the formula, Indicates whether the nth task in the mth task chain is on the path. When the task is on the path, Takes 1, otherwise takes 0, g is the total number of tasks in the task chain;

[0039] Extract the task chain concurrency data of each task chain and mark it as C m , C m It represents the number of tasks executed in parallel in the same time period in the mth task chain. The specific quantitative formula is as follows:

[0040]

[0041] In the formula, Indicates the concurrency value of the nth task in the mth task chain. If the task is executed in parallel with other tasks, the number of parallel tasks is taken. Represents the total execution time of the mth task chain;

[0042] Calculate the cumulative delay index of each task chain. The specific calculation formula is as follows:

[0043]

[0044] Where, CDI m is the cumulative delay index of the mth task chain.

[0045] Preferably, a task chain cumulative risk assessment model is constructed based on the task chain impact coefficient and cumulative delay index of each generated task chain to generate the cumulative risk assessment coefficient of each task chain, specifically including the following steps:

[0046] Collect several task chain impact coefficients, several cumulative delay indexes and corresponding cumulative risk assessment coefficients of the mth task chain generated in the past period of time, and calibrate them as and x represents the number of several task chain impact coefficients, several cumulative delay indexes and corresponding cumulative risk assessment coefficients of the mth task chain generated in the past period of time, x=3, 4, 5, ..., h, h is a positive integer, and the data collected in the past period of time is formed into a historical data set;

[0047] The multivariate regression model is selected as the task chain cumulative risk assessment model, and is trained through historical data sets to determine the value of the regression coefficient according to the formula:

[0048]

[0049] Where β0, β1 and β2 are regression coefficients;

[0050] By minimizing the error between the predicted value and the actual value, the regression coefficient is optimized, and finally the values ​​of the regression coefficients β0, β1 and β2 are determined;

[0051] Using the finalized regression coefficient, the constructed task chain cumulative risk assessment model is input with the task chain impact coefficient TCIC of the mth task chain generated in real time. m and cumulative delay index CDI m , real-time generation of the cumulative risk assessment coefficient RF of the mth task chain m .

[0052] Preferably, a predetermined cumulative risk assessment coefficient threshold interval [RF one , RF two ], and the cumulative risk assessment coefficient RF of each task chain generated m With this threshold interval [RF one , RF two ] to compare, evaluate the cumulative risk level of each task chain dependency during the project execution according to the comparison results, and classify the task chains in the project according to the evaluation results, identify high-risk, medium-risk and low-risk task chains, and generate corresponding risk assessment reports. The specific comparison and analysis are as follows:

[0053] If RF m <RF one , if the cumulative risk of the task chain dependency is low risk, then the task chain is a low-risk task chain, and a low-risk assessment report is generated;

[0054] If RF one ≤RF m ≤RF two , if the cumulative risk of the task chain dependency is medium risk, then the task chain is a medium risk task chain, and a medium risk assessment report is generated;

[0055] If RF m >RFtwo , if the cumulative risk of the task chain dependency is high risk, then the task chain is a high-risk task chain and a high-risk assessment report is generated.

[0056] Preferably, based on the risk assessment report, it is recommended that managers take corresponding risk mitigation measures for task chains of different risk levels, specifically including the following steps:

[0057] For low-risk assessment reports, managers are advised to maintain current resource allocation and conduct routine monitoring as planned without making additional adjustments;

[0058] For medium-risk assessment reports, managers are advised to increase the frequency of monitoring, conduct regular reviews of the progress of the task chain, and prepare emergency plans in advance;

[0059] For high-risk assessment reports, managers are advised to prioritize resource allocation, adjust task priorities, implement continuous monitoring of the task chain, record and regularly provide feedback on task execution, and adjust the overall project schedule in real time based on task execution.

[0060] Preferably, a project execution risk assessment system based on data mining comprises a task relationship definition and data collection module, a comprehensive data information acquisition and storage module, a task chain dependency analysis and risk assessment module, a risk mitigation measure suggestion module, and a real-time monitoring and dynamic risk adjustment module;

[0061] Task relationship definition and data collection module: At the beginning of the project, determine the key tasks of the project and their interrelationships, configure the data collection system, define information collection rules, and determine the data information structure for subsequent analysis;

[0062] A comprehensive data information acquisition and storage module acquires comprehensive data information related to project execution and stores it after acquisition;

[0063] The task chain dependency analysis and risk assessment module analyzes the stored comprehensive data information, evaluates the cumulative risk of each task chain dependency during the project execution process, and classifies each task chain in the project according to the evaluation results, identifies high-risk, medium-risk and low-risk task chains, and generates corresponding risk assessment reports;

[0064] The risk mitigation measures recommendation module recommends that managers take corresponding risk mitigation measures for task chains with different risk levels based on the risk assessment report;

[0065] The real-time monitoring and dynamic risk adjustment module updates comprehensive data information in real time during project execution, regularly recalculates the cumulative risk assessment coefficient, and dynamically adjusts the risk assessment results.

[0066] In the above technical solution, the technical effects and advantages provided by the present invention are:

[0067] 1. The present invention can accurately assess the cumulative risk of each task chain in the project by modeling the key tasks and their interrelationships in detail during the project execution process, and comprehensively analyzing multi-dimensional data such as task execution time, path dependency, critical path, delay, task chain length and concurrency. In particular, in scenarios where there are complex dependencies between multiple tasks, the system can adjust the risk assessment results in real time by using a multivariate regression model and a dynamic data update mechanism, ensuring that managers can obtain accurate risk warning information in a timely manner, thereby making scientific decisions and reducing the possibility of project delays or failures due to untimely risk identification.

[0068] 2. The present invention generates assessment reports of high, medium and low risk levels, and makes specific management suggestions according to different risk levels, so that managers can adjust resource allocation and monitoring efforts in a targeted manner. Low-risk task chains use conventional monitoring, medium-risk task chains increase monitoring frequency and formulate emergency plans, and high-risk task chains prioritize resource allocation, adjust priorities and provide real-time feedback on execution. This classification management method not only optimizes the efficiency of project resource utilization, but also enhances the flexibility and operability of risk response, thereby better ensuring the timely completion and overall success of the project.

[0069] 3. The present invention can dynamically respond to changes and uncertainties in project progress by updating comprehensive data information in real time and recalculating risk assessment coefficients regularly. Through automatic data collection, processing and analysis, the system can make timely adjustments to the management measures of the task chain according to the latest risk assessment results, thereby improving the intelligence of project management and the accuracy of decision support. This dynamic adjustment mechanism greatly reduces the risks caused by unforeseen changes during project execution and improves the stability and success rate of project management. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0071] Figure 1 The present invention is a flowchart of a project execution risk assessment system and method based on data mining.

[0072] Figure 2 The module diagram of a project execution risk assessment system and method based on data mining of the present invention. DETAILED DESCRIPTION

[0073] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these example embodiments are provided so that the description of the present disclosure will be more comprehensive and complete, and the concept of the example embodiments will be fully conveyed to those skilled in the art.

[0074] The present invention provides Figure 1 A project execution risk assessment method based on data mining is shown, which specifically includes the following steps:

[0075] At the beginning of the project, determine the key tasks of the project and their interrelationships, configure the data collection system, define the information collection rules, and determine the data information structure for subsequent analysis;

[0076] At the beginning of the project, the system first needs to automatically identify and extract the key tasks in the project through project management tools or preset project templates. These key tasks are usually tasks that have a significant impact on the overall progress and success of the project, such as milestones, the generation of key deliverables, etc. By analyzing the project's work breakdown structure (WBS) or Gantt chart, the system can identify these key tasks and their priorities in the project. Next, the system will analyze the interdependencies between tasks to determine which tasks are prerequisites and which tasks can only be started after the predecessor tasks are completed. This process can be completed automatically through algorithms to ensure that all key tasks and their dependencies are accurately identified, laying the foundation for subsequent risk assessments.

[0077] Once the key tasks and dependencies are identified, the system will automatically configure the data collection system to ensure that the data of these key tasks can be monitored in real time during the project execution. The process of configuring the data collection system includes: selecting appropriate sensors, loggers or project management software interfaces to obtain real-time data on task execution; setting the frequency and scope of data collection to ensure that the status and progress of key tasks can be fully captured; and establishing a connection channel between data collection and storage to ensure that all collected data can be stored in the system securely and efficiently. This configuration process is completed automatically by the software, ensuring that the data collection system can be seamlessly integrated into the project management process and can adapt to the scale and complexity of different projects.

[0078] In order to ensure the standardization and accuracy of data collection, the system needs to predefine a set of information collection rules. These rules include: clarifying which data needs to be collected (such as task start time, completion time, resource usage, etc.), and how to handle abnormal data (such as task delays, resource conflicts, etc.). In addition, the system needs to define the structure of data information for subsequent analysis. The data information structure includes the format, field name, data type, etc. of the data to ensure that all collected information can be processed and analyzed in a standardized manner. This step is to ensure the consistency and integrity of the data, so that subsequent risk analysis can be based on reliable data, and it also facilitates the system to conduct large-scale data management and analysis.

[0079] Acquire comprehensive data information related to project execution and store it after acquisition;

[0080] In this embodiment, the comprehensive data information includes task execution time data, task dependency path data, task critical path data, task delay data, task chain length data and task chain concurrency data. The comprehensive data information is acquired from the project management system, the real-time monitoring platform and the task log through the configured automatic data acquisition module, and is stored in a data storage unit for subsequent analysis after acquisition.

[0081] Comprehensive data information refers to quantitative data related to project tasks collected by the automation system during the project execution process. These data include: task execution time data (the difference between the actual start and completion time of the task and the planned time, used to analyze the progress deviation of the task), task dependency path data (the relationship between tasks, especially the dependency path between the predecessor task and the successor task, used to understand the mutual influence between tasks), task critical path data (task data on the critical path, these tasks have a direct impact on the overall completion time of the project), task delay data (the actual delay time of the task, used to evaluate the degree of delay and possible risk propagation), task chain length data (the length of the path from a task to the end of the project, measuring the position and influence of the task in the entire project), and task chain concurrency data (reflecting the simultaneous execution of multiple task chains in the same time period, used to analyze the impact of concurrent tasks on resources and time).

[0082] These comprehensive data information can be automatically obtained from different project management tools and platforms through the configured automatic data collection module. Specifically, project management systems (such as JIRA, Microsoft Project) provide task execution time data, task dependency path data, and critical path data, which are usually automatically collected through API interfaces or direct data export functions; real-time monitoring platforms (such as IoT-based sensor networks or online monitoring tools) are used to collect task delay data and task chain length data in real time, which are synchronized to the system regularly through real-time monitoring modules; task logs (such as daily work reports or automatically generated operation logs) record detailed task execution status and concurrent task information, and automatically extract task chain concurrency data through text analysis or log parsing tools. All these data are connected to different data sources and synchronized to the central system through the automatic collection module, without manual intervention, ensuring the accuracy and timeliness of the data.

[0083] After data acquisition, the comprehensive data information will be automatically stored in a specially designed data storage unit, which is usually a structured database or data warehouse that can support efficient query and analysis. The storage process includes data cleaning (removing outliers and duplicate data to ensure data quality), data formatting (converting data from different sources into a unified format for subsequent analysis), and data indexing (establishing indexes based on key fields such as task ID, timestamp, and task path to ensure fast access to data). Through these steps, the data is systematically stored in a unified database and provides a basis for subsequent risk assessment analysis. The storage process is automatically completed through a database management system (such as MySQL, PostgreSQL) to ensure data persistence and integrity.

[0084] Analyze the stored comprehensive data information, evaluate the cumulative risk of dependencies of various task chains during project execution, and classify the various task chains in the project according to the evaluation results, identify high-risk, medium-risk and low-risk task chains, and generate corresponding risk assessment reports;

[0085] In this embodiment, the stored comprehensive data information is analyzed to evaluate the cumulative risk of dependencies of various task chains during project execution, and the various task chains in the project are classified according to the evaluation results to identify high-risk, medium-risk and low-risk task chains, and generate a corresponding risk assessment report, which specifically includes the following steps:

[0086] Extracting task execution time data, task dependency path data, task critical path data, task delay data, task chain length data, and task chain concurrency data of each task chain from the data storage unit;

[0087] To extract the relevant data of each task chain from the data storage unit, this can be achieved through the configured automatic data extraction module. This module can automatically extract data from the data storage unit through database query, API interface call and batch data export. First, the system will generate the corresponding query statement according to the preset task chain identifier and data type (such as task execution time, task dependency path, task critical path, etc.), and execute the query operation through the database management system (such as SQL database) to extract the required quantitative data. For real-time monitoring data, the system can directly obtain it from the stored real-time data stream by calling the API interface, and automatically organize and structure the data before extracting it. All extracted data will be automatically aggregated into an intermediate cache area for subsequent preprocessing and analysis operations to ensure the integrity and consistency of the data.

[0088] Preprocess the extracted data of each task chain;

[0089] The purpose of preprocessing is to ensure that the data of each task chain extracted from the data storage unit is consistent, accurate and available in subsequent analysis. Preprocessing usually includes the following steps: data cleaning, format standardization and outlier processing. Data cleaning is to automatically detect and delete or correct missing, incomplete or duplicate data through software algorithms to ensure the integrity of the data. Format standardization involves converting data from different sources or formats into a standardized format (such as a unified timestamp format and consistent numerical units), which enables different data types to be compared and integrated under the same analysis framework. Outlier processing automatically identifies data points that may deviate from the normal range through statistical analysis algorithms, and selects to correct or eliminate these outliers according to the set rules to prevent them from misleading the subsequent analysis results. In software implementation, these preprocessing steps are usually integrated through automated data processing modules to ensure that the data has been cleaned and standardized before entering the next step of analysis, with high accuracy and consistency.

[0090] Analyze the preprocessed data to generate task chain influence coefficients and cumulative delay indexes for each task chain;

[0091] A task chain cumulative risk assessment model is constructed based on the task chain impact coefficient and cumulative delay index of each task chain generated, and a cumulative risk assessment coefficient of each task chain is generated;

[0092] Determine the pre-set cumulative risk assessment coefficient threshold interval, and compare the generated cumulative risk assessment coefficient of each task chain with the threshold interval. According to the comparison results, evaluate the cumulative risk level of each task chain dependency during the project execution process, and classify the task chains in the project according to the evaluation results, identify high-risk, medium-risk and low-risk task chains, and generate a corresponding risk assessment report.

[0093] To determine the pre-set cumulative risk assessment coefficient threshold interval, historical project data analysis, expert experience value setting, and machine learning model training can be used. First, the system can analyze a large number of historical data of similar projects to calculate the distribution of the cumulative risk assessment coefficient of the task chain, and use statistical analysis methods such as percentiles or standard deviation ranges to preliminarily define high, medium, and low risk threshold intervals. Secondly, combined with the experience of project management experts and industry standards, these preliminarily defined thresholds are adjusted to better meet the risk management needs of actual projects. Finally, the system can also optimize and dynamically adjust the threshold intervals based on known risk assessment results and historical data by training machine learning models, so that it has higher prediction accuracy in future projects. These threshold intervals are stored in the system in a parameterized form and can be flexibly adjusted according to project characteristics and needs to ensure the accuracy and applicability of risk assessment.

[0094] In this embodiment, the logic for obtaining the task chain influence coefficient and the cumulative delay index of each task chain is as follows:

[0095] Extract the actual execution time and planned execution time of each task in the task execution time data of each task chain, and mark the actual execution time and planned execution time of each task in each task chain as and Indicates the actual execution time of the nth task in the mth task chain, represents the planned execution time of the nth task in the mth task chain, m = 1, 2, 3, ..., k, n = 1, 2, 3, ..., g, k and g are both positive integers;

[0096] In order to extract the actual execution time and planned execution time of each task in the task execution time data of each task chain, the data can be extracted from the project management system, task scheduling system and real-time monitoring platform through the configured automatic data collection module. First, the project management system (such as JIRA, Microsoft Project, etc.) usually records the planned execution time of the task (including the planned start time and planned end time of the task). These data can be automatically collected through API interface calls or direct export of data files in table form. The actual execution time can be obtained from the task scheduling system or real-time monitoring platform, which contains the actual start and end time of the task. The data collection module automatically captures this data regularly or on demand through pre-set rules and task identifiers, and organizes it into structured data, which is uniformly stored in the data storage unit for subsequent analysis. In this way, the extraction process is fully automated, which can ensure the accuracy and real-time nature of the data.

[0097] Compare the actual execution time of each task in each task chain with the planned execution time, and calculate the time deviation of each task. The specific calculation formula is as follows:

[0098]

[0099] In the formula, is the time deviation of the nth task in the mth task chain;

[0100] When calculating the time deviation of each task in the task chain, the execution deviation of the task is obtained by comparing the actual execution time of each task with the planned execution time. Specifically, the time deviation is the actual execution time Planned execution time The time deviation not only reflects whether the task is completed as planned, but also directly affects the overall progress of the task chain. It should be noted that the time deviation here is actually closely related to the task delay data. Task delay data usually indicates the delay of the actual start or completion time of the task relative to the plan, while the time deviation further quantifies the manifestation of this delay in specific time. Therefore, time deviation It can be regarded as a concretization of task delay data. Through this calculation, the impact of each task on the overall execution of the task chain can be more intuitively evaluated, providing a data basis for the subsequent calculation of the cumulative delay index and task chain impact coefficient.

[0101] Extract the number of dependent tasks of each task in the task dependency path data of each task chain, and set the target number of dependent tasks of each task in the task dependency path data of each task chain as It represents the number of tasks that the nth task in the mth task chain directly depends on. The specific quantitative formula is as follows:

[0102]

[0103] In the formula, Indicates whether the nth task in the mth task chain depends on the pth task. If there is a dependency, Take 1, otherwise take 0, p = 1, 2, 3, ..., g, g is the total number of tasks in the task chain;

[0104] In the dependency relationship of the task chain, the number of dependent tasks of each task can be determined by extracting the task dependency path data. Indicates the number of other tasks that the nth task in the mth task chain directly depends on. When calculating specifically, the dependency matrix is introduced, where when the nth task depends on the pth task, otherwise, By traversing all possible task dependencies (i.e., p = 1, 2, 3, ..., g), the number of directly dependent tasks of the task can be accumulated. This calculation method can accurately quantify the dependencies of tasks in the task chain and provide a basis for the subsequent calculation of the task chain impact coefficient. This description clearly explains how to determine the number of dependencies of tasks through the dependency matrix and combine it with other key factors in the task chain for overall analysis.

[0105] Extract the critical path data of each task chain and mark it as Indicates whether the nth task in the mth task chain is on the critical path. When the task is on the critical path, Take 1, otherwise take 0. At the same time, extract the resource consumption of each task and the total resource consumption of all tasks in the project and mark them as and R total , Indicates the resource consumption of the nth task in the mth task chain;

[0106] To extract the critical path data of the task chain and the resource consumption of the task, you can obtain it from the project management system and resource management system through the configured automatic data collection module. First, the critical path data is usually generated through the built-in critical path analysis function of the project management system (such as Primavera, Microsoft Project, etc.). The system will identify which tasks are on the critical path of the project. The data of these tasks can be automatically extracted through the API interface or data export function and marked as When a task is on the critical path, the system sets its value to 1, otherwise it is set to 0. At the same time, resource consumption data It can be extracted through a resource management system (such as an ERP system). The system records the resources consumed by each task during execution (such as man-hours, funds, equipment usage, etc.) and can be directly obtained through a data interface. The total resource consumption of all tasks in the project is R total It can also be calculated and extracted in the resource management system. After being uniformly processed and stored, these data can be used for subsequent analysis and calculation to ensure the accuracy and timeliness of the data. Extracting data in this automated way can effectively support the calculation of the task chain impact coefficient and the cumulative delay index.

[0107] Calculate the comprehensive weight of each task in each task chain. The specific quantitative formula is as follows:

[0108]

[0109] In the formula, is the comprehensive weight of the nth task in the mth task chain;

[0110] The combined weight of each task in the task chain is calculated When considering whether a task is on the critical path and the relative proportion of the task's resource consumption in the project, it is necessary to consider whether the task is on the critical path and the relative proportion of the task's resource consumption in the project. The critical path system and resource consumption ratio Jointly determined. Critical path coefficient The critical path analysis function of the project management system determines that when a task is on the critical path, the value is 1, otherwise it is 0. The resource consumption ratio indicates the proportion of the resource consumption of the task in the resource consumption of all tasks in the project. This data is usually extracted from the resource management system. By combining the position of the task in the critical path and the importance of its resource consumption, the comprehensive weight of the task can be obtained. This quantifies the overall impact of the task in the task chain. This comprehensive weight provides key input data for the subsequent task chain impact coefficient calculation.

[0111] Calculate the task chain influence coefficient of each task chain. The specific calculation formula is as follows:

[0112]

[0113] Where TCIC m is the task chain influence coefficient of the mth task chain;

[0114] The task chain influence coefficient TCIC of the mth task chain m The size of directly reflects the degree of influence of the task chain on the overall progress during the project execution, and is closely related to the cumulative risk level of the dependencies of each task chain during the project execution. Specifically, a larger task chain impact coefficient indicates that there is a larger time deviation, higher resource consumption, and more critical path tasks in the task chain, which means that the task chain has a stronger dependency and has a more significant impact on the overall progress of the project. Therefore, the larger the task chain impact coefficient, the higher the probability that the task chain will be judged as high risk in the cumulative risk assessment. In the risk assessment process, based on the task chain impact coefficient, high-risk, medium-risk, and low-risk task chains can be effectively identified and classified, thereby providing managers with an accurate basis for risk level assessment.

[0115] Extract the task chain length data of each task chain and mark it as L m , L m It represents the number of path tasks from the starting task to the end task of the mth task chain. The specific quantitative formula is as follows:

[0116]

[0117] In the formula, Indicates whether the nth task in the mth task chain is on the path. When the task is on the path, Takes 1, otherwise takes 0, g is the total number of tasks in the task chain;

[0118] When extracting task chain length data L m The key is to identify and calculate the number of path tasks from the starting task to the end task. The specific method is to automatically identify all path tasks in the task chain through the path analysis function of the project management system and mark whether these tasks are on the path. For each task n in the mth task chain, the path data It is used to indicate whether the task is a task on the path. When the task is on the path, Otherwise, it is 0. By traversing all tasks in the task chain, the total number of path tasks L is accumulated. m This process can usually be achieved through the path tracking tool in the project management system. The system can automatically identify and record the key tasks in the path and generate the path length based on this data. In this way, the path task data is extracted to ensure the accuracy of the data, providing a reliable basis for the subsequent calculation of the cumulative delay index and task chain impact coefficient.

[0119] Extract the task chain concurrency data of each task chain and mark it as C m , C m It represents the number of tasks executed in parallel in the same time period in the mth task chain. The specific quantitative formula is as follows:

[0120]

[0121] In the formula, Indicates the concurrency value of the nth task in the mth task chain. If the task is executed in parallel with other tasks, the number of parallel tasks is taken. Represents the total execution time of the mth task chain;

[0122] When extracting the concurrency data C of the task chain m The key is to identify the number of tasks that are executed in parallel in the same time period in the task chain. In the specific calculation process, the time execution records of each task can be obtained through the project scheduling system or task management platform to determine the start and end time of the task. For the nth task in the mth task chain, its concurrency Indicates the number of other tasks executed in parallel with this task during the execution period. The data can be extracted from the scheduling system through the automated data collection module. The system will automatically calculate the number of parallel tasks based on the time records of the tasks and summarize the number of all parallel tasks. Then, by normalizing these data and combining them with the total execution time of the task chain, The concurrency C of the task chain can be obtained m This data extraction method ensures the accuracy of concurrency calculation and provides key support for the subsequent cumulative delay index calculation.

[0123] Calculate the cumulative delay index of each task chain. The specific calculation formula is as follows:

[0124]

[0125] Where, CDI m is the cumulative delay index of the mth task chain.

[0126] The cumulative delay index CDI of the mth task chain m The size of reflects the cumulative effect of delayed tasks in the task chain and the impact of task concurrency on the overall progress, which is directly related to the cumulative risk level of the dependencies of each task chain during the project execution. Specifically, a higher cumulative delay index indicates that there are more delayed tasks in the task chain, and these tasks may be executed under high concurrency, thereby increasing the overall delay risk of the task chain. Therefore, the larger the cumulative delay index, the higher the dependency and delay risk of the task chain in the cumulative risk assessment, and it may be assessed as a high-risk task chain. In the risk assessment of project execution, by analyzing the cumulative delay index, it is possible to more accurately identify which task chains have a potential significant impact on the overall progress of the project, thereby providing managers with targeted risk mitigation measures.

[0127] In this embodiment, a task chain cumulative risk assessment model is constructed based on the task chain impact coefficient and cumulative delay index of each task chain generated to generate a cumulative risk assessment coefficient of each task chain, specifically including the following steps:

[0128] Collect several task chain impact coefficients, several cumulative delay indexes and corresponding cumulative risk assessment coefficients of the mth task chain generated in the past period of time, and calibrate them as and x represents the number of several task chain impact coefficients, several cumulative delay indexes and corresponding cumulative risk assessment coefficients of the mth task chain generated in the past period of time, x=3, 4, 5, ..., h, h is a positive integer, and the data collected in the past period of time is formed into a historical data set;

[0129] To collect the task chain impact coefficient, cumulative delay index and corresponding cumulative risk assessment coefficient of each task chain generated in the past period of time, it can be achieved by configuring the automated data recording and monitoring module. Specifically, these data are usually generated in real time by the project management system, risk assessment platform or task scheduling system during the project execution, and are automatically stored in the central database or data warehouse. The data collection process can be carried out through regular scheduling tasks or triggered event-driven methods. The system will automatically extract relevant data of all historical task chains within a specified time interval, including task chain impact coefficient, cumulative delay index and cumulative risk assessment coefficient. After these data are automatically obtained through the data interface or API, they are organized into structured data tables to ensure that they can be directly called and batch processed during model training and regression analysis. This method ensures the accuracy and timeliness of data collection and provides a reliable data basis for subsequent analysis.

[0130] The value of x is greater than or equal to 3 because when performing multiple regression analysis, it is necessary to ensure that the model has sufficient dimensions and data points to achieve accurate regression fitting. Specifically, the value of x represents a number of task chain impact coefficients, cumulative delay indexes, and cumulative risk assessment coefficients generated in the past period of time. Selecting the setting of x≥3 can ensure that the data samples are rich enough when training the model, and the coefficients in the regression equation (such as β0, β1, and β2) can be fitted through multiple data points to generate a reliable regression model. In addition, considering that there are multiple variables involved in the equation, if the value of x is too small (for example, only 1 or 2), the predictive ability of the regression model will be insufficient, and it will be difficult to fully reflect the complex relationship between task chain risks and influencing factors. Therefore, selecting x≥3 can ensure the robustness and accuracy of the model while avoiding overfitting or underfitting problems caused by insufficient sample size.

[0131] The multivariate regression model is selected as the task chain cumulative risk assessment model, and is trained through historical data sets to determine the value of the regression coefficient according to the formula:

[0132]

[0133] Where β0, β1 and β2 are regression coefficients;

[0134] The reason for choosing the multivariate regression model as the cumulative risk assessment model for task chains is that the model can simultaneously consider the combined impact of multiple independent variables on the cumulative risk of dependencies. The multivariate regression model constructs a linear function by analyzing the relationship between these variables and the risk assessment coefficient to predict the cumulative risk level of each task chain. The advantage of this model is that it can handle complex multivariate relationships, and through training with historical data, the model can automatically adjust the weights of each variable (i.e., regression coefficients) to generate accurate risk assessment results. The multivariate regression model is a method in statistics that determines the best-fitting linear relationship by minimizing the error between the predicted value and the actual value. It is widely used in risk prediction, trend analysis and other fields.

[0135] The three regression coefficients β0, β1, and β2 represent the constant term of the model and the weights of the independent variables, respectively. Specifically, β0 is the intercept of the regression equation, which indicates the basic value of the cumulative risk assessment coefficient without any input of the task chain influence coefficient and the cumulative delay index; β1 is the weight of the task chain influence coefficient, which reflects the contribution of the task chain influence coefficient to the cumulative risk; β2 is the weight of the cumulative delay index, which indicates the impact of the delay on the cumulative risk of the task chain. Through training with historical data, the model can automatically adjust these coefficients, making the predicted cumulative risk assessment coefficient more accurate, and ultimately helping managers effectively assess and respond to potential risks in projects.

[0136] By minimizing the error between the predicted value and the actual value, the regression coefficient is optimized, and finally the values ​​of the regression coefficients β0, β1 and β2 are determined;

[0137] The purpose of optimizing the regression coefficient by minimizing the error between the predicted value and the actual value is to ensure that the regression model predicts the cumulative risk of the task chain more accurately. Specifically, the regression coefficients β0, β1, and β2 determine the weights of the task chain impact coefficient and the cumulative delay index in the final cumulative risk assessment. In order to obtain the best regression coefficient, it is necessary to train the model with historical data so that the error between the risk assessment value predicted by the model and the actual historical risk data is minimized. This process can be achieved through the iterative optimization algorithm in the software. Common methods include the least squares method (OLS) and the gradient descent method. The software will automatically adjust the regression coefficient so that the error function (usually the sum of squared errors) is minimized. This process is usually implemented in data analysis tools (such as Python's Scikit-learn library, R language, or SPSS). The system will continuously optimize the coefficients in multiple iterations and finally determine the optimal β0, β1, and β2 values ​​to ensure the reliability and accuracy of the model in actual project risk prediction.

[0138] Using the finalized regression coefficient, the constructed task chain cumulative risk assessment model is input with the task chain impact coefficient TCIC of the mth task chain generated in real time. m and cumulative delay index CDI m , real-time generation of the cumulative risk assessment coefficient RF of the mth task chain m .

[0139] In this embodiment, the predetermined cumulative risk assessment coefficient threshold interval [RF one , RF two ], and the cumulative risk assessment coefficient RF of each task chain generated m With this threshold interval [RF one , RF two Compare and evaluate the cumulative risk level of each task chain dependency during the project execution according to the comparison results. According to the evaluation results, classify the task chains in the project, identify high-risk, medium-risk and low-risk task chains, and generate corresponding risk assessment reports. The specific comparison and analysis are as follows:

[0140] If RF m <RF one , if the cumulative risk of the task chain dependency is low risk, then the task chain is a low-risk task chain and a low-risk assessment report is generated, which means that the dependency and delay of the task chain have little impact on the overall progress of the project. Usually, the delay and dependency of such task chains will not cause serious chain reactions. Managers do not need to take additional intervention measures and only need to monitor according to the normal process. The corresponding risk assessment report will mark the task chain as "low risk" and provide basic information and monitoring suggestions for the task chain, prompting managers to allocate more resources to high-risk task chains;

[0141] If RF one ≤RF m ≤RF two , if the cumulative risk of the task chain dependency is medium risk, then the task chain is a medium-risk task chain, and a medium-risk assessment report is generated, indicating that the dependency and delay of the task chain have a certain degree of impact on the project, which may produce a chain effect under certain conditions. In this case, managers are required to monitor the task chain more closely and formulate countermeasures in advance according to the progress of the project. The generated medium-risk assessment report will list the specific risk points of the task chain, the possible impacts, and the recommended mitigation measures to help managers make preventive adjustments at key nodes;

[0142] If RF m >RF two, if the cumulative risk of the dependency of the task chain is high risk, then the task chain is a high-risk task chain, and a high-risk assessment report is generated, indicating that the dependency and delay of the task chain pose a major threat to the overall progress of the project, and there is a large risk of chain reaction. This type of task chain is usually the key to project success. Delays or problems may cause serious deviations in the entire project. The generated high-risk assessment report will list in detail the key risks of the task chain and possible serious consequences, and strongly recommend that managers take emergency measures, such as prioritizing resource allocation, adjusting task priorities, etc., to ensure that the task chain can be completed on time and avoid irreversible impacts on the overall project.

[0143] Based on the risk assessment report, managers are advised to take appropriate risk mitigation measures for task chains with different risk levels;

[0144] In this embodiment, based on the risk assessment report, it is recommended that the manager take corresponding risk mitigation measures for task chains of different risk levels, which specifically include the following steps:

[0145] For low-risk assessment reports, managers are advised to maintain current resource allocation and conduct routine monitoring as planned without making additional adjustments;

[0146] For task chains that are assessed as low risk, the current resource allocation can be maintained through the project management system, and task progress can be tracked through the regular monitoring module. The specific operation method is that the management system uses the default monitoring frequency in the monitoring settings of low-risk task chains. The system will automatically collect and record the execution data of the task chain as planned, such as progress, resource consumption, etc., and generate regular status reports. These reports do not need to be specially processed, they only need to be archived for subsequent inspection. This method ensures the rational use of resources while reducing the management burden.

[0147] For medium-risk assessment reports, managers are advised to increase the frequency of monitoring, conduct regular reviews of the progress of the task chain, and prepare emergency plans in advance;

[0148] For task chains that are assessed as medium risk, the risk monitoring module in the project management system can be used to increase the monitoring frequency and set up an automatic warning mechanism. The specific method is to adjust the monitoring parameters in the task chain progress monitoring module. The system will collect and analyze task data more frequently and trigger warnings when possible delays or resource shortages are detected. At the same time, managers can use the emergency plan management module to formulate response plans in advance, such as resource allocation, task rescheduling, etc. The system will automatically execute these emergency measures according to predetermined rules to ensure that the project can be adjusted in time when problems arise.

[0149] For high-risk assessment reports, managers are advised to prioritize resource allocation, adjust task priorities, implement continuous monitoring of the task chain, record and regularly provide feedback on task execution, and adjust the overall project schedule in real time based on task execution.

[0150] For task chains that are assessed as high-risk, resources can be allocated preferentially and task priorities can be adjusted in real time through the resource optimization and dynamic scheduling modules in the project management system. The specific operation is that after risk identification, the system will automatically adjust the execution order of tasks according to the preset priority rules, and dynamically allocate resources through the resource management module to ensure the smooth completion of high-risk tasks. At the same time, the system will turn on the continuous monitoring mode, record the execution of tasks in real time, and make dynamic adjustments through the data feedback mechanism. When new problems are detected, the project progress will be recalculated and a new schedule will be generated. The entire process can achieve automated adjustment and real-time feedback to ensure that high-risk task chains are protected to the greatest extent.

[0151] During the project execution, comprehensive data information is updated in real time, and the cumulative risk assessment coefficient is recalculated regularly to dynamically adjust the risk assessment results.

[0152] During the project execution process, real-time updating of comprehensive data information and regular recalculation of cumulative risk assessment coefficients can be achieved by configuring the automatic data collection and real-time analysis module in the project management system. Specifically, the data collection module continuously obtains the latest key data such as task execution time, dependency path, resource consumption, etc. from the project management system, task scheduling platform and resource management system. The system will update these comprehensive data information in real time according to the preset time interval or trigger event (such as task status change, progress delay, etc.). Subsequently, the risk analysis module in the system will use the latest data to recalculate the task chain impact coefficient and cumulative delay index of the task chain, and dynamically adjust the cumulative risk assessment coefficient on this basis. The updated risk assessment results will automatically generate a report to prompt managers of changes in the risk level of the project so that they can take corresponding adjustment measures in time. The significance of this is that it can ensure that the project is always in a real-time monitoring state during the execution process, and dynamically optimize the risk assessment model according to the actual situation, so as to effectively deal with uncertain factors, minimize the risk of project failure, and improve the accuracy and flexibility of project management.

[0153] like Figure 2 A project execution risk assessment system based on data mining is shown, comprising a task relationship definition and data collection module, a comprehensive data information acquisition and storage module, a task chain dependency analysis and risk assessment module, a risk mitigation measure recommendation module, and a real-time monitoring and dynamic risk adjustment module;

[0154] Task relationship definition and data collection module: At the beginning of the project, determine the key tasks of the project and their interrelationships, configure the data collection system, define information collection rules, and determine the data information structure for subsequent analysis;

[0155] A comprehensive data information acquisition and storage module acquires comprehensive data information related to project execution and stores it after acquisition;

[0156] The task chain dependency analysis and risk assessment module analyzes the stored comprehensive data information, evaluates the cumulative risk of each task chain dependency during the project execution process, and classifies each task chain in the project according to the evaluation results, identifies high-risk, medium-risk and low-risk task chains, and generates corresponding risk assessment reports;

[0157] The risk mitigation measures recommendation module recommends that managers take appropriate risk mitigation measures for task chains at different risk levels based on the risk assessment report;

[0158] The real-time monitoring and dynamic risk adjustment module updates comprehensive data information in real time during project execution, regularly recalculates the cumulative risk assessment coefficient, and dynamically adjusts the risk assessment results.

[0159] The above formulas are all dimensionless and numerical calculations. The formula is a formula for the most recent real situation obtained by collecting a large amount of data and performing software simulation. The preset parameters in the formula are set by technicians in this field according to actual conditions.

[0160] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website site, computer, server or data center to another website site, computer, server or data center by wired or wireless (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center containing one or more available media sets. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state hard disk.

[0161] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0162] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0163] In the several embodiments provided in the present application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0164] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0165] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0166] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A project execution risk assessment method based on data mining, characterized in that: The specific steps include: At the beginning of the project, determine the key tasks of the project and their interrelationships, configure the data collection system, define the information collection rules, and determine the data information structure for subsequent analysis; Acquire comprehensive data information related to project execution and store it after acquisition; The comprehensive data information is obtained from the project management system, the real-time monitoring platform and the task log through the configured automatic data acquisition module, and is stored in a data storage unit for subsequent analysis after being obtained; Extracting data of each task chain from the data storage unit; Preprocess the extracted data of each task chain; Analyze the preprocessed data to generate task chain influence coefficients and cumulative delay indexes for each task chain; The details are as follows: The actual execution time and planned execution time of each task in the task execution time data of each task chain are extracted and marked as and Indicates the actual execution time of the nth task in the mth task chain, represents the planned execution time of the nth task in the mth task chain, m = 1, 2, 3, ..., k, n = 1, 2, 3, ..., g, k and g are both positive integers; Compare the actual execution time of each task in each task chain with the planned execution time, and calculate the time deviation of each task. The specific calculation formula is as follows: In the formula, is the time deviation of the nth task in the mth task chain; Extract the number of dependent tasks of each task in the task dependency path data of each task chain and mark it as It represents the number of tasks that the nth task in the mth task chain directly depends on. The formula is as follows: In the formula, Indicates whether the nth task in the mth task chain depends on the pth task. If there is a dependency, Take 1, otherwise take 0, p = 1, 2, 3, ..., g, g is the total number of tasks in the task chain; Extract the critical path data of each task chain and mark it as Indicates whether the nth task in the mth task chain is on the critical path. When the task is on the critical path, Take 1, otherwise take 0. At the same time, extract the resource consumption of each task and the total resource consumption of all tasks in the project and mark them as and R total , Indicates the resource consumption of the nth task in the mth task chain; Calculate the comprehensive weight of each task in each task chain, the formula is as follows: In the formula, is the comprehensive weight of the nth task in the mth task chain; Calculate the task chain influence coefficient of each task chain. The specific calculation formula is as follows: In the formula, TCIC m is the task chain influence coefficient of the mth task chain; Extract the task chain length data of each task chain and mark it as L m , L m It represents the number of path tasks from the starting task to the end task of the mth task chain. The formula is as follows: In the formula, Indicates whether the nth task in the mth task chain is on the path. When the task is on the path, Takes 1, otherwise takes 0, g is the total number of tasks in the task chain; Extract the task chain concurrency data of each task chain and mark it as C m , C m It represents the number of tasks executed in parallel in the same time period in the mth task chain. The formula is as follows: In the formula, Indicates the concurrency value of the nth task in the mth task chain. If the task is executed in parallel with other tasks, the number of parallel tasks is taken. Represents the total execution time of the mth task chain; Calculate the cumulative delay index of each task chain. The specific calculation formula is as follows: Where, CDI m is the cumulative delay index of the mth task chain; A task chain cumulative risk assessment model is constructed based on the task chain impact coefficient and cumulative delay index of each generated task chain to generate a cumulative risk assessment coefficient for each task chain; Determine a preset cumulative risk assessment coefficient threshold interval, and compare the generated cumulative risk assessment coefficients of each task chain with the threshold interval, identify high-risk, medium-risk and low-risk task chains, and generate corresponding risk assessment reports; Based on the risk assessment report, managers are advised to take appropriate risk mitigation measures for task chains with different risk levels; During the project execution, comprehensive data information is updated in real time, and the cumulative risk assessment coefficient is recalculated regularly to dynamically adjust the risk assessment results.

2. According to the project execution risk assessment method based on data mining according to claim 1, it is characterized in that: A task chain cumulative risk assessment model is constructed based on the task chain impact coefficient and cumulative delay index of each task chain generated, and a cumulative risk assessment coefficient of each task chain is generated, specifically including the following steps: Collect several task chain impact coefficients, several cumulative delay indexes and corresponding cumulative risk assessment coefficients of the mth task chain generated in the past period of time, and calibrate them as and x represents the number of several task chain impact coefficients, several cumulative delay indexes and corresponding cumulative risk assessment coefficients of the mth task chain generated in the past period of time, x=3, 4, 5, ..., h, h is a positive integer, and the data collected in the past period of time is formed into a historical data set; The multivariate regression model is selected as the task chain cumulative risk assessment model, and is trained through historical data sets to determine the value of the regression coefficient according to the formula: Where β0, β1 and β2 are regression coefficients; By minimizing the error between the predicted value and the actual value, the regression coefficient is optimized, and finally the values ​​of the regression coefficients β0, β1 and β2 are determined; Using the finalized regression coefficient, the constructed task chain cumulative risk assessment model is input with the task chain impact coefficient TCIC of the mth task chain generated in real time. m and cumulative delay index CDI m , generate the cumulative risk assessment coefficient RF of the mth task chain in real time m .

3. A project execution risk assessment method based on data mining according to claim 2, characterized in that: Determine the preset cumulative risk assessment factor threshold range [RF one , RF two ], and the cumulative risk assessment coefficient RF of each task chain generated m With this threshold interval [RF one , RF two ] to compare, evaluate the cumulative risk level of each task chain dependency during the project execution according to the comparison results, and classify the task chains in the project according to the evaluation results, identify high-risk, medium-risk and low-risk task chains, and generate corresponding risk assessment reports. The specific comparison and analysis are as follows: If RF m <RF one , if the cumulative risk of the task chain dependency is low risk, then the task chain is a low-risk task chain, and a low-risk assessment report is generated; If RF one ≤RF m ≤RF two , if the cumulative risk of the task chain dependency is medium risk, then the task chain is a medium risk task chain, and a medium risk assessment report is generated; If RF m >RF two , if the cumulative risk of the task chain dependency is high risk, then the task chain is a high-risk task chain and a high-risk assessment report is generated.

4. A project execution risk assessment method based on data mining according to claim 3, characterized in that: According to the risk assessment report, managers are advised to take corresponding risk mitigation measures for task chains with different risk levels, including the following steps: For low-risk assessment reports, managers are advised to maintain current resource allocation and conduct routine monitoring as planned without making additional adjustments; For medium-risk assessment reports, managers are advised to increase the frequency of monitoring, conduct regular reviews of the progress of the task chain, and prepare emergency plans in advance; For high-risk assessment reports, managers are advised to prioritize resource allocation, adjust task priorities, implement continuous monitoring of the task chain, record and regularly provide feedback on task execution, and adjust the overall project schedule in real time based on task execution.

5. A project execution risk assessment system based on data mining, used to implement a project execution risk assessment method based on data mining as described in any one of claims 1 to 4, characterized in that: It includes task relationship definition and data collection module, comprehensive data information acquisition and storage module, task chain dependency analysis and risk assessment module, risk mitigation measure recommendation module and real-time monitoring and dynamic risk adjustment module; Task relationship definition and data collection module: At the beginning of the project, determine the key tasks of the project and their interrelationships, configure the data collection system, define information collection rules, and determine the data information structure for subsequent analysis; A comprehensive data information acquisition and storage module acquires comprehensive data information related to project execution and stores it after acquisition; The task chain dependency analysis and risk assessment module analyzes the stored comprehensive data information, evaluates the cumulative risk of each task chain dependency during the project execution process, and classifies each task chain in the project according to the evaluation results, identifies high-risk, medium-risk and low-risk task chains, and generates corresponding risk assessment reports; The risk mitigation measures recommendation module recommends that managers take appropriate risk mitigation measures for task chains at different risk levels based on the risk assessment report; The real-time monitoring and dynamic risk adjustment module updates comprehensive data information in real time during project execution, regularly recalculates the cumulative risk assessment coefficient, and dynamically adjusts the risk assessment results.

Citation Information

Patent Citations

  • HIVE task execution engine selection method and system

    CN109634989A

  • Project management method and system based on WBS engineering entity splitting and hooking

    CN116934052A