Intelligent identification and early warning method for scientific and technological consultation project risk factors
By constructing a map of resource competition and data synchronization delays, we can identify risk factors in scientific and technological consulting projects, optimize resource allocation, solve the problems of resource competition and progress delays in existing technologies, and improve project management efficiency.
Patent Information
- Application Number
- CN202510758814.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-09-09
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In existing scientific and technological consulting projects, the risk factor identification methods such as computing resource competition, personnel task overlap and data interface blockage rely on manual experience or simple statistics, which cannot dynamically adapt to project progress, resulting in unbalanced resource allocation and delayed project progress.
By constructing a resource competition distribution feature map, identifying the causal path of model iteration delays, evaluating the degree of interference in multi-module parallel development, detecting the impact path of data synchronization delays, building a risk map, optimizing resource allocation plans, and formulating intervention measures.
Effectively identify and warn of risk factors in scientific and technological consulting projects, improve resource utilization efficiency and project management level, optimize resource allocation, and reduce schedule delays and collaboration conflicts.
Smart Images

Figure CN120611974A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and in particular to an intelligent identification and early warning method for risk factors of scientific and technological consulting projects. Background Art
[0002] In the field of scientific and technological consulting project management, research on how to effectively identify and manage risk factors is of paramount importance. This area is directly related to the efficient advancement of projects and the optimal allocation of resources, especially in complex environments involving multi-party collaboration. Risk management is crucial to ensuring project success. With the increasing reliance of scientific and technological projects on high-performance computing, data analysis, and intelligent tools, computing resources have become one of the core foundational resources in project implementation. Scientific and technological projects often involve extensive data processing, model training, and online collaboration, resulting in highly concentrated and time-sensitive computing resource demands. Providing accurate and scientifically sound scientific and technological project consulting advice to consulting clients requires identifying risk factors associated with competition for computing resources. However, current risk identification methods, which mostly rely on manual experience or simple statistical analysis, struggle to deeply explore the complex dependencies hidden in multi-party collaboration and are unable to dynamically adapt to sudden changes during project progress. This results in often delayed risk warnings and frequent imbalances in resource allocation. Against this backdrop, the core challenges facing project management are gradually becoming apparent. The primary challenge is identifying hidden dependency conflicts in multi-party collaboration. Due to the lack of transparent mapping between resource requirements and task schedules across different teams or departments, issues such as computing resource contention are often overlooked. For example, scheduling conflicts between algorithm and data teams regarding computing cluster usage directly lead to delays in task queuing. This resource contention further exacerbates the phenomenon of overlapping tasks, as team members often need to switch between multiple tasks, frequently disrupting their work flow and significantly reducing efficiency. Overlapping tasks in turn leads to deeper barriers to data collaboration. When data interface standards between different systems are inconsistent, delays in data synchronization become another bottleneck hindering overall project progress. These interconnected challenges form a complex web of project risk management. Therefore, how to accurately identify multiple dependency paths such as computing resource contention, overlapping tasks, and data interface blockages through intelligent means, and dynamically extract key risk factors that indicate resource allocation imbalance, has become a critical issue that needs to be addressed. Summary of the Invention
[0003] The present invention provides an intelligent identification and early warning method for risk factors of scientific and technological consulting projects, which mainly includes:
[0004] Obtain computing cluster usage records, training period reservation data, and departmental task allocation data in scientific research consulting projects, map cluster conflicts and training reservation data to graph nodes, and construct a resource competition distribution feature map that includes departmental collaboration conflicts;
[0005] Identify model iteration delay data in computing cluster usage records, combine the distribution characteristics of resource competition, analyze the causal path between training period reservations and model iteration delays, extract the peak GPU usage and delay duration for each period, and determine the time node corresponding to the delay;
[0006] Obtain multi-module parallel development task allocation data for delayed nodes, identify the correlation between task overlap and code submission interruptions, analyze the relationship between task switching frequency and submission interruptions through time series comparison, and assess the degree of interference of multi-module parallel development on progress;
[0007] When the interference level exceeds the threshold, the task priorities and resource requirements of each department are obtained, a task overlap dependency subgraph is constructed, and task allocation imbalance indicators are quantitatively analyzed;
[0008] Extract distributed data interface configuration information, combine it with task allocation imbalance indicators, detect inconsistencies between protocols and task allocation data, and identify the impact paths of data synchronization delays;
[0009] Analyze the dependency path between the impact path of data synchronization delay and the distribution characteristics of computing resource competition risk, construct a risk map that includes departmental collaboration conflicts and chain risk diffusion, and obtain the risk factor distribution;
[0010] Obtain historical collaboration data from various departments, combine it with the distribution of risk factors, analyze and optimize risk diffusion paths and node weights, determine the order of intervention for high-risk risk factors, optimize paths based on the degree of interference using knowledge graphs, analyze and calculate the dependencies of cluster conflicts, build a risk warning mechanism, formulate intervention measures, and output an optimized resource allocation plan.
[0011] The technical solution provided by the embodiment of the present invention may have the following beneficial effects:
[0012] The present invention discloses a method for intelligent identification and early warning of risk factors of scientific and technological consulting projects. By analyzing computing cluster usage records, training appointments and task allocation data, a resource competition distribution characteristic diagram is constructed, the causal path of model iteration delay is identified, and the degree of interference of multi-module parallel development on the progress is evaluated. When the interference exceeds the threshold, the task allocation imbalance index is analyzed, the data synchronization delay impact path is detected, and a map containing departmental collaboration conflicts and risk diffusion is constructed. Based on historical data, the risk diffusion path and node weights are optimized, the intervention order of high-risk risk factors is determined, and the knowledge graph is used to analyze dependencies such as computing cluster conflicts. A risk early warning mechanism is constructed and intervention measures are formulated, and finally an optimized resource allocation plan is output. The present invention can effectively solve problems such as resource competition, progress delays and collaboration conflicts in scientific research projects, improve resource utilization efficiency and project management level. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 The present invention provides a flow chart of an intelligent identification and early warning method for risk factors of scientific and technological consulting projects.
[0014] Figure 2 This is a schematic diagram of an intelligent identification and early warning method for risk factors of scientific and technological consulting projects according to the present invention. DETAILED DESCRIPTION
[0015] To further understand the content of the present invention, the present invention is described in detail with reference to the accompanying drawings and examples. The present application is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are intended only to illustrate the relevant invention and are not intended to limit the invention. It should also be noted that, for ease of description, only portions relevant to the invention are shown in the accompanying drawings.
[0016] like Figure 1-2 In this embodiment, a method for intelligently identifying and early warning risk factors of scientific and technological consulting projects may specifically include:
[0017] S101. Obtain computing cluster usage records, training period reservation data, and departmental task allocation data in scientific research consulting projects, map cluster conflicts and training reservation data to graph nodes, and construct a resource competition distribution feature map that includes departmental collaboration conflicts.
[0018] Obtain computing cluster usage records from scientific research consulting projects and extract resource utilization data for each department at different time periods. Count the number of CPU cores, memory capacity, and GPUs occupied by each department within each time window. Calculate the ratio of each department's resource utilization to the overall cluster load. Identify the intensity of resource competition between departments based on utilization differences. Construct a departmental collaboration matrix to record the frequency of conflicts between departments due to resource competition. Based on the conflict frequency data in the departmental collaboration matrix, analyze the task dependencies submitted by each department in the training period reservation records. Use a directed graph structure to represent the task dependency path between departments. Nodes represent department entities, edges represent task dependencies and resource transfer directions, and edge weights reflect the degree of collaborative conflict. If the conflict frequency between two departments exceeds a preset threshold, they are marked as a high-competition node pair in the dependency path graph. For the high-contention node pairs marked in the dependency path diagram, the overlap of their appointment time periods and the difference in allocation priorities are extracted. The time concentration of resource usage is measured by calculating the negative logarithm of the resource usage frequency in each time period as the time distribution entropy value. The distribution characteristics of resource competition in the time dimension are judged according to the size of the entropy value, and a resource competition distribution feature map is obtained, which includes the peak time period of resource occupancy, the frequency distribution of conflicts between departments, and the dispersion of task execution time.
[0019] Specifically, computing clusters in scientific research and consulting projects are shared resources, and resource competition often arises among various departments during their use. Calculating resource utilization involves a comprehensive evaluation of multiple indicators.
[0020] Specifically, CPU core utilization is calculated by dividing the number of cores currently used by the department by the total number of cores in the cluster. Memory utilization and GPU utilization are calculated similarly. This utilization data needs to be collected within a fixed time window, such as every hour or every four hours, to capture dynamic changes in resource usage.
[0021] In one possible implementation, the department collaboration matrix is constructed based on records of actual conflict events. When two departments simultaneously request resources that exceed the system capacity, this is recorded as a conflict event. The matrix elements represent the number of conflicts between the two departments during the statistical period.
[0022] For example, if the machine learning department and the data analytics department frequently compete for GPU resources, the matrix element values will be relatively high. This matrix structure clearly demonstrates the intensity of the competition between the departments.
[0023] It's important to note that task dependency analysis requires integration with the task chain information in appointment records. In a directed graph, if department A's output data needs to be transferred to department B for subsequent processing, a directed edge from A to B is established in the graph. Edge weights consider not only the closeness of task dependencies but also the conflict frequency in the collaboration matrix. The weight is calculated by multiplying the conflict frequency by the dependency strength, reflecting the degree of conflict within the collaboration.
[0024] Exemplarily, the identification of high-contention node pairs is achieved by setting a conflict frequency threshold. If the number of conflicts between two departments in a certain period exceeds 30% of the total number of appointments, they are marked as high-contention node pairs. These node pairs are often concentrated in specific time periods, such as before the project deadline or during an important experimental window. The overlap of appointment time periods is obtained by calculating the proportion of the intersection of the appointment times of the two departments to their respective total appointment times. The calculation process of the time distribution entropy value reflects the time concentration of resource usage. Divide a day into 24 time periods, and count the frequency of resource usage in each time period. The more uniform the frequency, the greater the entropy value, indicating that the resource competition is more dispersed in time. Conversely, if resource usage is concentrated in a few time periods, the entropy value is small, indicating that there is an obvious peak period of usage. In this way, the time distribution characteristics of resource competition can be accurately identified, providing data support for subsequent resource scheduling optimization.
[0025] S102. Identify model iteration delay data in computing cluster usage records, analyze the causal path between training period reservations and model iteration delays based on the distribution characteristics of resource competition, extract the GPU occupancy peak and delay duration for each period, and determine the time node corresponding to the delay.
[0026] The planned start time and actual start time of each training task are extracted from the computing cluster usage records. The difference between the two is calculated to obtain the delay duration. Model iteration tasks with delays exceeding a preset threshold are identified, and the GPU usage and memory requirements corresponding to these tasks are obtained to form a delay record set containing the delay duration, resource requirements, and occurrence period. Based on the occurrence period information in the delay record set, the peak GPU usage data of the corresponding period in the aforementioned resource competition distribution characteristics are matched. The correlation between the delay duration and the peak GPU usage is calculated using the Pearson correlation coefficient. If the correlation coefficient exceeds the preset threshold, a causal relationship is established between the training appointment conflict and the model iteration delay within the period, and a path mapping table containing causal strength values is obtained. For each causal association record in the path mapping table, the peak GPU usage of the corresponding period and the number of tasks waiting to be executed within the period are extracted. The specific time node when the delay occurred is determined based on the difference between the actual start time and the planned start time of the task, resulting in a delay feature dataset containing the peak GPU usage of each period, the delay duration, and the corresponding time node.
[0027] Specifically, the identification of model iteration delays relies on the accurate calculation of time differences.
[0028] In one possible implementation, the computing cluster records the planned and actual start timestamps of each training task. The difference between the two represents the delay duration. For example, if a deep learning model training task was originally scheduled to start at 9:00 AM but didn't actually start until 10:30 AM due to resource usage, the delay would be 90 minutes. This delay is often closely related to resource demand, and the number of GPUs requested for the task should also be recorded. For example, a task requesting 8 GPUs is more likely to experience delays due to resource shortages than one requesting 2 GPUs. Constructing a set of delay records requires integrating multi-dimensional information.
[0029] Specifically, each delay record includes five key elements: task ID, delay duration, GPU requirements, memory requirements, and the time period of occurrence. The time period is recorded at hourly granularity to facilitate subsequent matching of time periods with resource competition distribution characteristics. When a natural language processing model training task experienced a 45-minute delay between 2:00 PM and 3:00 PM, and a computer vision task also experienced a similar delay during the same period, this indicated a systemic resource bottleneck during that period.
[0030] It's important to note that the Pearson correlation coefficient quantifies the linear correlation between delay duration and peak GPU utilization. The correlation coefficient is calculated based on the delay duration series and the corresponding GPU utilization series for multiple tasks within the same time period. When the correlation coefficient exceeds 0.8, it indicates that GPU resource constraints are the primary cause of task delays. This strong correlation constitutes the core evidence for the causal path, and the correlation coefficient value for each time period is recorded in the path mapping table as a quantitative indicator of causal strength.
[0031] In one embodiment, establishing causal relationships also requires considering conflicts in training appointments. When the total number of GPUs reserved by multiple departments during the same time period exceeds the cluster capacity, queues will inevitably occur. The number of tasks waiting to be executed directly reflects the level of congestion during that time period.
[0032] For example, if there are 15 tasks waiting in the queue between 8:00 PM and 9:00 PM, while there are only two tasks waiting between 3:00 AM and 4:00 AM, the delay risk for the former is significantly higher than that for the latter. Time nodes are determined using timestamps accurate to the minute. By comparing the scheduled and actual start times of tasks, the exact moment of delay can be pinpointed. These time nodes often closely coincide with peak GPU usage, further validating the causal relationship between resource contention and delays. The resulting delay feature dataset provides a quantitative basis for subsequent resource scheduling optimization, enabling managers to identify high-risk periods and take preventative measures.
[0033] S103. Obtain multi-module parallel development task allocation data for delayed nodes, identify the correlation between task overlap and code submission interruption, analyze the relationship between task switching frequency and submission interruption through time series comparison, and evaluate the degree of interference of multi-module parallel development on progress.
[0034] The specific time of the delay is obtained from the aforementioned delay time node data. The corresponding multi-module parallel development task allocation records are queried for these times. The number of modules and task types each developer undertakes during the same period are counted. Developers with overlapping tasks and their overlapping periods are identified, and a task overlap data table containing personnel identification, number of modules, and overlapping periods is generated. Based on the personnel identification and overlapping periods in the task overlap data table, the code version control records are queried to obtain the code submission frequency and submission interval within the corresponding period. If the interval between two consecutive submissions exceeds a preset threshold, it is marked as a submission interruption. The ratio of the number of interruptions within the overlapping period to the number of interruptions within the non-overlapping period is calculated to obtain the correlation coefficient between task overlap and code submission interruption. The correlation coefficient is used to screen out highly correlated personnel records, extract their task switching time series and code submission time series, and use the sliding window method to calculate the number of task switches within each time window. The number of submission interruptions within the same window is compared, and the ratio between the two is used to evaluate the degree of interference of multi-module parallel development on progress.
[0035] Specifically, task overlap is a common phenomenon in multi-module parallel development, and its identification relies on accurate task allocation record analysis.
[0036] Specifically, when a developer needs to handle development tasks for three different modules at the same time: the front-end interface module, the back-end interface module, and the database module, typical task overlap occurs. This overlap is reflected not only in the number of tasks but also in the diversity of task types, such as simultaneously developing new features, fixing bugs, and optimizing performance.
[0037] In one possible implementation, the task overlap data table needs to be constructed by comprehensively considering both the time dimension and the task dimension. Each record contains the developer's unique identifier, the number of modules undertaken during a specific period, and the specific overlap period information.
[0038] For example, a developer may be responsible for the development of four modules simultaneously between 9:00 and 12:00 in the morning, but only one module between 2:00 and 5:00 in the afternoon. This difference directly affects his work rhythm and code submission pattern.
[0039] It should be noted that the criteria for determining code submission interruptions are based on analysis of commit records in the version control system. A submission interruption occurs when the time interval between two consecutive code submissions during a developer's normal working hours exceeds a preset threshold, such as more than three hours. This interruption is often closely related to task switching. When developers switch between modules, they need to re-understand the business logic and code structure, which interrupts the development rhythm. The calculation of the correlation coefficient reflects the impact of task overlap on code submission continuity. This impact can be quantified by comparing the number of interruptions during overlapping and non-overlapping periods.
[0040] For example, a developer experiences an average of four submission interruptions per day during the period of task overlap, but only one interruption when focusing on single module development. The ratio of 4 is the correlation coefficient, indicating that task overlap increases the risk of submission interruption by three times.
[0041] For example, the application of the sliding window method in time series analysis can capture the dynamic characteristics of task switching. The window size is set to 2 hours, sliding every 30 minutes, and the number of task switches in each window is counted. When the window switches from module A to module B and then to module C, it is recorded as 2 switches. At the same time, the number of code submission interruptions in the same window is counted, and the proportional relationship between the two intuitively reflects the interference of task switching on development continuity. The evaluation of the interference degree value comprehensively considers the correlation between the switching frequency and the number of interruptions. When the number of task switches is highly positively correlated with the number of submission interruptions, it means that the parallel development of multiple modules has seriously affected the continuity of the development progress. This quantitative evaluation provides an objective basis for project management, helping managers to reasonably arrange task allocation, reduce unnecessary module switching, and improve overall development efficiency.
[0042] S104. When the interference level exceeds the threshold, obtain the task priorities and resource requirements of each department, construct a task overlap dependency subgraph, and quantitatively analyze the task allocation imbalance indicator.
[0043] If the level of disruption to progress caused by multi-module parallel development exceeds a preset threshold, the priority values and resource requirements of each department's current tasks are retrieved from task management records, including the amount of computing resources, storage space, and estimated execution time required for each task. A directed graph structure is constructed based on the data transfer relationships between tasks, with nodes representing tasks and edges representing dependencies, forming a task overlap dependency subgraph. Based on the node and edge information in the task overlap dependency subgraph, the number of overlaps between multiple high-priority tasks on the same resource is calculated as the priority conflict degree. The resource demand gap is calculated by summing the difference between the total resource demand and the available resources in each period. The task completion rate is then extracted from historical records as the ratio of the number of tasks completed on schedule to the total number of tasks in each department, generating initial imbalance assessment data containing these three indicators. Based on the three indicator values in this initial imbalance assessment data, a weighted summation method is used to calculate a comprehensive imbalance index, where the priority conflict degree, resource demand gap, and the inverse of the task completion rate are assigned preset weight coefficients. This yields a quantitative task allocation imbalance index, completing an assessment of the degree of imbalance in task allocation within each department.
[0044] Specifically, the interference degree value serves as a trigger condition, reflecting the severity of the impact of multi-module parallel development on the overall progress.
[0045] Specifically, when this value exceeds a preset threshold, such as 0.7, it indicates that task switching and code submission interruptions during development have seriously impacted project progress, necessitating immediate and in-depth analysis of task allocation status. This threshold judgment mechanism ensures that subsequent complex analysis processes are only initiated when intervention is truly necessary.
[0046] In one possible implementation, task management records contain rich, multi-dimensional information. Priority values typically range from 1 to 5, with 5 representing the highest priority and 1 the lowest. Resource requirement details include computing resource requirements, such as 32 CPU cores and 128GB of memory; storage requirements, such as 2TB of data storage; and estimated execution times, such as a deep learning training task estimated to take 72 hours to complete. These details provide the foundational data for subsequent dependency analysis.
[0047] It's important to note that the construction of the task overlapping dependency subgraph is based on the data flow relationship between tasks. When the output data of task A needs to be used as input for task B, an edge is created in the directed graph from A to B. This dependency relationship is very common in scientific research projects, for example, when the output of a data preprocessing task directly affects the subsequent model training task. The subgraph structure clearly displays the dependency network between tasks, making the task execution order and resource conflicts clear at a glance. The calculation of priority conflict focuses on the core contradictions of resource competition.
[0048] For example, if three tasks with a priority of 5 simultaneously require GPU resources, but the system only has two available GPUs, the conflict level is 3. This metric directly reflects the intensity of resource competition among high-priority tasks. The resource demand gap is calculated through a simple difference calculation. For example, if the total demand for a certain period of time is 100 CPU cores, but only 60 are available, the gap is 40 cores.
[0049] For example, historical data on task completion rates provides an objective assessment of a department's execution capabilities. A department was assigned 20 tasks over the past month, completing 16 of them on schedule, for a completion rate of 0.8. This metric not only reflects the department's execution efficiency but also suggests potential resource allocation issues. Departments with low completion rates are often victims of insufficient resource allocation or irrational task scheduling. A weighted summation method reflects the relative importance of different indicators when calculating the comprehensive imbalance index. Priority conflict might be assigned a weight of 0.4 because it directly impacts the execution of critical tasks; resource gaps receive a weight of 0.4 to reflect the severity of resource shortages; and the inverse of task completion rates receive a weight of 0.2 to reflect historical execution. Through this weighted calculation, the resulting comprehensive imbalance index comprehensively reflects the degree of imbalance in task allocation, providing a quantitative basis for management decision-making.
[0050] S105: Extract the distributed data interface configuration information, combine it with the task allocation imbalance indicator, detect the inconsistency between the protocol and the task allocation data, and identify the impact path of the data synchronization delay.
[0051] The interface configuration information of each node is extracted from the distributed data management configuration file, including the data transmission protocol type, port number, timeout period, and buffer size parameters. The network topology and data flow relationship of each department's task execution nodes are obtained to form an interface configuration dataset containing node identification, protocol type, and configuration parameters. Based on the protocol type and task allocation imbalance indicator value in the interface configuration dataset, the data transmission protocol actually used by each node is compared with the protocol requirements preset in the task allocation plan. If a protocol type mismatch or a deviation between the parameter configuration and the task requirements is found, the node is marked as inconsistent. The proportion of inconsistent nodes to the total number of nodes is calculated to obtain the protocol consistency deviation value. The protocol consistency deviation value is used to filter out nodes with deviations exceeding the threshold. The timestamp records of these nodes during the data transmission process are tracked, and the difference between the actual transmission time and the expected transmission time is calculated as the synchronization delay duration. A delay propagation path diagram is constructed based on the data dependency relationship between nodes to identify the impact path of data synchronization delay.
[0052] Specifically, the distributed data management configuration file contains key parameter information for system operation.
[0053] Specifically, data transmission protocol types include TCP, UDP, HTTP, and other options, each corresponding to different transmission characteristics and applicable scenarios. The TCP protocol provides reliable data transmission and is suitable for mission-critical data; the UDP protocol has a fast transmission speed but does not guarantee reliability, making it suitable for monitoring data with high real-time requirements. The port number configuration determines the network location where the service listens, such as port 8080 for web services and port 3306 for database connections. The timeout parameter controls the maximum waiting time for the connection to avoid long-term resource occupation. The buffer size directly affects data transmission efficiency. If it is too small, transmission operations will be triggered frequently, while if it is too large, excessive memory resources will be occupied.
[0054] In one possible implementation, obtaining the network topology reveals the connectivity between the task execution nodes in each department. The master node, typically located at the center of the network, is responsible for task scheduling and resource allocation; compute nodes are distributed across departments, performing specific data processing tasks; and storage nodes provide data persistence services. Data flow relationships record the data transfer path during task execution, such as the flow of raw data from storage nodes to compute nodes and the return of processed results to storage nodes. This topology and flow information provide the basis for subsequent protocol matching checks.
[0055] It's important to note that protocol inconsistency can cause serious data transmission problems in distributed environments. This occurs when a task allocation plan requires the use of TCP to ensure data integrity, but the actual node configuration uses UDP. Parameter configuration deviation is equally important. For example, if a task requires large amounts of data transmission and a buffer size of 64KB is specified, but the node configuration is only 8KB, this can lead to inefficient transmission. The protocol consistency deviation value, calculated as the ratio of the number of inconsistent nodes to the total number of nodes, reflects the degree of misconfiguration in the entire system.
[0056] For example, timestamp recording plays a key role in tracking data synchronization delays. Each data packet records a precise timestamp when it is sent and received, and by comparing these timestamps, the actual transmission time can be calculated. When a data packet is expected to be transmitted within 100 milliseconds, but it actually takes 500 milliseconds, the synchronization delay is 400 milliseconds. This delay will propagate downstream along the data dependency relationship, forming a chain reaction. The construction of the delay propagation path graph is based on the data dependency relationship between nodes. If the delay of node A causes node B to wait for input data, and the delay of node B affects node C, a delay propagation path of A→B→C is formed. By identifying these paths, it is possible to locate the key bottlenecks in the system and provide precise guidance for optimizing configuration. The identification of such impact paths is of great significance for improving the operating efficiency of the entire distributed system.
[0057] S106. Analyze the dependency path between the resource competition risk in the distribution characteristics of data synchronization delay impact paths and computing resource competition, construct a risk map that includes departmental collaboration conflicts and chain risk diffusion, and obtain the risk factor distribution.
[0058] From the impact path of data synchronization delays, we extract the sequence of delayed nodes and delay duration data. We then match the resource utilization and conflict frequency information of the corresponding nodes in the resource contention distribution characteristics to identify overlapping nodes experiencing both data delays and resource contention. We then calculate the composite risk value for each overlapping node, multiplying the delay duration by the resource conflict frequency. This creates a multi-dependency path data table containing node identifiers, composite risk values, and dependency relationships. Based on the node dependencies and composite risk values in the multi-dependency path data table and combined with inter-departmental collaboration conflict frequency data, we construct a directed weighted graph structure. Nodes represent departments or task execution units, and edges represent risk propagation paths. Edge weights are determined by multiplying the composite risk value of the upstream node by the degree of inter-departmental collaboration conflict. Edges with weights exceeding a preset threshold are marked as high-risk propagation paths. Through high-risk propagation paths, we track the spread of risk across departments. A graph traversal method is used to calculate the cumulative risk impact value for each node. The distribution of risk factors across departments and task nodes is determined based on the magnitude and distribution density of the cumulative risk impact values, resulting in risk factor distribution data containing both numerical and spatial distribution characteristics.
[0059] Specifically, the calculation of the composite risk value reflects the quantitative idea of the superposition effect of multiple factors.
[0060] In one possible implementation, when a computing node experiences a data synchronization delay of 300 milliseconds and a resource conflict frequency of 5 per hour, its composite risk value is 1500. This value not only reflects the impact of a single risk factor but, more importantly, reveals the amplified effect of the interaction of multiple risk factors. Overlapping nodes are identified through spatiotemporal matching. That is, nodes that appear on the delay path and in a resource contention hotspot within the same time window are marked as overlapping nodes.
[0061] It's important to note that the construction of the directed weighted graph structure fully considers the directionality and intensity of risk propagation. Nodes in the graph can represent specific departments, such as the Data Analysis Department or the Algorithm R&D Department, or they can represent task execution units, such as a model training server group. Edges are oriented according to task dependencies and data flow, from upstream nodes to downstream nodes. The calculation of edge weights involves two key factors: the composite risk value of the upstream node reflects the severity of the source risk, while the degree of interdepartmental collaboration conflict reflects the resistance or support to risk propagation. If the composite risk value of the Algorithm R&D Department is 2000 and the degree of collaboration conflict with the Data Analysis Department is 0.8, the edge weight connecting the two is 1600.
[0062] Specifically, high-risk transmission paths are identified using a threshold judgment mechanism. The preset threshold is typically determined based on statistical analysis of historical data, such as the 75th percentile of all edge weights. Paths exceeding this threshold are marked as high-risk transmission paths, which are often the most vulnerable links in the system.
[0063] For example, if there are high compound risk values and collaboration conflicts in each link of the path from the data collection node through the preprocessing node to the model training node, the entire path will become a channel for concentrated risk outbreaks.
[0064] For example, graph traversal methods employ a depth-first or breadth-first strategy when calculating cumulative risk impact values. Starting from the high-risk source node, the risk is propagated layer by layer along directed edges, accumulating risk values for each edge according to the edge weight. A node may simultaneously receive risk inputs from multiple upstream nodes, and these inputs are summed to form its cumulative risk impact value. The data analysis department may be subject to both delay risk from the data acquisition department and resource competition risk from the computing resource scheduling department. The combined effect of these factors causes its cumulative risk value to peak. The spatial characteristics of the risk factor distribution are reflected by both the node's position in the network topology and the magnitude of its cumulative risk value. Core nodes tend to have higher risk factor values because they connect multiple upstream and downstream nodes and are prone to becoming risk convergence points. While edge nodes may present lower direct risks, they may be severely impacted when risks erupt due to a lack of redundant paths. Identifying these distribution characteristics provides a precise quantitative basis for subsequent risk management and resource optimization.
[0065] By identifying the servers, databases, and interface nodes where data synchronization is delayed, we can determine the paths by which the delays affect each business process. We can then analyze resource competition scenarios based on the path optimization results, analyze the priority conflicts of each department based on the scenarios, and determine the multiple dependencies of resource competition risks. Based on these multiple dependencies, we can analyze the potential conflicts caused by resource competition in cross-departmental collaboration. Combined with historical business peak data, we can generate the triggering conditions and impact range of risk points in the dependency paths.
[0066] Identify servers, databases, and interface nodes experiencing data synchronization delays, extract the delay duration and frequency of each node, track the affected downstream business processes, construct impact propagation paths based on the execution order and data dependencies of the business processes, calculate the cumulative delay time and number of affected tasks along each path, and generate a business impact dataset containing node type, impact path, and delay severity. Based on the affected task information in the business impact dataset, match the computing resource requirements and usage periods of each task, identify task combinations with overlapping resource requirements, and determine the degree of priority conflict by comparing the priority values of tasks across departments. If multiple high-priority tasks compete for the same resource, construct a resource contention dependency graph and determine a multi-dependency matrix that captures the resource contention relationships and priority differences between tasks. This multi-dependency matrix identifies resource contention nodes between cross-departmental tasks. Combined with historical resource usage records and task execution logs during peak business periods, extract peak resource saturation thresholds and task queue times. Compare current resource utilization with historical thresholds to determine risk trigger conditions. Calculate the length of the risk propagation path and the number of departments involved to obtain impact data.
[0067] Specifically, the identification of data synchronization delay nodes covers key components in the system architecture.
[0068] Specifically, server nodes include application servers, computing servers, and file servers, each with distinct latency characteristics. Application server latency primarily manifests as increased response times, such as request processing extending from 50 milliseconds to 200 milliseconds. Database node latency manifests as increased query execution and transaction submission times. Interface node latency is reflected in increased data transmission and protocol conversion time. These delays are calculated by diffing timestamps recorded by monitoring tools.
[0069] In one possible implementation, constructing impact propagation paths requires a deep understanding of the data dependencies of business processes. A 300-millisecond delay in the order processing server can have a knock-on impact on downstream processes like inventory updates, payment confirmations, and logistics scheduling. The cumulative delay is calculated by adding up the delays at each node. For example, if order processing is delayed by 300 milliseconds, inventory updates add 150 milliseconds due to the delay, and payment confirmation adds another 200 milliseconds, for a cumulative delay of 650 milliseconds along the entire path. The number of affected tasks is calculated by counting all waiting or delayed tasks along the path.
[0070] It's important to note that the resource contention dependency graph is constructed based on the resource contention relationships between tasks. When the data analysis department's batch processing tasks require 80% of CPU resources, while the algorithm training department's deep learning tasks also require 70%, these two tasks directly compete for resources. The degree of priority conflict is determined by comparing the numerical differences in task priorities, which are scaled from 1 to 10, with higher values indicating higher priorities. If two critical tasks, both with a priority of 9, compete for GPU resources simultaneously, the conflict reaches its highest level.
[0071] For example, a multi-dependency matrix illustrates the complex interactions between tasks. The rows and columns of the matrix represent different tasks, and the values of the matrix elements indicate the strength of inter-task dependencies and the degree of resource competition. When the output of task A is a required input for task B, and both compete for the same memory resources, the corresponding matrix element values comprehensively reflect the dual relationships of data dependency and resource competition. This matrix structure clearly illustrates potential conflicts in cross-departmental collaboration. Analysis of historical peak business data reveals critical conditions for risk triggering. Analysis of peak month-end settlement data from the past year revealed that the probability of system delays increases sharply when CPU utilization exceeds 85% and memory utilization exceeds 90%. Risk triggering conditions are determined using statistical analysis, comparing current resource utilization with historical thresholds. Exceeding the thresholds indicates a high-risk state. The scope of impact is determined by calculating the number of hops in the risk propagation path and the number of departments involved. The longer the path and the more departments involved, the wider the scope of impact, providing a quantitative basis for risk management.
[0072] S107. Obtain historical collaboration data from each department, combine it with the distribution of risk factors, analyze and optimize the risk diffusion path and node weights, determine the intervention order for high-risk risk factors, optimize the path based on the degree of interference using the knowledge graph, analyze and calculate the dependencies of cluster conflicts, build a risk warning mechanism, formulate intervention measures, and output an optimized resource allocation plan.
[0073] Using data mining techniques, the team extracted interdepartmental collaboration records, task flow data, and personnel interaction information from internal OA (Office Office) platforms, email servers, and project management tools. SQL queries were used to analyze the frequency and temporal distribution of these collaboration records, constructing a matrix of interdepartmental collaboration intensity. Based on this matrix, a departmental collaboration network topology was generated. Based on this network topology, a centrality algorithm from graph theory was used to calculate the importance weights of each department's nodes. If a department's collaboration intensity exceeded a preset threshold, the department was labeled a high-weight risk propagation node. Based on the distribution of these high-weight risk propagation nodes, the interdepartmental propagation paths of risk factors were determined. A departmental relationship knowledge graph was constructed using the Neo4j graph database. The risk factor propagation paths were used as edge connections within the graph. The Dijkstra algorithm was used to calculate the shortest diffusion path of risk from the source department to each target department, identifying dependency nodes and potential conflict locations along these shortest diffusion paths. Based on the distribution characteristics of these dependency nodes and potential conflict locations, a threshold-based risk warning judgment rule was established. If the risk propagation intensity along the shortest diffusion path reached a warning threshold, a corresponding warning signal and resource allocation adjustment plan were generated.
[0074] Specifically, the application of data mining technology in enterprise collaborative analysis is based on the ability to integrate and process multi-source heterogeneous data.
[0075] For example, the OA system of a large manufacturing company recorded 245 interactions between the Finance and Procurement departments during the monthly budget approval process. The email server showed 15.3 emails per week between the two departments regarding supplier evaluations. Project management tools tracked an average task turnover time of 4.2 days for cross-departmental collaborative projects. These dispersed data sources were uniformly extracted using SQL queries to form a standardized collaboration record dataset containing key fields such as timestamp, participating departments, interaction type, and duration. A collaboration intensity matrix constructed based on this collaboration record dataset quantifies the relationship between interdepartmental collaboration.
[0076] In one embodiment, the collaboration intensity value between the Finance Department and the Purchasing Department is 0.85, indicating a high frequency of business interactions between the two departments, while the collaboration intensity between the Finance Department and the R&D Department is only 0.23, indicating a relatively low frequency of collaboration. After graph theory transformation, the collaboration intensity matrix generates a departmental collaboration network topology structure. This structure uses departments as nodes and collaboration relationships as edges, forming a networked expression of collaboration within the enterprise. The core value of this transformation lies in converting discrete departmental relationships into a computable network model, providing a mathematical basis for subsequent importance analysis. The centrality algorithm plays a key role in department importance assessment. Its principle is to quantify the influence of nodes by analyzing their degree of connectivity and positional characteristics in the network.
[0077] Specifically, the degree centrality algorithm calculates the number of other departments that each department is directly connected to, while the closeness centrality algorithm evaluates the average distance of a department to all other departments in the network.
[0078] For example, although the Human Resources Department collaborates directly with eight other departments, its intermediary role in key business processes earns it a high centrality score of 0.78. When a department's centrality score exceeds the preset threshold of 0.7, it is identified as a high-weight risk propagation node, meaning that any abnormalities in that department could quickly affect the entire collaborative network. Determining the risk factor propagation path relies on in-depth analysis of the connectivity patterns between high-weight nodes.
[0079] It should be noted that risk factors are not single risk events, but a complex concept that includes multi-dimensional risk elements such as staff turnover, system failure, and process interruption.
[0080] For example, when a key employee leaves the Finance Department, this risk factor propagates along the Finance Department, Procurement Department, and Supply Chain Management Department, impacting interrelated business processes such as budget approval, procurement decisions, and supplier management. The identification of these propagation paths is based on dependency analysis within historical collaboration data. By analyzing the frequency and closeness of interdepartmental business flows, the likelihood and intensity of risk transmission are determined. The Neo4j graph database, as the technical foundation for knowledge graphs, provides efficient graph data storage and query capabilities.
[0081] In one possible implementation, each department node contains attribute information such as department number, department name, staff size, and business type, while edges record relationship attributes such as collaboration type, collaboration frequency, and degree of dependency. Based on this, the Dijkstra algorithm calculates the shortest path, not only considering path length but also incorporating edge weights to assess the cost and speed of risk transmission.
[0082] For example, there are two paths for risk diffusion from R&D to Marketing: R&D-Product-Marketing and R&D-Project Management-Marketing. The algorithm compares the total weight of each path and identifies the former as the shortest diffusion path. The identification of dependency nodes and potential conflict locations is based on an analysis of path bottlenecks.
[0083] Specifically, when a department simultaneously assumes the transit function of multiple risk transmission paths, it becomes a dependency node, and its normal operation directly affects the speed and scope of risk diffusion. Potential conflicts arise at departmental handover points where competition for resources is fierce or the boundaries of responsibilities are blurred, such as the overlapping areas of responsibilities between the sales department and the customer service department in the customer complaint handling process. The establishment of risk warning judgment rules is based on a threshold comparison mechanism, which triggers corresponding response measures by setting warning thresholds at different levels. When the risk transmission intensity on the shortest diffusion path reaches 0.6, a yellow warning is triggered, and when it reaches 0.8, a red warning is triggered. Accordingly, resource allocation adjustment plans of different levels are generated, realizing the automation and precision of risk management.
[0084] With the above embodiments of the present invention as inspiration, and through the above description, relevant personnel can make various changes and modifications without departing from the technical scope of this invention. The technical scope of this invention is not limited to the contents of the specification, but must be determined according to the scope of the claims.
Claims
1. An intelligent identification and early warning method for risk factors of scientific and technological consulting projects, characterized by: The method comprises: Mapping nodes of computing cluster usage records and training session reservation data in scientific research consulting projects, constructing dependency path diagrams that include departmental collaboration conflicts, and analyzing the distribution characteristics of resource competition; Identifying model iteration delay data in the computing cluster usage record, analyzing the causal path between the training period reservation data and the model iteration delay data in combination with the distribution characteristics of the resource competition, and determining the time node corresponding to the delay; Obtaining multi-module parallel development task allocation data for the time node corresponding to the delay, analyzing the correlation between task overlap and code submission interruption in the multi-module parallel development task allocation data, and evaluating the degree of interference of task switching frequency on progress; Obtain the task priorities and resource requirements of each department, construct a task overlap dependency subgraph, and quantitatively analyze task allocation imbalance indicators; Extracting distributed data interface configuration information, combining it with the task allocation imbalance indicator, detecting inconsistencies between the protocol and the departmental task allocation data, and identifying the impact path of data synchronization delays; Analyze the dependency path between the impact path of the data synchronization delay and the resource competition risk in the distribution characteristics of the resource competition, construct a risk map that includes departmental collaboration conflicts and chain risk diffusion, and obtain a risk factor distribution; Obtain historical collaboration data from various departments, combine it with the distribution of risk factors, analyze risk diffusion paths and node weights, determine the order of intervention for high-risk risk factors, build a risk early warning mechanism, formulate intervention measures, and output an optimized resource allocation plan.
2. The intelligent identification and early warning method for risk factors of scientific and technological consulting projects according to claim 1 is characterized in that: The construction of a dependency path diagram containing departmental collaboration conflicts and analysis of the distribution characteristics of resource competition include: The resource occupancy rate data of each department in different time periods in the computing cluster usage records are extracted, the computing resource and storage resource allocation ratios in the resource occupancy rate data are counted, and a department collaboration matrix is constructed to record the conflict frequency caused by resource competition between departments; according to the conflict frequency in the department collaboration matrix, the task dependency relationship in the training period reservation data is analyzed, and a directed graph structure is constructed to represent the task dependency path between departments, where the nodes represent the department entities and the edges represent the task dependency relationship and the resource transfer direction; for the high-conflict frequency node pairs in the directed graph structure, the time period overlap and priority difference in the training period reservation data are extracted, the entropy value of the resource usage time distribution is calculated, and the resource occupancy peak time period and the conflict frequency distribution in the distribution characteristics of the resource competition are determined.
3. The intelligent identification and early warning method for risk factors of scientific and technological consulting projects according to claim 1 is characterized in that: The identifying of the model iteration delay data in the computing cluster usage record, combining the distribution characteristics of the resource competition, analyzing the causal path between the training period reservation data and the model iteration delay data, and determining the time node corresponding to the delay includes: Extract the difference between the planned and actual start time of the training task in the computing cluster usage record to generate a delay record set; match the delay record set with the resource occupancy data of the corresponding time period in the distribution characteristics of the resource competition, calculate the correlation between the delay duration and the resource occupancy data, and establish a causal association path; and determine the specific time node when the delay occurred based on the causal association path.
4. The intelligent identification and early warning method for risk factors of scientific and technological consulting projects according to claim 1 is characterized in that: The analyzing the correlation between task overlap and code submission interruption in the multi-module parallel development task allocation data includes: Extract the task overlapping period in the multi-module parallel development task allocation data, and count the number of modules undertaken by developers; query the code submission records within the task overlapping period, calculate the submission interval time, and determine the number of submission interruptions; compare the number of submission interruptions in the task overlapping period and the non-overlapping period to generate a correlation coefficient.
5. The intelligent identification and early warning method for risk factors of scientific and technological consulting projects according to claim 1 is characterized in that: The construction of the task overlapping dependency subgraph and the quantitative analysis of task allocation imbalance indicators include: According to the task priorities and resource requirements in the departmental task allocation data, a directed graph structure is constructed to represent the task dependencies; the number of overlapping high-priority tasks in the directed graph structure is calculated, the difference between resource requirements and available resources is counted, the task completion rate is extracted, and a task allocation imbalance index value is generated.
6. The intelligent identification and early warning method for risk factors of scientific and technological consulting projects according to claim 1 is characterized in that: The detection protocol is inconsistent with the departmental task allocation data, and the impact path of data synchronization delay is identified, including: Extract the protocol type and parameters in the distributed data interface configuration information to generate an interface configuration data set; compare the interface configuration data set with the protocol requirements of the department task allocation data, and mark inconsistent nodes; and construct a data synchronization delay propagation path diagram based on the transmission time difference of the inconsistent nodes.
7. The intelligent identification and early warning method for risk factors of scientific and technological consulting projects according to claim 1 is characterized in that: The analyzing the dependency path between the impact path of the data synchronization delay and the resource contention risk in the distribution characteristics of the resource contention includes: Extract the delayed node sequence in the impact path of the data synchronization delay, match the resource occupancy data in the distribution characteristics of the resource competition, and identify overlapping nodes; generate a composite risk value based on the delay duration and conflict frequency of the overlapping nodes, and construct a directed weighted graph containing the risk propagation path.
8. The method according to claim 1, characterized in that The method also includes: identifying servers, databases, and interface nodes in the impact path of the data synchronization delay, determining the impact path of the servers, databases, and interface nodes on the business process, analyzing priority conflicts among departments, and determining multiple dependencies of resource competition risks.
9. The method according to claim 8, characterized in that The identifying of the servers, databases, and interface nodes in the impact path of the data synchronization delay, determining the impact path of the servers, databases, and interface nodes on the business process, analyzing the priority conflicts of various departments, and determining the multiple dependencies of resource competition risks include: Extract the delay duration and frequency of the servers, databases, and interface nodes in the impact path of the data synchronization delay, and construct a business process impact propagation path; match the task resource requirements in the business process impact propagation path, and identify task combinations with overlapping resource requirements; generate a resource competition dependency graph based on the priority differences of the task combinations, and determine the multiple dependency matrix of the resource competition risk; combine the multiple dependency matrix of the resource competition risk to analyze the resource competition nodes between the tasks of each department and generate priority conflict distribution data.
10. The intelligent identification and early warning method for risk factors of scientific and technological consulting projects according to claim 1 is characterized in that: The acquisition of historical collaboration data from various departments, combined with the risk factor distribution, analysis of risk diffusion paths and node weights, and determination of the intervention sequence for high-risk risk factors include: Extract the collaboration records and task flow data from the historical collaboration data of each department to construct a department collaboration network topology structure; calculate the importance weight values of department nodes based on the department collaboration network topology structure, and mark high-weight risk propagation nodes; based on the high-weight risk propagation nodes and the risk factor distribution, generate a risk diffusion path diagram to determine the risk factor intervention order.
Citation Information
Cited By
Task alarm processing method and system based on intelligent grading
CN121455650A
Enterprise operation state real-time analysis method and system based on multi-source data fusion
CN121480991A
Methods, apparatus, electronic devices and storage media of assisting testing
CN122489367A
Methods, apparatus, electronic devices and storage media of assisting testing
CN122489367B