A distributed computing power scheduling intelligent optimization method and system
By building a multi-level system dependency graph and causal relationship model, combining multi-dimensional anomaly detection and fault prediction technology, real-time monitoring and preventive resource scheduling, the problems of inaccurate fault positioning and lag in distributed systems are solved, and efficient fault handling and resource optimization are achieved.
Patent Information
- Application Number
- CN202510637762.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-05-19
AI Technical Summary
In the prior art, there are problems such as insufficient accuracy in fault positioning, passive lag in fault processing, separation of resource scheduling and fault processing, and serious dependence on manual intervention, resulting in inadequate fault diagnosis and repair efficiency of distributed systems.
Build a multi-level system dependency graph and causal relationship model, combine multi-dimensional anomaly detection and fault prediction technology, monitor the system status in real time, identify the fault propagation path, and implement preventive resource scheduling and isolation strategies to achieve intelligent repair and resource reorganization.
It improves fault prediction capabilities, shortens fault location time, improves system availability and resource utilization efficiency, reduces operation and maintenance costs, and enhances system resilience and reliability.
Smart Images

Figure CN120179507B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of distributed computing technology, and more specifically, to a distributed computing power scheduling intelligent optimization method and system. Background Art
[0002] As the scale and complexity of distributed systems continue to grow, fault diagnosis and resource scheduling in the system are becoming increasingly difficult. The main problems in the existing technology are as follows:
[0003] Insufficient fault location accuracy: Traditional fault diagnosis methods are typically based on single-dimensional threshold monitoring or simple statistical correlation analysis. These methods are difficult to accurately locate the root cause of faults in complex distributed systems, especially when system components have complex dependencies.
[0004] Passive and delayed fault handling: Current fault handling strategies are mostly reactive, with diagnosis and repair only beginning after a fault has occurred and impacted user experience. This lacks effective prediction and prevention mechanisms.
[0005] Resource scheduling and fault handling are separated: In existing technologies, resource scheduling and fault handling are often treated as two independent issues. Resource scheduling systems rarely consider the failure risk and health status of the system, resulting in inefficient resource utilization when faced with faults.
[0006] Heavy reliance on manual intervention: The diagnosis and repair of complex faults still rely heavily on the experience and intuition of operations personnel, making intelligent and automated processing difficult to achieve. This results in extended fault repair time and reduced system availability.
[0007] Lack of a systematic causal relationship model: Existing technologies mostly conduct analysis based on correlation rather than causality, which is easily affected by false correlations and makes it difficult to accurately grasp the propagation path and impact range of the fault.
[0008] Therefore, there is a need for a distributed computing power scheduling optimization method that can organically combine fault management with resource scheduling and achieve active prevention and intelligent processing based on a causal model. Summary of the Invention
[0009] The present invention provides a distributed computing power scheduling intelligent optimization method and system to solve the technical problems in related technologies such as insufficient fault location accuracy, passive and delayed fault processing, separation of resource scheduling and fault processing, and heavy reliance on manual intervention.
[0010] The present invention provides a distributed computing power scheduling intelligent optimization method, comprising the following steps:
[0011] Construct a multi-layer system dependency graph and causal relationship model. The multi-layer system dependency graph includes the hardware layer, virtualization layer, application service layer, and business function layer. The causal relationship model describes the dependencies and impact paths within and between each layer.
[0012] Based on a multi-level system dependency graph and causal relationship model, it uses multi-dimensional anomaly detection and fault prediction technology to monitor system status in real time and predict potential faults;
[0013] When an anomaly is detected or a potential fault is predicted, a critical path analysis is performed using a multi-level system dependency graph and causal relationship model to identify the fault propagation path and key impact nodes. Fault isolation strategies are then implemented to block fault propagation and implement preventive resource scheduling.
[0014] Implement intelligent repair and resource reorganization based on fault isolation results and business continuity requirements.
[0015] In a preferred embodiment, the step of constructing a multi-level system dependency graph and causal relationship model includes:
[0016] Collect component information and status data at all levels in the system;
[0017] Analyze historical data through data mining technology to identify the dependencies between components;
[0018] Combine expert knowledge and machine learning algorithms to build a causal relationship model between components;
[0019] Regularly update multi-level system dependency diagrams and causal models to adapt to system changes.
[0020] In a preferred embodiment, the multi-dimensional anomaly detection and fault prediction steps include:
[0021] Establish a multi-dimensional normal behavior baseline for key system indicators based on historical data and expert knowledge;
[0022] Monitor system status indicators in real time and use multiple anomaly detection algorithms to identify abnormal patterns that deviate from normal behavior;
[0023] Predict the probability, impact range, and severity of potential failures based on time series analysis and causal reasoning;
[0024] Generate anomaly detection and fault prediction reports to provide a basis for further decision-making.
[0025] In a preferred embodiment, the critical path analysis and fault isolation in the steps of blocking fault propagation and implementing preventive resource scheduling include:
[0026] Analyze the propagation path of anomalies or potential faults based on multi-level dependency graphs;
[0027] Calculate the risk assessment results of each propagation path and identify key nodes and links that require priority treatment;
[0028] Develop and implement targeted isolation measures to block the propagation of faults;
[0029] Implement a hierarchical isolation strategy for businesses requiring high availability to ensure the continuity of core businesses.
[0030] In a preferred embodiment, the preventive resource scheduling step includes:
[0031] Adjust resource allocation strategies based on failure risk assessment results;
[0032] Perform health scoring on tasks on high-risk nodes and calculate migration priority scores based on the scoring results and priorities;
[0033] Create shadow instances for critical tasks to achieve seamless switching;
[0034] Implement multi-level resource reservation and priority management to ensure resource requirements for high-priority tasks.
[0035] In a preferred embodiment, the intelligent repair and resource reorganization step includes:
[0036] Automatically select the appropriate repair strategy based on the fault type and system status;
[0037] Execute automatic repair processes for faults that can be repaired automatically;
[0038] Generate detailed fault reports and repair suggestions for complex faults that require manual intervention;
[0039] Reallocate and optimize resource configuration based on the repaired system status and business needs.
[0040] In a preferred embodiment, a distributed computing power scheduling intelligent optimization method further includes continuous learning and optimization steps:
[0041] Collect and analyze data and results during troubleshooting;
[0042] Updated causality models and anomaly detection baselines;
[0043] Optimize fault prediction algorithms and resource scheduling strategies;
[0044] Generate optimization reports for long-term system improvement.
[0045] In a preferred embodiment, a distributed computing power scheduling intelligent optimization system is used to execute a distributed computing power scheduling intelligent optimization method, including:
[0046] The system perception and modeling module is used to collect component information and status data at all levels of the system, build multi-level system dependency graphs and causal relationship models, and monitor system status in real time to identify anomalies and predict potential failures;
[0047] Fault analysis and isolation module, used to analyze fault propagation paths and identify key nodes, develop and implement targeted isolation strategies, and block fault propagation;
[0048] Resource scheduling and repair module, which is used to perform preventive resource scheduling based on risk assessment results and implement intelligent repair strategies according to fault type and system status;
[0049] The continuous learning and optimization module is used to collect and analyze fault handling data, update the causal relationship model and anomaly detection baseline, and continuously optimize the system model and scheduling strategy.
[0050] In a preferred embodiment, a distributed computing power scheduling intelligent optimization system further includes a decision support module for:
[0051] Comprehensive analysis of anomaly detection, fault prediction, and resource status information;
[0052] Evaluate the costs and benefits of different treatment strategies;
[0053] Make processing decisions automatically or semi-automatically based on preset strategies and system status;
[0054] Record the decision-making process and results for subsequent learning and optimization.
[0055] In a preferred embodiment, the distributed computing power scheduling intelligent optimization system is deployed using a distributed architecture, including:
[0056] Central control node, responsible for global decision-making and coordination;
[0057] Multiple edge nodes, distributed throughout the system, are responsible for local data collection and preprocessing;
[0058] Data sharing mechanism ensures real-time synchronization of information between nodes;
[0059] Fault-tolerant mechanism ensures that the system can still operate normally when some nodes fail.
[0060] The beneficial effects of the present invention are:
[0061] Improved fault prediction capabilities: By building a multi-level system dependency graph and causal relationship model, combined with multi-dimensional anomaly detection technology, potential faults in the system can be identified in advance, transforming fault handling from passive response to active prevention, reducing the occurrence rate of faults.
[0062] Improved fault location accuracy: The root cause analysis method based on the causal relationship model can accurately trace the source and propagation path of the fault, shortening the fault location time from several hours with traditional methods to minutes, thereby improving the location accuracy.
[0063] Improved system availability: Preventive resource scheduling and fault isolation technologies effectively block fault propagation and improve the availability of key services.
[0064] Optimizing resource utilization efficiency: Integrating fault risk perception into resource scheduling decisions enables dynamic and optimized resource allocation, improving resource utilization while ensuring system stability.
[0065] Improved automation of operations and maintenance: Intelligent repair and resource reorganization functions significantly reduce manual intervention and lower operation and maintenance costs.
[0066] Enhanced system resilience: Multi-level resource reservation and priority management mechanisms ensure the continuity of critical services in the event of failures, reducing system recovery time.
[0067] Continuous optimization capability: By continuously learning from fault handling experience and system operation modes, the system's fault handling capabilities and resource scheduling strategies will continue to improve, forming an adaptive optimization closed loop. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] Figure 1 This is a flow chart of a distributed computing power scheduling intelligent optimization method of the present invention;
[0069] Figure 2 It is a detailed flow chart of the present invention for constructing a multi-level system dependency graph and a causal relationship model;
[0070] Figure 3 It is a detailed flow chart of the present invention for real-time monitoring of system status and prediction of potential failures;
[0071] Figure 4 It is a detailed flow chart of the present invention for blocking fault propagation and implementing preventive resource scheduling;
[0072] Figure 5 It is a detailed flow chart of the implementation of intelligent repair and resource reorganization of the present invention. DETAILED DESCRIPTION
[0073] The subject matter described herein will now be discussed with reference to example embodiments. It should be understood that these embodiments are discussed solely to enable those skilled in the art to better understand and implement the subject matter described herein, and that the functions and arrangements of the elements discussed may be varied without departing from the scope of this specification. Various examples may omit, substitute, or add various processes or components as needed. Furthermore, features described in some examples may be combined in other examples.
[0074] At least one embodiment of the present invention discloses a distributed computing power scheduling intelligent optimization method, such as Figures 1 to 5 As shown, the following steps are included:
[0075] Step 1: Build a multi-layer system dependency graph and causal relationship model. The multi-layer system dependency graph includes the hardware layer, virtualization layer, application service layer, and business function layer. The causal relationship model describes the dependencies and impact paths within and between each layer.
[0076] It includes the following sub-steps:
[0077] Step 1.1, system component dependency data collection;
[0078] Utilize automatic discovery tools and monitoring system APIs to collect static configuration information and runtime dependency data of components in the distributed system. The collected data includes:
[0079] Hardware component dependencies: topological connections between physical servers, network devices, and storage devices;
[0080] Infrastructure dependencies: the mapping between virtual machines, containers, and physical hosts;
[0081] Platform component dependencies: the calling relationships between databases, message queues, caches, and other middleware;
[0082] Application service dependencies: dependencies between application layer components such as microservices and API call chains.
[0083] In some embodiments, by analyzing network traffic capture data, dynamic dependencies not recorded in the configuration management database, such as temporarily established data flow channels or implicit service calls, can be automatically discovered. For example, in a microservices architecture, service discovery may not rely entirely on a central registry, but instead identify actual service call patterns through network traffic analysis, capturing a comprehensive dependency graph that includes occasional calls.
[0084] Step 1.2, multi-level dependency graph construction;
[0085] Build a multi-level system dependency graph based on the collected dependency data :
[0086] ;
[0087] in, Represents all components in the system; Represents dependencies between components; Indicates the dependency strength between components. A larger value indicates a stronger dependency.
[0088] The dependency graph is organized through a layered architecture, from the underlying hardware to the upper-level applications, capturing the vertical dependencies between components at different levels and the horizontal dependencies between components at the same level, forming a complete dependency network.
[0089] Step 1.3, causal relationship model construction;
[0090] Based on the Bayesian network and Structural Causal Model (SCM) theory, the dependency graph is upgraded to a causal model. The causal model is constructed through the following steps:
[0091] Apply causal discovery algorithms to historical monitoring data to preliminarily identify candidate causal relationships between components;
[0092] Verify and adjust causal relationships by combining expert knowledge rule base to eliminate false correlations;
[0093] Calculate causal strength to quantify the magnitude of causal influence between components;
[0094] Establish a formal causal relationship representation:
[0095] ;
[0096] in, is the outcome variable, representing the observed or predicted state or performance indicator of the system component, yes The set of all direct causal parent nodes of All upstream components or factors, It is a causal mechanism function that describes the mathematical mapping relationship of how the parent node affects the outcome variable through a specific mechanism. It is an external interference factor, which means that it cannot be observed in the model but will affect random variables, such as environmental noise, unknown interference, etc.
[0097] The causal relationship model provided in this application is specifically implemented as follows:
[0098] Hybrid causal discovery algorithm: This application adopts a hybrid causal discovery method that combines a constraint-based algorithm (such as the PC algorithm) and a scoring-based algorithm (such as the GES algorithm).
[0099] Among them, the PC algorithm first identifies possible causal relationships through conditional independence tests and constructs an undirected skeleton structure; then applies directional rules to determine the direction of the edges and form a partially directed graph.
[0100] The GES algorithm evaluates the likelihood of different graph structures through a Bayesian scoring function and searches for the optimal model in the search space.
[0101] Step 1.4, causal discovery of multidimensional time series data;
[0102] Apply specific causal discovery algorithms to the time series data generated by the system to identify causal relationships in the time dimension:
[0103] Use Granger Causality Test to analyze the causal relationship between time series data;
[0104] Apply Conditional Transfer Entropy to calculate the direction and intensity of information flow;
[0105] Construct a temporal causal graph (TemporalCausalGraph) to represent the changes in causal relationships under different time windows;
[0106] Identify common causal chain patterns and summarize them into a fault propagation law library.
[0107] The output of step 1 is a complete multi-level system dependency graph and causal relationship model, which can accurately represent the dependency relationships and causal strengths between components in the distributed system, providing a basis for subsequent fault root cause analysis and propagation prediction.
[0108] Step 2: Based on a multi-level system dependency graph and causal relationship model, multi-dimensional anomaly detection and fault prediction technology is used to monitor the system status in real time and predict potential faults;
[0109] It includes the following sub-steps:
[0110] Step 2.1: Distributed monitoring probe deployment and data collection;
[0111] Deploy lightweight monitoring probes at key nodes in the distributed system to continuously collect multi-dimensional runtime data:
[0112] Hardware telemetry data: CPU temperature, fan speed, hard disk SMART information, etc.;
[0113] System performance indicators: CPU utilization, memory usage, I / O latency, network throughput, etc.
[0114] System logs: operating system logs, application logs, error and warning information;
[0115] Performance counters: cache hit rate, lock contention times, garbage collection frequency, etc.
[0116] Network topology and traffic data: inter-node communication delay, packet loss rate, and connection status changes.
[0117] In addition, the probe adopts an adaptive sampling strategy to dynamically adjust the sampling frequency according to the system load and anomaly probability, minimizing the monitoring overhead while ensuring data quality.
[0118] In a resource-constrained edge computing environment, probes can adopt a hierarchical activation mechanism, usually remaining in a low-energy basic monitoring mode, and only activating the full data collection function when abnormal signs are detected or a trigger signal is received from the central coordinator.
[0119] Step 2.2, multi-model anomaly detection system;
[0120] An ensemble learning approach is used to combine multiple anomaly detection algorithms to capture different types of anomaly patterns:
[0121] Apply statistical models to detect anomalies in numerical indicators;
[0122] Density-based models and distance-based models are used to identify spatial outliers;
[0123] Use deep learning models to capture complex pattern anomalies in time series data and apply the Isolation Forest algorithm to quickly identify global anomalies.
[0124] The outputs of each model are confidence-weighted and combined to generate a comprehensive anomaly score to reduce false positives and false negatives.
[0125] Step 2.3, fault propagation dynamics model construction;
[0126] This application provides a fault propagation model based on infectious disease propagation theory to quantify the spread of faults in the system:
[0127] Abstract system components into nodes with different states (healthy, latent, faulty);
[0128] Define the state transition function:
[0129] ;
[0130] in, Presentation Component and in time status, 、 Represent components respectively and components In time status, Is with components A collection of directly connected components, is a parameter set; is a state transition function that describes how the component state evolves over time.
[0131] Establishing the fault propagation rate matrix ,matrix It can be expressed as:
[0132] ;
[0133] in, 、 、 Respectively indicate the fault slave component Propagate to components , components and components rate; 、 、 Respectively represent the fault slave components Propagate to components , components and components rate; 、 、 Respectively indicate the fault slave component Propagate to components , components and components rate; is the total number of components in the system. This rate value reflects the speed of fault propagation, and a larger value indicates faster propagation.
[0134] The fault propagation dynamics model provided in this application is specifically implemented as follows:
[0135] State transfer function implementation: The state transfer function of this application adopts an improved SEIR (susceptible-latent-infected-recovered) model framework and is customized according to the characteristics of the computing system.
[0136] The complete state transition function is:
[0137] ;
[0138] in, represents the state value of component i at time t+1, represents the current state value of component i at time t, is the global propagation parameter, which controls the coefficient of the overall fault propagation speed, represents the set of all components directly connected to component i, is the specific propagation rate from node j to node i, indicating the specific rate at which a fault propagates from j to i, is the fault state intensity of node j at time t. The larger the value, the more serious the fault. is the spontaneous failure rate, which represents the probability coefficient of the component itself causing failure. is the load factor of node i, reflecting the current load level of the component. The higher the load, the greater the possibility of spontaneous failure.
[0139] Step 2.4, counterfactual reasoning and failure risk assessment;
[0140] Apply counterfactual reasoning techniques to assess the causal impact between system components through hypothetical interventions:
[0141] Constructing a counterfactual model based on structural equations:
[0142] ;
[0143] in, Indicates that when the variable Intervention as value When the variable The counterfactual value of ; is the counterfactual function used to calculate the outcome after the intervention; It is a variable intervention value; Indicates background conditions Next, the variable All parent nodes except The values of other variables besides ; Representation and variables relevant exogenous factors or random disturbances;
[0144] For each important component, calculate the probability of its impact on other components under failure conditions:
[0145] ;
[0146] in, represents the probability function; Indicates the target component to be affected; Indicates the component Intervene to place it in a faulty state; Indicates in the component The component is intervened in the fault state. Status; Indicates in the component In case of failure, components There is also a probability of failure.
[0147] Based on the impact probability and component importance, a risk score is calculated for each component:
[0148] ;
[0149] in, Presentation Component A risk score that quantifies the potential impact of component failure on the entire system; Represents the set of all components in the system; Indicates that the system may be affected by components Any component affected by the failure; Presentation Component the business importance of the Indicates in the component In the event of a failure, the component The conditional probability of failure also occurring; Indicates the component An operation that causes human intervention to cause it to fail.
[0150] In some embodiments, counterfactual reasoning can be extended to a multi-step reasoning process to assess the potential impact depth of a fault chain reaction.
[0151] For example, in addition to calculating the direct impact, we can also evaluate the secondary and tertiary chain reactions, which can be formalized through recursive counterfactual queries as follows:
[0152] ;
[0153] in, represents the risk score of component X after considering the n-step chain reaction; Represents any component in the system that may be affected; Represents the set of all components in the system; Indicates the business importance of component Y; represents the conditional probability that component Y will fail after n steps of chain reaction after component X fails; Indicates the operation of human intervention on component X to put it into a fault state; Indicates the number of steps in the chain reaction, used to distinguish direct effects (n = 1) from deeper indirect effects (n ≥ 2).
[0154] In practical applications, this multi-step reasoning can discover deep-level risk links that are difficult to identify with traditional methods, especially in highly complex microservice architectures.
[0155] The output of step 2 includes: anomaly detection results, fault propagation model, node risk score, and system risk heat map. This information together constitutes a comprehensive assessment of the system health status and provides a basis for subsequent fault prevention and resource scheduling decisions.
[0156] Step 3: When an anomaly is detected or a potential fault is predicted, a critical path analysis is performed using a multi-level system dependency graph and causal relationship model to identify the fault propagation path and key impact nodes. A fault isolation strategy is then implemented to block fault propagation and implement preventive resource scheduling.
[0157] It includes the following sub-steps:
[0158] Step 3.1, critical path analysis and fault isolation;
[0159] Based on the multi-level dependency graph and risk assessment results, we identify key nodes and links in the system and implement targeted isolation measures:
[0160] Apply graph theory algorithms (such as the minimum cut set algorithm) to identify critical paths and bottlenecks in fault propagation;
[0161] Calculate the isolation value index of each node based on the connectivity and business importance of the components:
[0162] ;
[0163] in, Representation node The isolation value index is used to quantify the cost-effectiveness of isolating the node; represents the i-th node or component in the system; Indicates isolated nodes The number of nodes that can be prevented from being affected by the failure, that is, the total number of other nodes that can be protected after isolating the node; This ratio indicates the number of critical business nodes that may be affected by isolation, that is, the scope of business interruption that may result from isolating the node. A higher ratio indicates that isolating the node can achieve greater fault protection with less business impact.
[0164] Precisely isolate components with high IVI values to form a "firewall" to block fault propagation paths;
[0165] A circuit breaker mechanism is used to proactively cut off non-critical dependencies that communicate with a component when a component anomaly is detected but not yet a complete failure, preventing the abnormal state from spreading.
[0166] In some implementations, critical path analysis can apply community detection algorithms to partition system components into multiple semi-autonomous functional communities, and then enforce isolation measures at the community boundaries. For example, in a large-scale microservice architecture, modularity maximization algorithms (such as the Louvain method) can be used to identify closely collaborating service groups and construct a multi-level fault isolation strategy based on the community structure.
[0167] Furthermore, in financial or medical systems requiring high availability, isolation mechanisms can adopt a progressive strategy, starting with "observation mode," gradually escalating to "throttling mode," and finally to full "isolation mode." Each stage is based on the anomaly confidence level and business impact assessment based on real-time monitoring. This progressive isolation balances fault protection with business continuity, reducing unnecessary service interruptions caused by misjudgments.
[0168] Step 3.2, risk-aware resource allocation decision;
[0169] This application provides a method to integrate risk assessment results into the resource scheduling decision-making process to avoid assigning tasks to high-risk nodes:
[0170] Define the node health score:
[0171] ;
[0172] in, Representation node Health score, used to quantify the reliability and stability of the node; represents the i-th node or component in the system; Representation node The failure risk score of the node quantifies the possibility of failure of the node. The health score is inversely proportional to the risk score. The higher the health score, the more stable and reliable the node. The health score ranges from 0 to 1, where 1 indicates complete health (zero risk) and 0 indicates extremely high risk.
[0173] According to the criticality and resource demand characteristics of the task type, a task node matching strategy is formulated to maintain a "healthy node whitelist" for key tasks, ensuring that key businesses are only assigned to nodes with a health score exceeding the threshold. Node;
[0174] Design a load balancing algorithm that takes failure risks into account, balancing system load while avoiding overloading high-risk areas;
[0175] Achieve a balance between task affinity and risk aversion to optimize task allocation decisions;
[0176] Step 3.3, preventive resource migration;
[0177] For tasks already running on high-risk nodes, proactively implement preventive migration and transfer them to healthy nodes:
[0178] Calculate the migration priority score based on resource utilization and task priority:
[0179] ;
[0180] in, Indicates that the task Slave nodes Migrate to Node Priority score; represents the i-th task in the system; Indicates the source node where the task is currently located; Indicates the target node to which the task will be migrated; Indicates a task Business priority, where a higher value indicates a more important task; Represents the source node The risk score of the node is higher, where a higher value indicates a greater risk of node failure. Indicates the target node The health of the node. A higher value indicates a more stable and reliable node. Represents the migration cost coefficient, ranging from 0 to 1. A higher value indicates a greater migration cost. An inverse indicator of migration cost, where higher values indicate more economically feasible migration.
[0181] Implement a lightweight task state preservation and recovery mechanism to support seamless migration of tasks between different nodes;
[0182] Design an incremental migration strategy. For large tasks, migrate critical status first, then gradually migrate non-critical data to minimize the impact of the migration on services.
[0183] Establish a migration storm protection mechanism to limit the number of concurrent migration tasks and avoid system instability caused by a large number of simultaneous migrations.
[0184] In some implementations, preventative resource migration can leverage historical migration performance data to build a migration time prediction model, accurately assessing the migration window in advance. For example, in stateful service migrations, the system can use a regression model to predict the time required to complete the migration based on factors such as task type, data volume, and network bandwidth. This model can then optimize the migration schedule accordingly, ensuring that critical tasks are migrated before risk thresholds are breached.
[0185] In addition, for critical tasks that cannot be quickly migrated (such as long-running data processing jobs), the system can adopt a "shadow instance" strategy;
[0186] A replica instance of the task is launched on a healthy node. It runs in parallel with the original instance for a period of time, but does not provide external services. After data and status synchronization are complete, traffic is quickly switched at an appropriate time. This strategy is particularly suitable for scenarios with extremely high service continuity requirements, such as database master-slave switchover.
[0187] Step 3.4, multi-level resource reservation and priority management;
[0188] Build a multi-level resource reservation pool to allocate resource guarantees of different qualities to tasks of different priorities;
[0189] Implement strict resource reservation for the highest priority tasks to ensure that there are sufficient resources to execute critical business in any situation;
[0190] Adopt a flexible reservation strategy for medium-priority tasks, dynamically adjusting the proportion of reserved resources based on system load and risk conditions;
[0191] Implement resource reservation recovery and reallocation mechanism. When the idleness of reserved resources exceeds the threshold, they are temporarily released to low-priority tasks, but the preemption rights are retained.
[0192] Build a resource scheduling priority map to quickly determine task priorities and make reasonable allocation decisions when resources compete.
[0193] The output of step 3 is a set of proactive protection measures, including fault isolation strategies, risk-aware resource allocation decisions, preventative task migration plans, and multi-level resource reservation schemes. Together, these measures form a multi-layered defense to prevent fault propagation and ensure critical business continuity.
[0194] Step 4: Implement intelligent repair and resource reorganization based on fault isolation results and business continuity requirements;
[0195] It includes the following sub-steps:
[0196] Step 4.1, fault root cause analysis;
[0197] Use the causal relationship model to locate the root cause of the fault:
[0198] Construct a causal graph (CausalGraph) of failure events based on a multi-level system dependency graph;
[0199] Apply causal inference algorithms to observed performance anomalies and alarm events to calculate the posterior probability of each potential root cause;
[0200] Intervention analysis was performed using an optimized Pearl's do-calculus to assess the strength of the causal relationship between potential root causes and observed symptoms through a virtual intervention (do-operator);
[0201] Introducing prior knowledge of time series, considering event time series constraints, and improving the accuracy of root cause inference;
[0202] Combine expert rules and data-driven approaches to handle complex cascading failures and identify initial triggering events
[0203] In some embodiments, root cause analysis can integrate uncertainty representation and reasoning capabilities to generate multiple hypothesized root causes and their probability distributions in the initial diagnosis stage, and continuously update the probabilities as new evidence is collected.
[0204] For example, the system can use Bayesian belief networks or Markov logic networks to represent the uncertainty of fault diagnosis and use causal probabilistic reasoning to calculate conditional posterior probabilities. , supporting decision making under incomplete information conditions.
[0205] in, Indicates that evidence is observed In the case of The conditional probability of being the real cause of the failure; represents the possible failure root dependent variable; Indicates the A specific root cause hypothesis; Represents the fault evidence variable collected by the system; Indicates the specific observed fault evidence value in the test environment.
[0206] Furthermore, the system implements adaptive diagnostic paths, automatically selecting the most effective diagnostic strategy based on initial symptom characteristics. For example, for network failures, it prioritizes checking connection status and routing information; for resource exhaustion failures, it prioritizes analyzing resource usage trends and mutation points. This context-aware diagnostic strategy selection significantly reduces mean time to diagnosis. In large distributed systems, root cause identification can be reduced from tens of minutes with traditional methods to just a few minutes.
[0207] Step 4.2, intelligent repair strategy generation;
[0208] Retrieve potential solutions from a library of remediation strategies for the identified root causes;
[0209] Use a combination of case-based reasoning (CBR) and reinforcement learning to generate specific repair steps based on historical cases and real-time system status;
[0210] Establish a repair strategy evaluation model and calculate the comprehensive score for each candidate strategy:
[0211] ;
[0212] in, Indicates repair strategy The comprehensive evaluation score of the strategy, the higher the score, the better the strategy; represents the i-th candidate repair strategy; Indicates repair strategy The expected effect, the higher the value, the better the repair effect; Indicates repair strategy Implementation costs, including consumption of computing resources and network bandwidth; Indicates repair strategy The execution risk of the product is higher, and the higher the value, the greater the negative impact. Indicates repair strategy Execution time, which is the time required from the start to the completion of the repair; 、 、 、 They represent the weight coefficients of expected effect, implementation cost, execution risk and execution time respectively.
[0213] Considering the current system status and business needs, the evaluation weights are dynamically adjusted, giving priority to quick repair strategies in time-sensitive scenarios;
[0214] Design a parallel repair task planning algorithm to identify repair steps that can be executed in parallel and shorten the total repair time;
[0215] Step 4.3, monitoring and adjustment of the repair process;
[0216] Establish a key performance indicator (KPI) monitoring system for the repair process to evaluate the repair progress in real time;
[0217] Define the expected model of repair effect and set the expected system status improvement curve at each stage;
[0218] Design an adaptive adjustment mechanism for repair solutions, automatically starting solution optimization when it detects that the actual effect deviates from expectations;
[0219] Build a repair impact prediction model to assess the risk of secondary failures that may be caused by repair activities;
[0220] Implement a repair rollback trigger to enable timely rollback when a repair operation is found to have a negative impact.
[0221] In some embodiments, the system can implement a semi-automatic repair framework for human-machine collaboration.
[0222] For low-risk, standardized repair steps (such as restarting services and clearing caches), the system can perform them fully automatically;
[0223] For high-risk or non-standard operations, the system generates detailed operational recommendations, which are reviewed and executed by human experts.
[0224] The system can also learn from the manual operations of human experts, using operation logs and interactive feedback to gradually improve its automatic repair capabilities. This progressive automation strategy can gradually improve the level of repair automation while ensuring safety.
[0225] Furthermore, the system repair process enables multi-granular rollback checkpoints, automatically creating system state snapshots before critical repair steps and supporting fine-grained selective rollback. For example, before high-risk operations like database schema changes, not only is a data backup saved, but the configuration status of all services that rely on the database is also recorded, ensuring a full recovery to a safe state when necessary.
[0226] Step 4.4, resource reorganization and system optimization;
[0227] Identify system vulnerabilities and optimization opportunities based on fault characteristics and root cause analysis results;
[0228] Implement resource rebalancing to adjust load distribution and avoid resource hot spots and single point failure risks;
[0229] Optimize the system dependency structure, reduce unnecessary cross-node dependencies, and improve the system's modularity and fault isolation capabilities;
[0230] Implement resource redundancy or functional reconstruction for components that frequently fail based on historical failure data;
[0231] Establish a flexible resource pool adjustment mechanism to dynamically adjust the allocation ratio of various resources based on failure modes and business needs.
[0232] In some implementations, the system can perform simulation-based "what-if" analysis, validating the effectiveness of optimization proposals within a digital twin environment before implementing major structural adjustments. For example, for a planned microservices architecture reorganization, the system can build a simulation model based on historical load and failure data to predict post-reorganization performance and fault recovery capabilities. This simulation-driven optimization decision-making can significantly reduce the risk of architectural changes.
[0233] Furthermore, in long-running systems, resource reorganization can adopt an incremental refactoring strategy, breaking down large-scale reorganizations into multiple, independent, incrementally implementable small changes. After each change, system performance is observed, feedback is collected, and subsequent plans are adjusted. This incremental optimization approach is particularly suitable for critical business systems that cannot withstand large-scale downtime. It can gradually improve system resilience and fault tolerance while ensuring business continuity.
[0234] The output of step 4 is a fault root cause identification report, a customized repair strategy, repair process monitoring data, and a resource reorganization optimization plan, which together achieve rapid recovery from system failures and continuous optimization of long-term health.
[0235] Application examples of this implementation:
[0236] For example, a large-scale financial services cloud platform comprises approximately 3,000 computing nodes, supporting multiple critical business systems such as transaction settlement, risk analysis, and user services. The platform faces major challenges: instability caused by surges in system load during peak trading periods, the risk of cascading failures caused by complex dependencies between components, and prolonged service interruptions caused by traditional fault handling methods.
[0237] Implementation process example:
[0238] Examples of building multi-level system dependency graphs and causal relationship models:
[0239] First, the system collected dependency data for more than 15,000 components, including hardware devices, virtual machines, middleware, microservices, etc. Through API calls, log analysis, and network traffic monitoring, it discovered many dynamic dependencies that were not recorded in the configuration database.
[0240] When building a multi-level dependency graph, the system divides components into four main levels:
[0241] Hardware layer: 540 physical servers, 42 network devices, and 38 storage devices;
[0242] Infrastructure layer: 2,500 virtual machines, 780 container clusters;
[0243] Platform layer: 120 database instances, 85 message queues, and 60 cache clusters;
[0244] Application layer: 650 microservices, 1,800 API endpoints.
[0245] The system applied a hybrid causal discovery algorithm to analyze the system's operational data over the past 90 days, successfully identifying causal relationships between approximately 28,000 pairs of components. During the expert knowledge integration phase, the system incorporated 326 expert rules provided by the operations team and eliminated approximately 4,200 pairs of components that appeared to be correlated but not causally related.
[0246] Specifically, within the transaction system microservice cluster, the system identified a critical causal chain: exhaustion of the payment gateway service's connection pool led to increased response latency in the transaction processing service, which in turn triggered timeout retries in the user session management service, ultimately leading to a surge in database connections. While traditional correlation analysis often treats this causal chain as independent issues, this system uses temporal causal analysis to clarify the propagation sequence and impact paths of these events.
[0247] Multi-dimensional anomaly detection and fault prediction examples:
[0248] The system deploys lightweight monitoring probes on all 3,000 computing nodes. Each probe collects 42 different performance indicators, including hardware indicators such as CPU, memory, I / O, and network, as well as software indicators such as process status and service response time.
[0249] In the anomaly detection phase, the system integrates four different anomaly detection algorithms:
[0250] Statistical models: used to identify anomalies in numerical indicators such as CPU usage and memory consumption;
[0251] DBSCAN density clustering: used to find outliers in multidimensional indicator space;
[0252] Autoencoder: Captures complex pattern changes in time-series data of performance indicators of key services;
[0253] Isolation Forest: quickly screens global outliers and provides preliminary anomaly signals;
[0254] For transaction processing clusters, the system uses an adaptive sampling strategy to increase the sampling frequency (once every 5 seconds) during peak transaction periods and reduce the frequency (once every 30 seconds) during off-peak periods, balancing monitoring accuracy and system overhead.
[0255] In one real-world case, the system detected a small but sustained increase in query latency in a database cluster. While this didn't reach traditional alert thresholds, the multi-model anomaly detection system identified a pattern highly similar to historical failure precursors. The system immediately initiated predictive fault propagation analysis, finding that without intervention, the issue would affect transaction processing services that rely on the database within approximately 15 minutes.
[0256] Examples of fault propagation blocking and preventive resource scheduling:
[0257] Based on the predictions, the system first calculated the critical path for fault propagation and identified the connection pool management component between the database cluster and the transaction processing service as the bottleneck, with an Isolation Value Index (IVI) of 3.8. The system immediately implemented precise circuit breaking measures, routing non-critical query requests to backup read-only instances while maintaining core transaction functionality.
[0258] The system then activated risk-aware resource scheduling, assigning all new transaction processing tasks to nodes with a health score above 0.95. Simultaneously, the system calculated migration priority scores for critical tasks already running on nodes near the affected area, identifying two key account reconciliation tasks for immediate migration.
[0259] The system employed an incremental migration strategy, first migrating the task's memory state and critical data (approximately 120MB) to enable rapid service recovery on the new node. The remaining 4GB of data was then migrated in the background. Throughout the entire process, service interruption lasted no more than 800 milliseconds, well below the 2-second limit allowed by the Service Level Agreement (SLA).
[0260] At the same time, the system reserves an additional 20% of computing resources for high-priority transaction processing tasks and sets preemption permissions to ensure that resources can be quickly acquired in the event of a sudden increase in load.
[0261] Examples of intelligent repair and resource reorganization:
[0262] After completing fault isolation and preventive migration, the system conducted an in-depth root cause analysis. Using a cause-and-effect graph, the system identified the root cause of the increased database query latency as degraded I / O performance of a specific disk partition in the storage subsystem. The system calculated the posterior probabilities of multiple possible root causes and determined that degraded disk performance (with a probability of 0.87) was the most likely cause, significantly higher than connection pool configuration issues (with a probability of 0.12) and network congestion (with a probability of 0.08).
[0263] Based on the root cause analysis results, the system retrieved three potential solutions from the remediation strategy library and calculated the comprehensive scores using the remediation strategy evaluation model:
[0264] Data redistribution plan: (High effect, low risk, but long time);
[0265] Storage hot migration solution: (medium effect, low risk, medium duration);
[0266] Hardware replacement plan: (High effect, high risk, and longest time).
[0267] The system selected a data redistribution solution and broke it down into parallelizable subtasks, while continuously monitoring the performance. During the repair process, the system detected that some table migrations had caused changes in query plans and promptly adjusted the indexing strategy, avoiding potential performance degradation.
[0268] After the repair was completed, the system optimized and reorganized the resource structure, adjusted the storage resource allocation strategy, increased resource redundancy for I / O-intensive workloads, and optimized the dependency structure between the database connection pool and microservices, reducing cross-node calls.
[0269] Technical effect verification:
[0270] This implementation method has achieved significant technical results in the actual application of the financial services cloud platform, mainly in the following two aspects:
[0271] Fault prediction accuracy and lead time: During the six-month operation period, the system successfully predicted 85% of potential system faults and identified fault signs 18.7 minutes in advance on average, providing ample time for preventive measures. Table 1 shows the prediction results for different types of faults:
[0272] Table 1: Statistics of prediction effects of different types of faults;
[0273]
[0274] System availability and resource efficiency: Through proactive prevention and precise isolation strategies, the system significantly improves service availability while optimizing resource utilization efficiency. Table 2 shows the comparison data before and after the application of this technical solution:
[0275] Table 2: Comparative data before and after application of this technical solution;
[0276]
[0277] During a critical business peak, the system successfully predicted and isolated a potential cascading failure chain, controlling a serious incident that could have affected 85% of transaction processing capabilities to a local area, affecting only 2.3% of non-core businesses, ensuring the continuity of core transaction business, and avoiding potential losses estimated to exceed one million US dollars.
[0278] It can be seen from the above real application examples that the distributed computing power scheduling intelligent optimization method based on causal reasoning and multi-level dependency graphs in this embodiment can effectively predict and prevent system failures in actual production environments, significantly improve system availability and optimize resource utilization efficiency, and provide a powerful fault management and resource scheduling solution for large-scale distributed computing environments.
[0279] The above describes an embodiment of the present invention, but this embodiment is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Ordinary technicians in this field can also make more forms of equivalent embodiments based on the inspiration of this embodiment, all of which are protected by this embodiment.
Claims
1. A distributed computing power scheduling intelligent optimization method, characterized in that: The following steps are involved: Construct a multi-layer system dependency graph and causal relationship model. The multi-layer system dependency graph includes the hardware layer, virtualization layer, application service layer, and business function layer. The causal relationship model describes the dependencies and impact paths within and between each layer. The causal relationship model is constructed using a hybrid causal discovery algorithm that combines the constraint-based PC algorithm and the scoring-based GES algorithm. The PC algorithm identifies possible causal relationships through conditional independence tests and constructs an undirected skeleton structure. The GES algorithm uses a Bayesian scoring function to evaluate the likelihood of different graph structures and search for the optimal model in the search space. Based on a multi-level system dependency graph and causal relationship model, multi-dimensional anomaly detection and fault prediction technology are used to monitor system status in real time and predict potential faults. Among them, multi-dimensional anomaly detection uses an integrated multi-model anomaly detection system that combines statistical models, DBSCAN density clustering, autoencoders, and isolation forest algorithms. The outputs of each model are combined through confidence weighting to generate a comprehensive anomaly score. Fault prediction uses an improved SEIR fault propagation dynamics model, and the state transition function is: ; in, represents the state value of component i at time t+1, represents the current state value of component i at time t, is the global propagation parameter, which controls the coefficient of the overall fault propagation speed, represents the set of all components directly connected to component i, is the specific propagation rate from node j to node i, indicating the specific rate at which a fault propagates from j to i, is the fault state intensity of node j at time t. The larger the value, the more serious the fault. is the spontaneous failure rate, which represents the probability coefficient of the component itself causing failure. is the load factor of node i, reflecting the current load level of the component. The higher the load, the greater the possibility of spontaneous failure. When an anomaly is detected or a potential fault is predicted, a critical path analysis is performed using a multi-level system dependency graph and causal relationship model to identify the fault propagation path and key impact nodes. Fault isolation strategies are then implemented to block fault propagation and implement preventive resource scheduling. Fault isolation strategies are determined by calculating an isolation value index: ; in, Representation node Isolation Value Index; represents the i-th node or component in the system; Indicates isolated nodes The number of nodes that can be protected from failures; Indicates the number of important business nodes that may be affected due to isolation; Preventive resource scheduling uses counterfactual reasoning techniques to assess the probability of component failure affecting other components: ; in, represents the probability function; Indicates the target component to be affected; Indicates the component Intervene to place it in a faulty state; Indicates in the component The component is intervened in the fault state. Status; Indicates in the component In case of failure, components Also the probability of failure; Based on the fault isolation results and business continuity requirements, intelligent repair and resource reorganization are implemented. Intelligent repair uses a repair strategy evaluation model to calculate a comprehensive score: ; in, Indicates repair strategy Comprehensive evaluation score; represents the i-th candidate repair strategy; Indicates repair strategy The expected effect, the higher the value, the better the repair effect; Indicates repair strategy implementation costs; Indicates repair strategy The execution risk of the product is higher, and the higher the value, the greater the negative impact. Indicates repair strategy Execution time, which is the time required from the start to the completion of the repair; 、 、 、 They represent the weight coefficients of expected effect, implementation cost, execution risk and execution time respectively.
2. A distributed computing power scheduling intelligent optimization method according to claim 1, characterized in that: The steps to build a multi-level system dependency graph and causal relationship model include: Collect component information and status data at all levels in the system; Analyze historical data through data mining technology to identify the dependencies between components; Combine expert knowledge and machine learning algorithms to build a causal relationship model between components; Regularly update multi-level system dependency diagrams and causal models to adapt to system changes.
3. The distributed computing power scheduling intelligent optimization method according to claim 1, characterized in that: The multi-dimensional anomaly detection and fault prediction steps include: Establish a multi-dimensional normal behavior baseline for key system indicators based on historical data and expert knowledge; Monitor system status indicators in real time and use multiple anomaly detection algorithms to identify abnormal patterns that deviate from normal behavior; Predict the probability, impact range, and severity of potential failures based on time series analysis and causal reasoning; Generate anomaly detection and fault prediction reports to provide a basis for further decision-making.
4. The distributed computing power scheduling intelligent optimization method according to claim 1, characterized in that: The steps for blocking fault propagation and implementing preventive resource scheduling include critical path analysis and fault isolation: Analyze the propagation path of anomalies or potential faults based on multi-level dependency graphs; Calculate the risk assessment results of each propagation path and identify key nodes and links that require priority treatment; Develop and implement targeted isolation measures to block the propagation of faults; Implement a hierarchical isolation strategy for businesses requiring high availability to ensure the continuity of core businesses.
5. The distributed computing power scheduling intelligent optimization method according to claim 1, characterized in that: The preventive resource scheduling steps include: Adjust resource allocation strategies based on failure risk assessment results; Perform health scoring on tasks on high-risk nodes and calculate migration priority scores based on the scoring results and priorities; Create shadow instances for critical tasks to achieve seamless switching; Implement multi-level resource reservation and priority management to ensure resource requirements for high-priority tasks.
6. The distributed computing power scheduling intelligent optimization method according to claim 1, characterized in that: The steps of intelligent repair and resource reorganization include: Automatically select the appropriate repair strategy based on the fault type and system status; Execute automatic repair processes for faults that can be repaired automatically; Generate detailed fault reports and repair suggestions for complex faults that require manual intervention; Reallocate and optimize resource configuration based on the repaired system status and business needs.
7. The distributed computing power scheduling intelligent optimization method according to claim 1 is characterized in that: It also includes continuous learning and optimization steps: Collect and analyze data and results during troubleshooting; Updated causality models and anomaly detection baselines; Optimize fault prediction algorithms and resource scheduling strategies; Generate optimization reports for long-term system improvement.
8. A distributed computing power scheduling intelligent optimization system, used to execute a distributed computing power scheduling intelligent optimization method according to any one of claims 1 to 7, characterized in that: include: The system perception and modeling module is used to collect component information and status data at all levels of the system, build multi-level system dependency graphs and causal relationship models, and monitor system status in real time to identify anomalies and predict potential failures; Fault analysis and isolation module, used to analyze fault propagation paths and identify key nodes, develop and implement targeted isolation strategies, and block fault propagation; Resource scheduling and repair module, which is used to perform preventive resource scheduling based on risk assessment results and implement intelligent repair strategies according to fault type and system status; The continuous learning and optimization module is used to collect and analyze fault handling data, update the causal relationship model and anomaly detection baseline, and continuously optimize the system model and scheduling strategy.
9. A distributed computing power scheduling intelligent optimization system according to claim 8, characterized in that: Also includes decision support modules for: Comprehensive analysis of anomaly detection, fault prediction, and resource status information; Evaluate the costs and benefits of different treatment strategies; Make processing decisions automatically or semi-automatically based on preset strategies and system status; Record the decision-making process and results for subsequent learning and optimization.
10. A distributed computing power scheduling intelligent optimization system according to claim 8, characterized in that: The distributed computing power scheduling intelligent optimization system is deployed using a distributed architecture and includes: Central control node, responsible for global decision-making and coordination; Multiple edge nodes, distributed throughout the system, are responsible for local data collection and preprocessing; Data sharing mechanism ensures real-time synchronization of information between nodes; Fault-tolerant mechanism ensures that the system can still operate normally when some nodes fail.
Citation Information
Patent Citations
Microservice intelligent operation and maintenance system and method oriented to cloud native and application
CN117009119A
Automatic management method and system for disaster recovery process of intelligent calculation center
CN119003249A