Cloud monitoring service operation and maintenance dynamic optimization system and method based on AI intelligent agent
By introducing AI agents into the cloud monitoring system, unifying the acquisition and analysis of cross-layer indicators, identifying implicit fault associations and optimizing resource scheduling and fault repair strategies, the problem of insufficient indicator fragmentation and fault association identification capabilities in existing cloud monitoring solutions is solved, and more efficient fault warning and root cause positioning are achieved.
Patent Information
- Application Number
- CN202510698823.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-28
AI Technical Summary
Existing cloud monitoring solutions rely on static rule bases or single data sources, resulting in cross-layer indicator fragmentation, insufficient ability to identify implicit fault associations, high false alarm rate, due to positioning delay, and lack of coordination between resource scheduling strategies and fault repair actions, which can easily cause cascading failure risks.
The dynamic optimization system for cloud monitoring service operation and maintenance based on AI agents is adopted. Through the cross-layer index perception module, the spatio-temporal correlation analysis module, the causal verification and coupling degree analysis module, the dynamic priority decision module and the policy execution and feedback module, the unified data collection and analysis of the physical layer, the virtual layer and the application layer are realized, implicit fault associations are identified, and resource scheduling and fault repair strategies are optimized.
It improves the accuracy of fault warning and the accuracy of root cause positioning, reduces the false alarm rate and cascading fault risk, and enhances the resilience of the system and the adaptability of operation and maintenance strategies.
Smart Images

Figure CN120223501A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cloud computing intelligent operation and maintenance, and particularly to a cloud monitoring service operation and maintenance dynamic optimization system and method based on an AI intelligent agent. Background Art
[0002] Under the background of the rapid development of current information technology, cloud computing has become one of the core infrastructure for enterprise digital transformation. Especially in the trend of the increasing popularity of hybrid clouds and multi-cloud environments, the types of services carried by cloud platforms are becoming more complex, the service dependency chains are significantly lengthened, and the system operation status shows characteristics of high dynamics, high concurrency, and multi-level coupling. To ensure the continuity and stability of critical business systems, cloud platform operation and maintenance management is gradually evolving from static monitoring to intelligent, automated, and dynamic optimization. There is an urgent need to actively identify and respond to scheduling potential fault risks in complex systems through means such as AI intelligent agents, so as to improve the overall service quality guarantee level.
[0003] However, existing cloud monitoring solutions mostly rely on static rule libraries or single data source drivers, and there are problems such as cross-layer index fragmentation and insufficient ability to identify implicit fault associations. For example, it is difficult to effectively verify the causal link between physical layer resource fragmentation and application layer service performance deviation, resulting in a high false alarm rate and a delay in root cause location. The resource scheduling strategy and the fault repair action lack coordination, and it is easy to increase the risk of cascading failures due to blind capacity expansion. In addition, the contradiction between short-term emergency response and long-term optimization goals lacks a dynamic trade-off mechanism, resulting in poor adaptability of operation and maintenance strategies to the actual scenario and difficulty in meeting the elastic requirements in the hybrid cloud environment. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the present invention provides a cloud monitoring service operation and maintenance dynamic optimization system and method based on an AI intelligent agent, which solves the problems in the above background art.
[0005] To achieve the above objectives, the present invention is realized through the following technical solutions: A cloud monitoring service operation and maintenance dynamic optimization system based on an AI agent, comprising the following steps: a cross-layer metric perception module, a spatio-temporal correlation analysis module, a causality verification and coupling degree analysis module, a dynamic priority decision module, and a policy execution and feedback module; The cross-layer metric perception module is used to deploy a containerized probe cluster to collect in real time the fragmentation metrics of physical layer hardware resources, container lifecycle events in the virtual layer, and performance deviation data of microservice call chains in the application layer in a hybrid cloud environment. Through hierarchical tagging preprocessing and outlier filtering, a standardized cross-layer metric set is output; The spatio-temporal correlation analysis module is used to connect to the cross-layer metric perception module, eliminate the temporal deviation between the low-frequency data of the physical layer and the high-frequency data of the application layer through an incremental time series alignment algorithm, construct a fault propagation probability map based on a dynamic time window, and identify hidden correlation nodes with spatio-temporal coupling characteristics among cross-layer metrics; The causality verification and coupling degree analysis module is used to receive the fault propagation map output by the spatio-temporal correlation analysis module, verify the authenticity of cross-layer causality through directed perturbation injection, fit the coupling degree index according to the correlation matrix between resource scheduling operations and the fault propagation chain, and construct coupling degree analysis rules to quantify the influence weight of operations on fault propagation, generating root cause location instructions for negative coupling scenarios; The dynamic priority decision module is used to analyze the historical fault repair time and resource waste rate to fit the propagation cost gradient according to the fault propagation cost model, input the coupling degree index, the real-time service level agreement default rate, and the fault propagation cost gradient into a weight function, calculate the collaborative priority of short-term suppression actions and long-term eradication actions through a dynamic weight function, and generate an execution sequence of the two types of actions through an asymmetric game strategy; The policy execution and feedback module is used to call the cloud platform interface to atomically execute decision-making actions, collect the changed metric data after execution, and dynamically update the fault propagation cost model and coupling degree analysis rules.
[0006] Further, the cross-layer metric perception module specifically includes: The containerized probe cluster is deployed in a microservice architecture on hybrid cloud nodes, including physical layer probes to collect hardware resource fragmentation metrics, virtual layer probes to monitor container lifecycle events and resource contention behaviors, and application layer probes to track microservice call chain topologies and performance deviations; The collection frequency is adjusted according to the dynamic change rate of the metrics. The physical layer uses low-frequency trigger sampling, and the application layer uses event-driven high-frequency tracking, and instantaneous noise is suppressed through a sliding window statistic; The hierarchical tagging preprocessing unit attaches environment context tags of the cloud platform type and service dependency level to the original data, filters outliers based on the isolation forest algorithm, and outputs a standardized cross-layer metric set.
[0007] Furthermore, the time sequence deviation between the low-frequency data of the physical layer and the high-frequency data of the application layer is eliminated through the incremental time sequence alignment algorithm. The specific process of constructing the fault propagation probability map based on the dynamic time window is as follows: Dynamically interpolate the low-frequency data of the physical layer and the high-frequency data of the application layer, and generate a continuous time sequence weighted according to the data confidence; Detect the potential phase difference between cross-layer indicators through sliding correlation analysis, and dynamically adjust the interpolation anchor points to eliminate the time sequence offset; Calculate the conditional transition probability between cross-layer indicators based on the aligned time sequence data, dynamically expand the time window range, and shrink the window when an indicator mutation is detected, generating a fault propagation probability map with weighted edges.
[0008] Furthermore, the identification logic for identifying latent association nodes with spatio-temporal coupling characteristics between cross-layer indicators is as follows: Extract the propagation paths across the physical layer, virtual layer, and application layer in the fault propagation probability map, and screen the candidate nodes with transfer probabilities exceeding the dynamic threshold; Conduct mutual information entropy analysis on the candidate nodes to quantify their dependence intensity on upstream and downstream indicators, and eliminate weak association interference terms; Verify the spatio-temporal causality directionality of the candidate paths through Granger causality test, and inject directional perturbations into the high-probability paths to observe the response amplitude of downstream indicators to confirm the effectiveness of spatio-temporal coupling.
[0009] Furthermore, the specific process of verifying the authenticity of cross-layer causality through directional perturbation injection and fitting the coupling degree index according to the correlation matrix between resource scheduling operations and the fault propagation chain is as follows: Select high-probability association paths in the fault propagation map, and inject controllable perturbations such as a sudden increase in simulated physical layer storage delay or restricting the virtual layer container network bandwidth into the source nodes of the paths; Observe the downstream indicator responses and record the perturbation propagation paths and amplitudes, and compare the consistency with the predicted paths in the original map; Extract the time sequence relationship between historical resource scheduling operations and the fault propagation chain, construct the correlation matrix between resource scheduling operations and the fault propagation chain, and based on the influence weights of resource scheduling operations on the fault chain length and repair time in the matrix, fit the coupling degree index through the gradient descent method.
[0010] Furthermore, the specific process of constructing the coupling degree analysis rule to quantify the influence weight of operations on fault propagation and generating the root cause location instruction for the negative coupling scenario is as follows: Based on the correlation matrix between resource scheduling operations and the fault propagation chain, calculate the contribution weights of operations to the fault chain length and repair time, and generate the initial coupling degree analysis rule; Set the dynamic determination threshold, and when the contribution weight of the operation to the fault propagation chain exceeds the threshold, mark it as a negative coupling operation; Trace back the nodes associated with the negative coupling operation in the fault propagation map, and screen the high-causality intensity nodes not covered by historical scheduling operations as candidate root causes; Conduct a reverse blocking test on the candidate root causes, and verify their interruption effect on the downstream fault chain by restricting their resource access or traffic distribution; Generate the location instruction containing the root cause node identifier, influence path, and repair suggestions, and push it to the operation and maintenance terminal.
[0011] Furthermore, the specific process of fitting the propagation cost gradient of the historical fault repair time and the resource waste rate according to the fault propagation cost model is as follows: Extract the historical fault repair time data and the resource waste rate caused by resource scheduling operations, and construct an initial propagation cost function; Iteratively optimize the cost function parameters by the gradient descent method, and dynamically adjust the repair time weight and the resource waste penalty factor; Calculate the current propagation cost gradient according to the change of the real-time fault propagation path length and the resource utilization rate; Continuously optimize the gradient parameters through the policy execution feedback data to adapt to the dynamic changes of the hybrid cloud environment.
[0012] Furthermore, the specific process of inputting the coupling degree index, the real-time service level agreement default rate, and the fault propagation cost gradient into the weight function, calculating the collaborative priority of the short-term suppression action and the long-term eradication action through the dynamic weight function, and generating the execution sequence of the two types of actions through the asymmetric game strategy is as follows: Introduce a coupling degree penalty factor in the dynamic weight function to suppress the priority of scheduling operations that may exacerbate fault propagation in high-coupling scenarios; Calculate the urgency weight of the short-term suppression action based on the real-time service level agreement default rate, and calculate the benefit weight of the long-term eradication action in combination with the propagation cost gradient; Define the short-term suppression action and the long-term eradication action as asymmetric game participants, and construct a benefit function that quantifies their contributions to service availability improvement and fault propagation suppression; Solve the optimal collaborative strategy through the dynamic Nash equilibrium, give priority to executing the short-term suppression action to quickly stop losses, and asynchronously trigger the long-term eradication action.
[0013] Furthermore, the policy execution and feedback module specifically includes: The atomic execution engine disassembles the decision-making action into independently executable atomic operations that call the cloud platform API to trigger the traffic limiting policy or initiate the storage volume migration task, and ensures the atomicity and consistency of cross-platform operations through the transaction lock mechanism; Monitor the cross-layer metric changes after execution, capture the suppression effect of the action on the fault propagation chain and the impact on the resource utilization rate, adjust the weight parameters in the fault propagation cost model according to the feedback data, optimize the coupling degree analysis rule, and enhance the early recognition ability for negative coupling scenarios.
[0014] A dynamic optimization method for cloud monitoring service operation and maintenance based on an AI agent, comprising the following steps: S1. Deploy a containerized probe cluster to collect in real time fragmentation metrics of physical layer hardware resources, container lifecycle events in the virtual layer, and performance deviation data of microservice call chains in the application layer. Through hierarchical tagging preprocessing and outlier filtering, output a standardized cross-layer metric set; S2. Connect a cross-layer metric perception module, eliminate the timing deviation between low-frequency data in the physical layer and high-frequency data in the application layer through an incremental timing alignment algorithm, construct a fault propagation probability map based on a dynamic time window, and identify hidden association nodes with spatio-temporal coupling characteristics between cross-layer metrics; S3. Receive the fault propagation map output by the spatio-temporal association analysis module, verify the authenticity of cross-layer causal relationships through directed perturbation injection, fit a coupling degree index according to the association matrix between resource scheduling operations and the fault propagation chain, and construct a coupling degree analysis rule to quantify the influence weight of operations on fault propagation, generating root cause location instructions for negative coupling scenarios; S4. Analyze the historical fault repair time and resource waste rate to fit the propagation cost gradient according to the fault propagation cost model, input the coupling degree index, the real-time service level agreement default rate, and the fault propagation cost gradient into a weight function, calculate the collaborative priority of short-term suppression actions and long-term eradication actions through a dynamic weight function, and generate an execution sequence of the two types of actions through an asymmetric game strategy; S5. Call the cloud platform interface to atomically execute decision-making actions, and collect the changed metric data after execution, dynamically updating the fault propagation cost model and the coupling degree analysis rule.
[0015] The present invention has the following beneficial effects:
[0016] (1) A dynamic optimization system for cloud monitoring service operation and maintenance based on an AI agent realizes unified collection and standardized processing of heterogeneous data in the physical layer, virtual layer, and application layer through a containerized probe cluster, breaks through the single-layer data limitation of traditional monitoring tools, and improves the cross-platform resource collaboration ability. Construct a fault propagation map based on an incremental timing alignment algorithm and a dynamic time window, accurately identify the spatio-temporal coupling characteristics between cross-layer metrics, and solve the problem of mis-association caused by static rules. Verify the causal authenticity through directed perturbation injection and a coupling degree analysis rule, reduce the manual troubleshooting cost, and improve the positioning accuracy in complex fault scenarios. Dynamically balance the priorities of short-term suppression actions and long-term eradication actions in combination with an asymmetric game strategy, avoid conflicts between resource scheduling and fault repair, and improve the overall resilience of the system. Dynamically update the model and rules through execution feedback data, realize the continuous iteration of operation and maintenance strategies, and adapt to the dynamic changes of the hybrid cloud environment.
[0017] (2) A dynamic optimization method for the operation and maintenance of cloud monitoring services based on AI agents, which realizes the full-life cycle management of cross-layer metrics from data collection, time series alignment to graph construction, and eliminates the interference of data islands on operation and maintenance decisions. Through the analysis of dynamic time windows and spatio-temporal coupling characteristics, potential cascading fault paths are identified in advance to improve the fault warning ability. By combining directional perturbation and correlation matrix to quantify the influence weight of resource scheduling on fault propagation, the reliability and interpretability of root cause location are enhanced. Based on the dynamic weight calculation of fault propagation cost gradient and real-time service level agreement default rate, a collaborative execution strategy suitable for complex scenarios is generated to reduce the risk of service interruption. Through the execution feedback closed-loop drive model and the adaptive update of rules, it is ensured that the operation and maintenance strategy always fits the actual environment requirements and improves the long-term operation and maintenance efficiency.
[0018] Of course, it is not necessary for any product implementing the present invention to achieve all the above-mentioned advantages simultaneously. Brief Description of the Drawings
[0019] Figure 1 It is a flow chart of a dynamic optimization system for the operation and maintenance of cloud monitoring services based on AI agents of the present invention.
[0020] Figure 2 It is a flow chart of a dynamic optimization method for the operation and maintenance of cloud monitoring services based on AI agents of the present invention. Detailed Embodiments
[0021] In the embodiments of the present application, through a dynamic optimization system and method for the operation and maintenance of cloud monitoring services based on AI agents, aiming at the complex coupling relationship between cross-layer resources and services in a hybrid cloud environment, by integrating heterogeneous data in the physical layer, virtual layer and application layer, and combining dynamic causal verification and game decision-making mechanisms, problems such as data fragmentation, implicit fault association failure, resource scheduling and fault repair action conflicts in traditional operation and maintenance solutions are solved, and a closed-loop control of fault warning, root cause location, strategy optimization and self-healing execution is realized, so as to improve the stability and resource utilization efficiency of the cloud service system.
[0022] The general idea of the solution in the embodiments of the present application is as follows:
[0023] Cross-layer data integration and implicit association mining: Heterogeneous data of physical layer hardware resources, virtual layer container clusters, and application layer microservice links are collected in real time through a containerized probe cluster. The data semantic and collection frequency differences are eliminated by using hierarchical tagging and dynamic time series alignment technologies, and a cross-layer unified standardized index set is constructed; further, based on dynamic time windows and incremental association analysis, a fault propagation probability graph is generated to identify implicit fault paths with spatio-temporal coupling characteristics (such as the cascading effect of storage performance degradation causing microservice call delay), breaking through the data fragmentation limitation of traditional single-layer monitoring.
[0024] Dynamic Causal Verification and Coupling Degree Quantification: Inject directional perturbations (such as simulating a sudden increase in physical layer I / O delay) in the fault propagation graph to verify the authenticity of the causal relationship between cross-layer metrics; combine the historical resource scheduling operations with the correlation matrix of the fault propagation chain, fit the coupling degree index and construct analysis rules to quantify the influence weight of operations on fault propagation (such as the probability of a cascading fault recurrence caused by an expansion operation), and generate root cause location instructions for negative coupling scenarios to solve the problems of low manual troubleshooting efficiency and high false alarm rate.
[0025] Multi-objective Collaborative Decision-making and Closed-loop Optimization: Based on the fault propagation cost model and the real-time service level agreement default rate, calculate the collaborative priorities of short-term suppression actions (flow control, degradation) and long-term eradication actions (storage migration, service reconstruction) through a dynamic weight function; introduce an asymmetric game strategy to balance the resource contention conflicts between the two types of actions and generate the optimal execution sequence. After the strategy is executed, dynamically update the weights of the propagation cost model and the coupling degree analysis rules through feedback data to form an adaptive closed-loop of "perception - decision - execution - optimization", continuously improving the fault self-healing efficiency and resource utilization rate in the hybrid cloud environment.
[0026] Please refer to Figure 1, an embodiment of the present invention provides a technical solution: a cloud monitoring service operation and maintenance dynamic optimization system based on an AI agent, including the following steps: a cross-layer metric perception module, a spatio-temporal correlation analysis module, a causality verification and coupling degree analysis module, a dynamic priority decision module, and a policy execution and feedback module; the cross-layer metric perception module is used to deploy a containerized probe cluster to collect in real time the fragmentation metrics of physical layer hardware resources, container lifecycle events in the virtual layer, and performance deviation data of microservice call chains in the application layer in a hybrid cloud environment, and output a standardized cross-layer metric set through hierarchical tagging preprocessing and outlier filtering; the spatio-temporal correlation analysis module is used to connect to the cross-layer metric perception module, eliminate the temporal deviation between the low-frequency data in the physical layer and the high-frequency data in the application layer through an incremental time series alignment algorithm, construct a fault propagation probability map based on a dynamic time window, and identify hidden association nodes with spatio-temporal coupling characteristics among cross-layer metrics; the causality verification and coupling degree analysis module is used to receive the fault propagation map output by the spatio-temporal correlation analysis module, verify the authenticity of cross-layer causality through directed perturbation injection, fit the coupling degree index according to the association matrix between resource scheduling operations and the fault propagation chain, and construct a coupling degree analysis rule to quantify the influence weight of operations on fault propagation, and generate root cause location instructions for negative coupling scenarios; the dynamic priority decision module is used to analyze the historical fault repair time and resource waste rate to fit the propagation cost gradient according to the fault propagation cost model, input the coupling degree index, the real-time service level agreement default rate, and the fault propagation cost gradient into a weight function, calculate the collaborative priority of short-term suppression actions and long-term eradication actions through a dynamic weight function, and generate an execution sequence of the two types of actions through an asymmetric game strategy; the policy execution and feedback module is used to call the cloud platform interface to atomically execute decision-making actions, collect the changed metric data after execution, and dynamically update the fault propagation cost model and the coupling degree analysis rule.
[0027] In this implementation plan, the cross-layer metric perception module is used to deploy a containerized probe cluster to continuously perceive the key operating status metrics at different levels (physical layer, virtual layer, application layer) in the hybrid cloud environment, perform preprocessing and standardization, and output a set of cross-layer comparable metrics. The containerized probe cluster refers to a group of lightweight monitoring programs deployed based on container technology (such as Docker), which can flexibly adapt to multi-cloud environments and achieve high scalability and low-intrusive data collection. Physical layer hardware resource fragmentation metrics, such as CPU idle fragments and discontinuous memory occupancy, are used to reflect the fragmentation degree of resource usage and will affect the performance scheduling efficiency. Container lifecycle events include status change events such as container startup, pause, migration, and destruction. Micro-service call chain performance offset data refers to the dynamic changes of metrics such as response latency and error rate on the call path between services and is used to analyze performance bottlenecks in the application layer. Hierarchical tagging preprocessing classifies and labels the collected data according to the source level, service category, etc. to improve the context matching ability of subsequent analysis. Outlier filtering eliminates invalid data caused by instantaneous fluctuations, collection errors, or network delays to ensure the quality of input metrics. The spatio-temporal correlation analysis module is responsible for eliminating the time series error between data at different levels and mining the implicit correlation relationships between various metrics under a dynamic time window to construct a fault propagation path. The incremental time series alignment algorithm is a time series data alignment method based on a sliding window, which can dynamically adjust the alignment method of data points during the continuous update of data to achieve synchronous matching between high-frequency and low-frequency metrics. The fault propagation probability graph is a graph structure based on a probability model, where nodes represent metric items, edges represent potential fault transfer paths, and weights represent propagation probabilities. Implicit correlation nodes refer to metric or event nodes that seemingly have no direct correlation on the surface but exhibit highly correlated behaviors under certain conditions or time periods. The causality verification and coupling degree analysis module verifies whether there is a real causal relationship between cross-layer metrics and quantifies the impact of resource operation behaviors on fault propagation, so as to locate the root cause of the fault. Directed perturbation injection injects small changes (such as resource up / downscaling) into specific components or parameters without affecting system stability, observes the response changes of the overall metric chain, and identifies causal relationships. The association matrix between resource scheduling operations and the propagation chain constructs an impact matrix between resource scheduling actions (such as container migration) and metric changes to fit the causal path. The coupling degree index is a numerical index that quantifies the mutual influence intensity between metrics or operations. The higher the value, the stronger the coupling relationship. Negative coupling scenarios refer to situations where certain operations or resource scheduling exacerbate system failures or performance degradation. Root cause location instructions generate instructions based on the above analysis to accurately locate the core nodes or operations that trigger negative propagation in the system. The dynamic priority decision module dynamically ranks the candidate actions considering multiple factors, weighs short-term loss prevention and long-term governance, and formulates an optimal operation and maintenance strategy sequence. The fault propagation cost model analyzes the repair time and resource waste in past fault events and quantifies the evolution cost of faults.Service Level Agreement (SLA) default rate: It refers to the proportion of failure to provide service quality (such as latency, availability) as required by the contract or platform. Propagation cost gradient: It refers to the marginal contribution rate of different nodes or paths to the growth of the overall system cost. Dynamic weight function: It is used to dynamically adjust the influence weights of various input factors according to actual situations (such as service level changes). Asymmetric game strategy: There is an asymmetry in benefits between different actions. Through the game model, the optimal response strategies of both parties are solved to achieve globally optimal collaborative actions. Strategy execution and feedback module: It calls the underlying interfaces of the cloud platform to automatically execute strategy actions and collects feedback data for model iteration to achieve closed-loop optimization. Atomic execution: The strategy is disassembled into the smallest independently executable operation units (such as restarting services, migrating containers, etc.) to ensure system controllability and security. Post-execution metric change data: It includes the metric differences and change rates before and after the execution actions, and is used to evaluate the effectiveness of the actions. Dynamically update the cost model and rules: The execution effects are fed back to adjust the decision-making model and the weights of causal rules in real time to achieve the adaptive evolution of the model.
[0028] Specifically, the cross-layer metric perception module specifically includes: The containerized probe cluster is deployed in a microservices architecture on hybrid cloud nodes. It includes physical layer probes to collect hardware resource fragmentation metrics, virtual layer probes to monitor container lifecycle events and resource contention behaviors, and application layer probes to track microservice call chain topologies and performance offsets; The collection frequency is adjusted according to the dynamic change rate of the metrics. The physical layer uses low-frequency trigger sampling, and the application layer uses event-driven high-frequency tracking. Instantaneous noise is suppressed through sliding window statistics; The hierarchical tagging preprocessing unit attaches environmental context tags of the cloud platform type and service dependency level to the original data, filters outliers based on the isolation forest algorithm, and outputs a standardized set of cross-layer metrics.
[0029] In this implementation solution, the cross-layer metric perception module encapsulates probes in a containerized manner and deploys them on multiple nodes in a hybrid cloud environment according to the microservices architecture. This deployment method enables the probes to be flexibly scaled according to actual monitoring needs, adapt to different cloud platforms (such as private clouds, public clouds, or edge computing nodes), and can independently update and dynamically expand the probe functions. The physical layer probes are mainly deployed in physical servers or bare metal environments to collect the fine-grained usage status of underlying hardware resources, including but not limited to the distributed idle core information of the CPU, the number of non-contiguous allocated areas of memory, and the instantaneous fluctuations of disk access. This information helps to identify the degree of resource fragmentation and provides a basic basis for efficient scheduling. The virtual layer probes are deployed in virtualization platforms or container orchestration systems (such as Kubernetes) to monitor the life cycle events of containers (such as creation, start, termination) and their resource competition behaviors with other containers or processes. By recording situations such as container scheduling failures and resource preemption, it helps to identify potential system bottlenecks or problems with unreasonable resource configurations. The application layer probes are integrated into the service gateways or middleware of each microservice and are responsible for capturing the call chain paths between services and key performance indicators such as the response time and latency changes of each call. This mechanism can restore the service dependency relationship graph and monitor the performance drift between services, helping to quickly locate the root cause of faults or performance anomaly nodes. The cross-layer metric perception module adaptively adjusts the sampling frequency according to the change trends of various metrics. For physical layer metrics with small fluctuations or slow change trends, a low-frequency sampling strategy is adopted to save computing resources; while for application layer metrics with drastic changes such as response time, a high-frequency sampling or event-driven method is adopted to ensure data timeliness and accuracy. The physical layer usually samples periodically at fixed intervals, such as collecting resource utilization once a minute; while the application layer adopts an event-triggered mode, for example, sampling is immediately triggered and encrypted for storage when microservice request response timeouts, error rates surge, etc. This hierarchical sampling strategy takes into account both system resource consumption and monitoring accuracy. To reduce misjudgments caused by instantaneous data anomalies, the module introduces a sliding window mechanism for smoothing. This mechanism aggregates data within a continuous period of time (such as taking the average or median), effectively suppressing noise interference caused by short-term disturbances such as network jitter and load spikes, thereby improving the stability of anomaly recognition. To achieve cross-platform and cross-layer fusion processing of data, the module embeds rich environmental context tags in the original monitoring data, including meta-information such as cloud platform type, service layer to which the node belongs, application name, and geographical deployment location. This tag system helps with subsequent cluster analysis, location analysis, and intra-system correlation analysis. In the data cleaning stage, the module uses the isolation forest algorithm to perform unsupervised elimination of abnormal samples. This method can identify extreme or abnormal data based on the difficulty of isolating samples without prior knowledge, further improving the overall quality and robustness of cross-layer metrics and avoiding interference with system judgment due to single-point anomalies.Finally, the module outputs the processed cross-layer metric data in a unified format, forming a standardized metric set that can be called by downstream systems. The standardization content includes unified metric names, data unit conversion, timestamp format alignment, label structure unification, etc., ensuring that multi-source heterogeneous data can be consistently parsed and utilized.
[0030] Specifically, the incremental time series alignment algorithm is used to eliminate the time series deviation between the low-frequency data of the physical layer and the high-frequency data of the application layer. The specific process of constructing the fault propagation probability map based on the dynamic time window is as follows: perform dynamic interpolation on the low-frequency data of the physical layer and the high-frequency data of the application layer, and generate a continuous time series by weighting according to the data confidence; detect the potential phase difference between cross-layer metrics through sliding correlation analysis, and dynamically adjust the interpolation anchor points to eliminate the time series offset; calculate the conditional transition probability between cross-layer metrics based on the aligned time series data, dynamically expand the time window range, and shrink the window when detecting metric mutations, generating a fault propagation probability map with weighted edges.
[0031] In this implementation scheme, there is usually a time series deviation between the low-frequency data of the physical layer and the high-frequency data of the application layer due to the sampling frequency difference. To eliminate this deviation, it is necessary to perform dynamic interpolation on the low-frequency data and the high-frequency data to generate a continuous time series. Data confidence weighting: The confidence of each data point is determined based on the reliability of the data source. For example, data from different probes may vary in accuracy, and data with higher confidence will be given greater weight. During the interpolation process, these weights are used to adjust the influence of each data point. Interpolation: The interpolation method usually adopts linear interpolation or more complex interpolation methods (such as spline interpolation) to generate continuous time series data according to the weights and data confidence. The goal of this step is to fill the gaps between data and ensure a smooth transition of the data sequence. Once continuous time series data is obtained through interpolation, the next step is to perform sliding correlation analysis to detect the potential time series deviation (i.e., phase difference) between cross-layer metrics. The core idea of this step is to find the time difference between the physical layer and application layer data and dynamically adjust the data alignment. Sliding correlation analysis: This analysis method detects potential phase differences by calculating the correlation between cross-layer metric data. Specifically, the sliding window calculates the correlation at different time points and gradually adjusts the data alignment point (i.e., the interpolation anchor point). If a large offset in the metric values within the time window is found, the interpolation anchor point is adjusted to reduce this deviation. Formula: ; where: and are the metric values of the physical layer and the application layer respectively. is the time delay, representing the possible phase difference. and are and 's mean values. is the size of the time window, which determines the range of the sliding window. Through the correlation calculated by the sliding window, the phase difference between data can be identified and corrected. After completing the alignment of time series data, the next step is to calculate the conditional transition probability between cross-layer metrics, that is, to predict the occurrence probability of the next state based on the current state of the metric. Conditional transition probability: The conditional transition probability between cross-layer metrics describes the probability that a metric transfers from one state to another given the current state. When calculating, the correlation at different time points is considered, and a transition probability model is fitted through historical data. Formula: ; where: is the state of the physical layer, is the next state of the application layer. is the physical layer state and the application layer state occur simultaneously. The denominator is the total number of all possible values, which is used to standardize the probability. Expansion and contraction of the dynamic time window: According to the change of data, the range of the time window will be dynamically adjusted. If a cross-layer metric mutates (for example, a fault or performance anomaly occurs), the time window is narrowed to focus on analyzing recent data to ensure that the instantaneous change at the time of the fault occurrence can be accurately captured. Conversely, when the system is stable, the time window expands to cover data over a longer time period to capture the overall trend of the system. Formula: ; where: is the size of the time window. and are the minimum and maximum sizes of the time window respectively. Generating a fault propagation probability graph with weighted edges Through the calculation of the conditional transition probability and the adjustment of the dynamic time window, a fault propagation probability graph with weighted edges is finally generated. In this graph, nodes represent metrics at different levels, edges represent the correlation and fault propagation probability between them. The weight represents the fault propagation intensity between different metrics. Graph with weighted edges: The size of the weight reflects the possibility of a fault propagating from one level to another. The larger the weight, the stronger the fault propagation between the two metrics, and the higher the probability of a fault occurrence. Formula: ; where: is the weight of the edge from node to node . is the conditional transition probability from node to node . is a reliability factor, indicating the nodes and The connection quality between them may be evaluated based on factors such as system stability and device health status. Through the above steps, an accurate fault propagation probability map is constructed, which helps to implement more precise fault warning and repair strategies in cloud monitoring services.
[0032] Specifically, the identification logic for identifying hidden association nodes with spatio-temporal coupling characteristics among cross-layer metrics is as follows: Extract the propagation paths across the physical layer, virtual layer, and application layer in the fault propagation probability map, and screen out candidate nodes whose transition probability exceeds the dynamic threshold; Conduct mutual information entropy analysis on the candidate nodes to quantify their dependence strength on upstream and downstream metrics, and eliminate weak association interference terms; Verify the spatio-temporal causal relationship directionality of the candidate paths through Granger causality test, and inject directional perturbations into the high-probability paths to observe the response amplitude of downstream metrics to confirm the effectiveness of spatio-temporal coupling.
[0033] In this implementation plan, to screen out candidate nodes whose transition probability exceeds the dynamic threshold, in the fault propagation probability map, the propagation paths between the cross-layer physical layer, virtual layer, and application layer are regarded as potential association paths. First, it is necessary to screen out those candidate nodes whose transition probability exceeds a certain dynamic threshold. Transition probability: The transition probability represents the probability of state change from one node to another. When this probability exceeds the set dynamic threshold, it indicates that this path may involve a strong correlation. Therefore, these nodes are considered candidate nodes. Dynamic threshold: The setting of the threshold is dynamic and is adjusted based on historical data and system state factors. Adjust the threshold according to the fluctuations of different system loads or fault history to improve the sensitivity of identification. For the candidate nodes screened out, next, conduct mutual information entropy analysis to quantify their dependence strength on upstream and downstream metrics. Mutual information entropy can help us evaluate the dependence relationship between two variables and quantify the correlation between cross-layer metrics. Mutual information entropy: Mutual information entropy quantifies the common information between two variables and measures the degree of information sharing between them. The larger the mutual information entropy, the stronger the dependence relationship between the two variables. Formula: ; where: and are two variables between cross-layer metrics (metrics of the physical layer and the application layer). is and 's joint probability distribution. and are respectively and 's marginal probability distributions. represents and The mutual information between. Through the calculation of mutual information entropy, it is possible to identify which nodes have strong correlations, and then eliminate those interference items with weak dependencies and retain important correlation paths. Granger causality test After the mutual information entropy analysis, we use the Granger causality test on the candidate paths to verify the directionality of their spatio-temporal causal relationships. The Granger causality test can help us determine whether one variable can predict the future changes of another variable. Granger causality test: By analyzing historical data, if the historical data of a variable can significantly improve the predictive ability for the current value of variable , we can say that Granger causality affects . This can help us understand the causal relationships between different levels in the system. The formula is as follows: ; where: is the target variable (such as the performance metric at the application layer). is the explanatory variable (such as the resource consumption situation at the physical layer). is the constant term, and are the lag coefficients, indicating the influence of historical data on the current value. is the number of lags, indicating the time range of the variable's history. is the error term, representing the part that the model cannot explain. The Granger causality test can help us identify whether there is a temporal causal relationship between cross-layer metrics, that is, to determine whether the metrics at one level will affect the metrics at other levels. Directed perturbation and downstream metric response: Once we have confirmed the directionality of the causal relationships between the candidate paths, the next step is to inject directed perturbations into the high-probability paths and observe the response amplitude of the downstream metrics. Directed perturbation: Directed perturbation refers to artificially changing the states of certain nodes (for example, by simulating load changes or resource consumption) to test the system's response to these changes. The purpose of directed perturbation is to verify whether the causal relationship has practical influence in the system. Downstream metric response amplitude: We observe the changes that occur in the downstream metrics after the directed perturbation, specifically manifested as the response amplitude. The larger the response amplitude, the stronger the spatio-temporal coupling of this causal path. If the changes in the downstream metrics are significant, it indicates a strong effectiveness of spatio-temporal coupling. Finally, after the directed perturbation and response tests, we confirm the effectiveness of spatio-temporal coupling. If the perturbation has a significant impact on the downstream metrics, then we can confirm that these paths are effective spatio-temporal coupling paths. Spatio-temporal coupling effectiveness: The effectiveness of spatio-temporal coupling means that there are strong causal relationships between cross-layer metrics, and these relationships have a practical impact on the system performance. This relationship can be further used to optimize the system, detect potential faults, and perform resource allocation.
[0034] Specifically, the specific process of verifying the authenticity of cross-layer causality through directional perturbation injection and fitting the coupling degree index according to the correlation matrix of resource scheduling operations and the fault propagation chain is as follows: Select a high-probability association path in the fault propagation graph, and inject a controllable perturbation that simulates a sudden increase in physical layer storage delay or restricts the virtual layer container network bandwidth into the path source node; Observe the downstream indicator response and record the perturbation propagation path and amplitude, and compare the prediction path consistency with the original graph; Extract the temporal relationship between historical resource scheduling operations and the fault propagation chain, construct the correlation matrix of resource scheduling operations and the fault propagation chain, and based on the influence weight of resource scheduling operations on the fault chain length and repair time in the matrix, fit the coupling degree index through the gradient descent method.
[0035] In this implementation plan, first, in the constructed fault propagation graph, we need to select high-probability association paths. A high-probability path represents the path that is most likely to have a causal relationship across the physical layer, virtual layer, and application layer. After selecting these paths, we will inject simulated perturbations into the source nodes of the paths to verify the authenticity of the causal relationship. Directional perturbation injection: We inject controlled perturbations at the source nodes. Common perturbations include a sudden increase in storage latency at the physical layer or restricting the network bandwidth of virtual layer containers, etc. For example, the physical layer may inject a sudden increase in storage latency, and the virtual layer can restrict the network bandwidth of the container to simulate resource bottlenecks or faults. Through directional perturbations, we verify whether the impact of the source node on downstream metrics conforms to the predicted path we obtained from the fault propagation graph. If the response of the downstream metrics after perturbation conforms to the expectation, the causal relationship on this path can be verified. After injecting perturbations into the source nodes, we need to observe the responses of the downstream metrics and record the propagation path and response amplitude of the perturbations. Response observation: We observe the changes in other hierarchical metrics of the system (such as application layer performance, virtual layer resource usage, etc.) after perturbation injection. If the response amplitude of the downstream metrics is large, it indicates that the causal relationship on the path is effective, and the impact of the perturbation is transmitted along the propagation path. Comparing the predicted path: Compare the actually observed perturbation propagation path with the predicted path in the original graph. If the two are consistent, it indicates that this causal relationship path is real and effective. To further analyze the association between fault propagation and resource scheduling, we need to extract the temporal relationship between historical resource scheduling operations and fault propagation chains. The purpose of this step is to find the potential association between resource scheduling operations and fault propagation and understand how these operations affect the length and repair time of the fault propagation chain. Temporal relationship: By analyzing historical data, we construct the temporal relationship between resource scheduling operations and fault propagation chains. For example, whether resource scheduling operations advance or delay the occurrence of fault propagation, or whether certain scheduling operations can shorten the repair time, etc. Based on the extracted historical data, we construct an association matrix between resource scheduling operations and fault propagation chains. This matrix contains the influence weights of each resource scheduling operation on the length and repair time of the fault propagation chain. Association matrix: In the association matrix, each element represents the impact of a certain resource scheduling operation on the fault propagation chain. Specifically, if an operation significantly shortens the length or repair time of the fault propagation chain, the value of this element in the matrix is large. Based on the data in the association matrix, we use the gradient descent method to fit the coupling degree index to quantify the coupling degree between resource scheduling operations and fault propagation chains. Definition of the coupling degree index The coupling degree index ( ) reflects the relationship between a certain resource scheduling operation and the fault propagation chain, considering the impact of resource scheduling operations on aspects such as the length of the fault propagation chain, repair time, and propagation path. To fit the coupling degree index, multiple factors can be introduced and calculated in combination with weights. Coupling degree index formula: ; where: : represents the coupling degree index of the a-th resource scheduling operation and the fault propagation chain. : the total number of resource scheduling operations. : the weight of the a-th resource scheduling operation, indicating the influence proportion it occupies in the entire fault propagation chain. : the influence of the a-th resource scheduling operation on the length of the fault propagation chain : the influence of the a-th resource scheduling operation on the repair time of the fault propagation chain : the influence of the a-th resource scheduling operation on the propagation path of the fault propagation chain : represent the influence weight coefficients of the chain length, repair time, and propagation path respectively. Definitions of various influencing factors, influence of chain length : represents the influence of the resource scheduling operation on the length of the fault propagation chain, usually determined by the expansion or compression of nodes or chains caused by the scheduling operation. For example, optimized resource scheduling may reduce the chain length, thereby reducing the delay of the propagation chain. ; influence on repair time : represents the influence of the resource scheduling operation on the repair time of the fault propagation chain. Changes in the scheduling operation may affect the timeliness of the recovery process, and thus affect the repair time. ; influence on propagation path : represents the influence of the resource scheduling operation on the propagation path of the fault propagation chain. The scheduling operation may activate new paths or affect the propagation effect of existing paths. ; Fitting the coupling degree index with the gradient descent method To optimize the coupling degree index, we can use the gradient descent method for parameter tuning to ensure that the model can better adapt to different resource scheduling and fault propagation scenarios. Gradient descent method for optimizing the coupling degree index parameters and The formula is as follows: ; where: is the parameter set to be optimized. is the learning rate. is the loss function, which measures the error between the coupling degree index predicted by the model and the actual coupling degree. Definition of the loss function is used to measure the error in the fitting process, which takes into account the difference between the actually observed coupling degree and the calculated coupling degree. The definition is as follows: ; where: is the coupling degree index predicted according to the resource scheduling operation and the model. is the actually observed coupling degree index. By minimizing the loss function , we can find the optimal parameter combination, so that the coupling degree index between the resource scheduling operation and the fault propagation chain can be fitted more accurately. The dynamic update of the model is through the iterative update of the gradient descent method, and each optimization will adjust the parameters , , and , making the calculated coupling degree index more conform to the characteristics of the actual fault propagation chain. As the number of iterations increases, the coupling degree index will be gradually optimized to the best configuration. Coupling degree index: The coupling degree index represents the degree of influence of the resource scheduling operation on the fault propagation chain. Specifically, it describes to what extent the resource scheduling operation can affect the length and repair time of the fault propagation chain. A higher coupling degree index indicates a stronger association between the resource scheduling operation and the fault propagation chain.
[0036] Specifically, the specific process of constructing the coupling degree analysis rule to quantify the influence weight of the operation on the fault propagation and generating the root cause location instruction for the negative coupling scenario is as follows: Based on the association matrix between the resource scheduling operation and the fault propagation chain, calculate the contribution weight of the operation to the fault chain length and repair time, and generate the initial coupling degree analysis rule; Set the dynamic decision threshold, and when the contribution weight of the operation to the fault propagation chain exceeds the threshold, mark it as a negative coupling operation; Trace back the nodes associated with the negative coupling operation in the fault propagation graph, and screen the high causal strength nodes not covered by the historical scheduling operation as candidate root causes; Conduct a reverse blocking test on the candidate root cause, and verify its interruption effect on the downstream fault chain by restricting its resource access or traffic distribution; Generate the location instruction including the root cause node identifier, the influence path and the repair suggestion, and push it to the operation and maintenance terminal.
[0037] In this implementation plan, first, an association matrix is constructed based on the relationship between resource scheduling operations and the fault propagation chain. This matrix reflects the impact of each resource scheduling operation on the length of the fault propagation chain, the repair time, and the propagation path. Through this matrix, the contribution of each resource scheduling operation to the fault chain can be quantified. Generate coupling degree analysis rules: After the association matrix between resource scheduling operations and the fault propagation chain is constructed, the coupling degree analysis rules are generated next. The main objective of this rule is to quantify the influence weight of each resource scheduling operation on the fault propagation chain and identify the negative coupling of the operation. Contribution weight calculation: Calculate the weight of the relationship between each resource scheduling operation and the fault propagation chain. The weight reflects the degree of influence of this operation on the entire fault chain length, repair time, and propagation path. The higher the contribution weight of the resource scheduling operation, the more significant its impact on the fault chain. Dynamic decision threshold: To identify negative coupling operations, a dynamic decision threshold needs to be set. When the contribution weight of a certain resource scheduling operation to the fault propagation chain exceeds this threshold, it is considered a negative coupling operation. Negative coupling operations will exacerbate the propagation of faults, so special attention and handling are required. Generate initial rules: Based on the above analysis, the coupling degree analysis rules are initially formed, which define what situations are determined to be negative coupling operations, how to calculate the contribution weight, and how to screen the resource scheduling operations that need to be focused on. Identification of negative coupling operations: After the coupling degree analysis rules are generated, the nodes in the fault propagation graph are analyzed next to identify those nodes related to negative coupling operations. Trace back the nodes associated with negative coupling operations: Use the marked negative coupling operations to trace back and analyze the nodes they are associated with. These nodes are potential sources in the fault propagation graph and may be the most critical trigger points in the fault propagation chain. Screen nodes with high causal strength: Among the traced-back nodes, screen those nodes with relatively high causal strength. These nodes are the parts with the strongest association with negative coupling operations. The causal strength can be determined by measuring the influence of the node during the fault propagation process. Nodes not covered by historical scheduling operations: Special attention needs to be paid to those nodes not covered by historical scheduling operations because these nodes may be the key reasons for the fault outbreak. Root cause location and verification: Further verify the selected candidate root cause nodes to determine whether they actually have an impact on the fault propagation chain. Reverse blocking test: Perform a reverse blocking test on the candidate root cause nodes. By restricting their resource access or traffic distribution, observe whether the downstream fault chain is interrupted. If it is interrupted, it proves that this node is the root cause in the fault propagation chain. Verification method: During the reverse blocking test, restricting the access to resources or traffic will change the state of the fault propagation chain. If the fault propagation path is effectively interrupted in this way, it indicates that this node is a valid root cause. Generate root cause location instructions: After verifying the candidate root cause, generate location instructions containing the root cause node identifier, the influence path, and repair suggestions. These instructions will be pushed to the operation and maintenance terminal for timely fault recovery and repair.Root cause localization instruction content: Root cause node identifier: including the key node identifiers in the fault propagation chain. Impact path: Determine the fault propagation path where the root cause node is located, and analyze the impact of this path on the overall system. Repair suggestion: Provide repair suggestions according to the characteristics of the root cause node, such as resource optimization, traffic restriction, path adjustment, etc. Push to the operation and maintenance terminal: Finally, the root cause localization instruction will be pushed to the operation and maintenance terminal through the automation system for operation and maintenance personnel to refer to and perform subsequent operations. The operation and maintenance personnel take corresponding repair measures according to the localization instruction to ensure the stable operation of the system.
[0038] Specifically, the specific process of analyzing the fitting propagation cost gradient of the historical fault repair time and the resource waste rate according to the fault propagation cost model is as follows: Extract the historical fault repair time data and the resource waste rate caused by resource scheduling operations, and construct an initial propagation cost function; Iteratively optimize the cost function parameters by the gradient descent method, and dynamically adjust the repair time weight and the resource waste penalty factor; Calculate the current propagation cost gradient according to the real-time fault propagation path length and the change of resource utilization rate; Continuously optimize the gradient parameters through the policy execution feedback data to adapt to the dynamic changes of the hybrid cloud environment.
[0039] In this implementation plan, construct a propagation cost model: First, by extracting historical data, construct an initial propagation cost function, which mainly includes two key factors: historical fault repair time and resource waste rate. These factors will serve as the basis for the propagation cost to help quantify the time cost and resource consumption in the fault repair process. Repair time and resource waste rate: Repair time: It represents the time consumed from the occurrence of the fault to the complete repair, which is usually affected by factors such as system response, fault location, and repair means. Resource waste rate: It represents the ineffective utilization or waste of system resources during the occurrence of the fault, which may be caused by redundant calculations, idle resources, over-allocation of resources, etc. The expression of the initial propagation cost function can be as follows: ; where: : Initial propagation cost. : Historical fault repair time. : Resource waste rate. and : Weight coefficients, which respectively control the contributions of repair time and resource waste to the propagation cost. Use the gradient descent method to optimize the propagation cost function: After constructing the initial propagation cost function, use the gradient descent method to optimize the parameters of the propagation cost function, especially the adjustment of the repair time weight and the resource waste penalty factor. This step is to continuously calculate the gradient iteratively and optimize the weights to reduce the total propagation cost. The optimization process aims to make the propagation cost function better reflect the actual fault repair process, dynamically adjust the weights of repair time and resource waste, so as to achieve more efficient resource utilization and shorter repair time. Implementation of the gradient descent method: The gradient descent method updates the parameters according to the gradient of the current propagation cost function, and optimizes the parameters and The update rules are as follows: ; ; where: and are the repair time weight and resource waste penalty factor for the f-th iteration respectively. η: learning rate, which controls the magnitude of each update. and are the partial derivatives of the propagation cost function with respect to the weight parameters respectively. Calculate the current propagation cost gradient: Once the optimization process is completed and the parameters of the propagation cost function have been adjusted to appropriate values, real-time data can be used to calculate the current propagation cost gradient. At this time, the influence weights of repair time and resource waste have been dynamically adjusted, and the propagation cost function can also accurately reflect the actual situation. Real-time propagation path and resource utilization rate: Propagation path length: It represents the span of fault propagation and can be regarded as the number of links from the fault source to the end point. The longer the fault propagation path, the longer the repair time usually required and the greater the resource consumption. Resource utilization rate: It represents the actual utilization degree of system resources during the fault repair process. High resource utilization rate can reduce resource waste and thus reduce the propagation cost. The calculation formula of the propagation cost gradient can be expressed as: ; where: : The gradient of the current propagation cost. : The change in real-time fault repair time. : The change in real-time resource waste rate. Policy execution feedback and gradient optimization: In order to adapt to the dynamic changes in the hybrid cloud environment, it is necessary to continuously optimize the gradient parameters through the feedback data of policy execution. The feedback data can reflect the improvement or deterioration of the system performance after the policy adjustment, and thus affect the change in the propagation cost. Feedback data: Policy execution effect: Obtain the feedback of the system's dynamic changes by monitoring the fault propagation path length and resource utilization rate of the system after the policy execution. Gradient adjustment: Adjust the gradient parameters in the propagation cost function according to the feedback data, so as to more precisely optimize the repair time and resource waste in future fault repair processes. Policy adjustment: Adjust the parameters in the gradient descent process according to the real-time feedback, so that the model adapts to the changes in the hybrid cloud environment, ensuring the efficient utilization of system resources and rapid fault repair. Summary of the optimization process: By optimizing the propagation cost model through the gradient descent method, the system can dynamically adjust the influence weights of repair time and resource waste, so as to achieve the goal of real-time adaptation to the changes in the hybrid cloud environment. During the continuous optimization process, the model can not only accurately predict the cost of fault repair, but also automatically adjust with the changes in the environment, improving the efficiency and response ability of the system.
[0040] Specifically, the process of inputting the coupling degree index, real-time service level agreement default rate, and fault propagation cost gradient into the weight function, calculating the collaborative priorities of short-term suppression actions and long-term eradication actions through the dynamic weight function, and generating the execution sequences of the two types of actions through the asymmetric game strategy is as follows: Introduce a coupling degree penalty factor into the dynamic weight function to suppress the priority of scheduling operations that may exacerbate fault propagation in high-coupling scenarios; calculate the urgency weight of short-term suppression actions based on the real-time service level agreement default rate, and calculate the benefit weight of long-term eradication actions in combination with the propagation cost gradient; define the short-term suppression actions and long-term eradication actions as asymmetric game participants, construct a benefit function that quantifies their contributions to service availability improvement and fault propagation suppression; solve the optimal collaborative strategy through dynamic Nash equilibrium, give priority to executing short-term suppression actions to quickly stop losses, and asynchronously trigger long-term eradication actions.
[0041] In this implementation plan, introduce a coupling degree penalty factor and adjust the priority: When calculating the collaborative priorities of short-term suppression actions and long-term eradication actions, the coupling degree index needs to be considered first, which helps to identify operations that may exacerbate fault propagation. Scheduling operations in high-coupling scenarios may have a negative impact on system stability and should therefore be given a lower priority. Coupling degree penalty factor: Introduce a coupling degree penalty factor γcoupling to penalize high-coupling scenarios. The coupling degree index ζ reflects the tight relationship between the indicators at each level in fault propagation. A higher coupling degree usually indicates a higher risk of fault propagation. The formula is expressed as: ; where: : The priority weight of the scheduling operation, considering the influence of the coupling degree penalty factor. : The basic weight coefficient, used to adjust the influence degree of the priority. : The coupling degree penalty factor, indicating the penalty intensity for high-coupling scenarios, with a value range between [0, 1]. : The coupling degree index, reflecting the correlation strength between layers. Through the above formula, scheduling operations that may exacerbate fault propagation can be effectively suppressed. Calculate the urgency weight of short-term suppression actions: Short-term suppression actions refer to measures that respond quickly and slow down fault propagation, usually used for intervention when the service default rate is high. Therefore, the urgency weight of short-term suppression actions should be calculated based on the real-time service level agreement (SLA) default rate. SLA default rate: The default rate reflects the probability that the current system fails to meet service commitments on time. The higher the default rate, the stronger the urgency of short-term suppression actions. The formula is expressed as: ; where: : The urgency weight of short-term suppression actions. : The basic weight coefficient, used to adjust the urgency weight. : The real-time SLA default rate. : The impact index of the default rate on the urgency weight, usually , such that when the default rate increases, the urgency weight increases significantly. Through this formula, the execution urgency of short-term suppression actions can be dynamically adjusted according to the SLA default rate. Calculate the benefit weight of long-term eradication actions: Long-term eradication actions aim to reduce the probability of future failures and improve the overall stability of the system by eliminating the root causes of failures. Therefore, the benefit weight of long-term eradication actions should be associated with the current propagation cost gradient. Propagation cost gradient: The propagation cost gradient reflects the cost changes during the current failure propagation process. The higher the propagation cost, the greater the benefit of long-term eradication actions. Formula representation: ; where: : The benefit weight of long-term eradication actions. : The basic weight coefficient, used to adjust the benefit weight of long-term eradication actions. : The current propagation cost gradient, reflecting the economic cost of failure propagation. : The impact index of propagation cost on the benefit weight, usually , such that the higher the cost, the greater the priority and benefit of the eradication action. Construct a game model and calculate the collaborative priority: In this step, short-term suppression actions and long-term eradication actions are regarded as two players in the game model. These two types of actions have different goals and benefits respectively. Short-term suppression actions focus on quickly reducing the impact of failures, while long-term eradication actions focus on eliminating the root causes of failures. Asymmetric game model: The contributions and benefits of short-term suppression actions and long-term eradication actions are different, so an asymmetric game strategy is needed to calculate the collaborative priority. Through the game model, the contributions of both to failure propagation suppression and service availability can be quantified, and finally the optimal execution order can be determined. Formula representation: ; where: and represent the execution strategies of short-term suppression actions and long-term eradication actions respectively. and represent the benefit functions of short-term suppression actions and long-term eradication actions respectively, considering the urgency weight, benefit weight, and priority. Solve the dynamic Nash equilibrium to obtain the optimal collaborative strategy: By solving the Nash equilibrium of the game model, the optimal collaborative strategy can be obtained, that is, under the given conditions, how to select short-term suppression actions and long-term eradication actions to maximize the improvement of system service availability and the suppression effect of failure propagation. Execution order: Based on the Nash equilibrium solution of the game, the system will first execute short-term suppression actions to quickly stop losses; then, asynchronously trigger long-term eradication actions to ensure the fundamental solution of failures. Through this method, the operation strategy of the system can be dynamically adjusted according to the actual failure propagation situation, resource utilization situation, and service default rate to achieve efficient failure suppression and service guarantee.
[0042] Specifically, the policy execution and feedback module specifically includes: The atomic execution engine disassembles the decision-making actions into independently executable atomic operations that call the cloud platform API to trigger the traffic limiting policy or initiate the storage volume migration task, and ensures the atomicity and consistency of cross-platform operations through the transaction lock mechanism; monitors the cross-layer metric changes after execution, captures the inhibitory effect of the actions on the fault propagation chain and the impact on resource utilization, adjusts the weight parameters in the fault propagation cost model according to the feedback data, optimizes the coupling degree analysis rules, and enhances the early recognition ability for negative coupling scenarios.
[0043] In this implementation plan, the module consists of an atomized execution engine, a transaction control mechanism, a cross-layer metric monitoring unit, and a feedback parameter tuning mechanism, aiming to achieve reliable execution and dynamic optimization of the scheduling strategy. Atomized Execution Engine: The atomized execution engine is used to refine unstructured high-level policy instructions into atomic operations at the cloud resource operation level. Each policy action is decomposed into several independently executable atomic tasks that directly control the cloud platform resource management interface (API). Examples of atomic operations: Invoke the rate-limiting policy API interface to trigger service instance degradation. Initiate the migration task of storage volumes (volumes) between different nodes to alleviate I / O bottlenecks or achieve load balancing. Atomicity and Consistency Guarantee Mechanism: To avoid the problem of inconsistent intermediate states during cross-platform or cross-service execution, a distributed transaction lock mechanism is introduced to ensure the following two points: Atomicity of operations: Any composite operation either succeeds completely or rolls back completely. Consistency of operations: After the operation is completed, the resource states of each subsystem are in a consistent and controllable state. Monitoring content includes: The change trend of key cross-layer metrics (such as system load, service latency, link traffic, instance failure rate, etc.) before and after policy execution; The state response of the nodes in the fault propagation chain to quantify the truncation or mitigation effect of the scheduling action on the potential propagation path; The change in resource utilization rate to evaluate the allocation efficiency of the action execution for computing power, storage, or bandwidth resources. Capture method: Collect metrics at each layer through an integrated monitoring system; Use a propagation chain analyzer to mark the metric response delay and slope change between the "controlled nodes" and the "potential propagation path". Model optimization goal: According to the feedback of the monitoring results, dynamically adjust the fault propagation cost model and coupling degree analysis model used in the decision-making process to improve the pertinence and foresight of policy responses. Tuning content includes: Adjustment of propagation cost weights: If a certain type of action is significantly effective in blocking the propagation chain in multiple execution rounds, its cost performance weight in the cost model can be increased (for example, increase the propagation reduction benefit value corresponding to its unit execution cost). Revision of coupling degree analysis rules: Analyze the co-variation relationship of cross-layer metrics before and after the action, and identify "negative coupling scenarios" that are not covered by the initial rules (that is, an operation at one layer causes reverse interference to another layer) to expand the rule base and improve the sensitivity of coupling discrimination. Enhancement of early warning capabilities: Feed back the key metric change patterns identified in the feedback to the training set to update the early identification model (such as the attention mechanism feature extraction network for negative coupling identification).
[0044] Please refer to Figure 2, A dynamic optimization method for cloud monitoring service operation and maintenance based on AI agents, comprising the following steps: S1. Deploy a containerized probe cluster to collect in real time the fragmentation metrics of physical layer hardware resources, container lifecycle events in the virtual layer, and performance deviation data of microservice call chains in the application layer in a hybrid cloud environment. Through hierarchical tagging preprocessing and outlier filtering, output a standardized cross-layer metric set; S2. Connect the cross-layer metric perception module, eliminate the time series deviation between the low-frequency data in the physical layer and the high-frequency data in the application layer through an incremental time series alignment algorithm, construct a fault propagation probability map based on a dynamic time window, and identify hidden association nodes with spatio-temporal coupling characteristics among cross-layer metrics; S3. Receive the fault propagation map output by the spatio-temporal association analysis module, verify the authenticity of cross-layer causal relationships through directed perturbation injection, fit the coupling degree index according to the association matrix between resource scheduling operations and the fault propagation chain, and construct a coupling degree analysis rule to quantify the influence weight of operations on fault propagation, generating root cause location instructions for negative coupling scenarios; S4. Analyze the historical fault repair time and resource waste rate to fit the propagation cost gradient according to the fault propagation cost model, input the coupling degree index, the real-time service level agreement default rate, and the fault propagation cost gradient into the weight function, calculate the collaborative priority of short-term suppression actions and long-term eradication actions through a dynamic weight function, and generate an execution sequence for the two types of actions through an asymmetric game strategy; S5. Invoke the cloud platform interface to atomically execute decision-making actions, and collect the changed metric data after execution to dynamically update the fault propagation cost model and the coupling degree analysis rule.
[0045] In this implementation scheme: S1. In this step, low-intrusive acquisition of multi-source heterogeneous data is achieved by deploying a "containerized probe cluster", and a "cross-layer metric structure system" consisting of a physical layer, a virtual layer, and an application layer is constructed. Acquisition content of the physical layer: fragmentation of server CPU usage rate, dispersion of memory allocation, fluctuations in hard disk I / O, etc.; acquisition content of the virtual layer: container restart, migration, life cycle events; acquisition content of the application layer: microservice link call latency, failure rate, throughput mutation, etc.; the "tagging preprocessing mechanism" is used to weight and classify metrics at different levels, and an outlier filter is integrated to improve the quality of training samples; real-time extraction and normalization processing of cross-layer fault symptoms are realized, and a high-dimensional and low-coupling data feature space is constructed to provide structured input for subsequent propagation analysis and causal modeling. S2. An incremental time series alignment algorithm and a dynamic time window graph construction strategy are proposed to achieve unified modeling of cross-layer data streams under asynchronous sampling, and capture potential propagation paths and spatio-temporal coupling relationships. The synchronization of data between the physical layer (such as 5-minute sampling) and the application layer (such as 5-second sampling) is adjusted through the incremental sliding window algorithm; a propagation probability graph is constructed using the metric change rate and delay response relationship; the "latent association node discovery mechanism" is introduced to mine latent nodes that do not show obvious anomalies in the initial stage of propagation; the "cross-layer causal diffusion chain" that cannot be described by traditional single-layer logs is revealed, significantly enhancing the system's interpretability of mixed fault paths in complex scenarios. S3. In this step, scheduling operation behaviors are incorporated into causal chain modeling, and the "scheduling-propagation chain coupling index" and the "negative coupling behavior recognition mechanism" are proposed to construct a root cause location rule set. Calculation of the coupling degree index: An association matrix between scheduling operations and propagation chain changes is established to measure the gain factor of scheduling operations on the chain extension length and repair duration; determination of negative coupling operations: A dynamic contribution threshold is introduced, and when the operation behavior significantly amplifies the propagation chain or causes recovery delay, it is marked as a negative coupling source; root cause screening: Trace back to the negative operation association nodes, exclude historical scheduling overlapping interference nodes, and only retain nodes with high causal strength; reverse blocking verification: Perform resource blocking tests on candidate root cause nodes to verify their interruption effect on downstream propagation; effectively avoid the over-reliance of traditional root cause location on the degree of anomaly, and realize proactive, multi-path, and multi-source root cause reasoning from the perspectives of causal structure and system behavior. S4. In this step, for the first time, the "coupling degree index, SLA default rate, and propagation cost gradient" are incorporated into a unified dynamic weight function, and the "asymmetric game mechanism" is introduced to achieve coordinated sorting of two types of actions.Design of dynamic weight function: Introduce a coupling penalty factor to prevent negative feedback in high-coupling areas during scheduling operations; Urgency weight: Measure the priority of short-term suppression actions through the degree of SLA violation; Long-term benefit weight: Calculate the benefits of eradication actions through the propagation cost gradient function; Asymmetric game modeling: Regard the two types of actions as game participants; Construct a two-objective benefit function with service availability and propagation reduction ability; Through the dynamic Nash equilibrium solution, obtain an optimal action combination strategy that can evolve with the environment; Break through the problem of rigidification in traditional rule-driven operation and maintenance, achieve coordinated control of short-term rapid loss prevention and long-term risk elimination, and have extremely high operation and maintenance flexibility. S5. In this step, introduce the "cloud platform API-level atomic execution mechanism" and the "execution-monitoring-optimization feedback loop" to construct a decision-making closed loop. Atomic operation decomposition: Decompose policy actions into API-level tasks (such as bandwidth throttling, computing resource rescheduling, volume migration, etc.); Adopt a distributed transaction lock mechanism to ensure execution consistency and uninterruptibility in multi-cloud heterogeneous systems; Feedback collection and model optimization: Real-time monitor the changes in key indicators after policy execution; Use the feedback results to adjust the coupling degree weight factor and cost gradient function parameters; Update the negative coupling determination rule and the early identification strategy model; Achieve a complete closed loop of policy → execution → perception → feedback → re-optimization, support dynamic adaptation and rapid response in an unstable resource environment.
[0046] In summary, the present application has at least the following effects:
[0047] A cloud monitoring service operation and maintenance dynamic optimization system and method based on AI agents realizes real-time collection and labeled preprocessing of multi-dimensional data at the physical layer, virtual layer, and application layer through a containerized probe cluster, forms a standardized cross-layer index set, and provides a unified data basis for subsequent analysis. Through the incremental time series alignment and dynamic propagation graph construction algorithm, accurately capture the cross-layer spatio-temporal coupling relationship, reveal hidden fault links, and improve the accuracy and foresight of operation and maintenance decisions. By fitting the coupling degree index and constructing an operation-propagation chain correlation matrix, realize the identification and root cause location of negative scheduling behaviors, and avoid the system instability risk brought by traditional experience-driven scheduling. Multidimensionally integrate the service level agreement violation rate, coupling degree index, and fault propagation cost gradient, and output a coordinated optimization action sequence through an asymmetric game strategy to improve the overall benefit ratio and response efficiency of policy execution. Through the cloud platform API-level atomic execution engine and feedback tuning mechanism, realize the full-process closed-loop control from policy generation to execution monitoring, have good system evolution ability and dynamic adaptability, and significantly enhance the intelligence and resilience of the cloud monitoring system.
[0048] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0049] The present invention is described with reference to the flowcharts and / or block diagrams of systems, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0050] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0051] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0052] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present invention.
[0053] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention also intends to include these modifications and variations.
Claims
1. A cloud monitoring service operation and maintenance dynamic optimization system based on an AI agent, characterized in that, It includes the following steps: Cross-layer metric perception module, spatio-temporal correlation analysis module, causal verification and coupling degree analysis module, dynamic priority decision-making module, policy execution and feedback module; The cross-layer metric perception module is used to deploy a containerized probe cluster to collect in real time the fragmentation metrics of physical layer hardware resources, container lifecycle events in the virtual layer, and performance deviation data of microservice call chains in the application layer. Through hierarchical tagging preprocessing and outlier filtering, it outputs a standardized cross-layer metric set; The spatio-temporal correlation analysis module is used to connect to the cross-layer metric perception module, eliminate the temporal deviation between the low-frequency data of the physical layer and the high-frequency data of the application layer through an incremental time series alignment algorithm, construct a fault propagation probability map based on a dynamic time window, and identify latent correlation nodes with spatio-temporal coupling characteristics among cross-layer metrics; The causal verification and coupling degree analysis module is used to receive the fault propagation map output by the spatio-temporal correlation analysis module, verify the authenticity of cross-layer causal relationships through directed perturbation injection, fit the coupling degree index according to the correlation matrix between resource scheduling operations and the fault propagation chain, and construct coupling degree analysis rules to quantify the influence weight of operations on fault propagation, generating root cause location instructions for negative coupling scenarios; The dynamic priority decision-making module is used to analyze the historical fault repair time and resource waste rate to fit the propagation cost gradient according to the fault propagation cost model, input the coupling degree index, real-time service level agreement default rate, and fault propagation cost gradient into the weight function, calculate the collaborative priority of short-term suppression actions and long-term eradication actions through the dynamic weight function, and generate the execution sequence of the two types of actions through an asymmetric game strategy; The policy execution and feedback module is used to call the cloud platform interface to atomically execute decision-making actions, collect the changed metric data after execution, and dynamically update the fault propagation cost model and coupling degree analysis rules.
2. The dynamic optimization system for cloud monitoring service operation and maintenance based on an AI agent according to claim 1, wherein: The cross-layer metric perception module specifically includes: The containerized probe cluster is deployed in a microservice architecture on hybrid cloud nodes, including physical layer probes to collect hardware resource fragmentation metrics, virtual layer probes to monitor container lifecycle events and resource contention behaviors, and application layer probes to track microservice call chain topologies and performance deviations; Adjust the collection frequency according to the dynamic change rate of metrics. The physical layer uses low-frequency trigger sampling, and the application layer uses event-driven high-frequency tracking, and suppress instantaneous noise through sliding window statistics; The hierarchical tagging preprocessing unit attaches environment context tags of cloud platform types and service dependency levels to the original data, filters outliers based on the isolation forest algorithm, and outputs a standardized cross-layer metric set.
3. The dynamic optimization system for cloud monitoring service operation and maintenance based on an AI agent according to claim 2, characterized in that: The specific process of eliminating the temporal deviation between the low-frequency data of the physical layer and the high-frequency data of the application layer through the incremental time series alignment algorithm and constructing a fault propagation probability map based on a dynamic time window is as follows: Perform dynamic interpolation on the low-frequency data of the physical layer and the high-frequency data of the application layer, and generate a continuous time series by weighting according to data confidence; Detect the potential phase difference between cross-layer metrics through sliding correlation analysis, and dynamically adjust the interpolation anchor points to eliminate temporal offset; Calculate the conditional transition probability between cross-layer metrics based on the aligned time-series data, dynamically expand the time window range, and shrink the window when detecting metric mutations to generate a fault propagation probability graph with weighted edges.
4. The dynamic optimization system for cloud monitoring service operation and maintenance based on an AI agent according to claim 3, wherein: The recognition logic for identifying hidden association nodes with spatio-temporal coupling characteristics between cross-layer metrics is as follows: Extract the propagation paths across the physical layer, virtual layer, and application layer in the fault propagation probability graph, and filter out candidate nodes with transition probabilities exceeding the dynamic threshold; Conduct mutual information entropy analysis on the candidate nodes to quantify their dependence strength with upstream and downstream metrics, and eliminate weak association interference terms; Verify the spatio-temporal causality directionality of the candidate paths through Granger causality test, and inject directional perturbations into the high-probability paths to observe the response amplitude of downstream metrics to confirm the effectiveness of spatio-temporal coupling.
5. An AI intelligent agent-based cloud monitoring service operation and maintenance dynamic optimization system according to claim 4, characterized in that: Verify the authenticity of cross-layer causal relationships through directional perturbation injection, and the specific process of fitting the coupling degree index according to the association matrix between resource scheduling operations and the fault propagation chain is as follows: Select high-probability association paths in the fault propagation graph, and inject controllable perturbations such as simulating a sudden increase in physical layer storage delay or restricting the virtual layer container network bandwidth into the source nodes of the paths; Observe the response of downstream metrics and record the perturbation propagation path and amplitude, and compare with the predicted path consistency of the original graph; Extract the time-series relationship between historical resource scheduling operations and the fault propagation chain, construct the association matrix between resource scheduling operations and the fault propagation chain, and fit the coupling degree index through the gradient descent method based on the influence weights of resource scheduling operations on the fault chain length and repair time in the matrix.
6. The dynamic optimization system for cloud monitoring service operation and maintenance based on an AI agent according to claim 5, characterized in that: And construct coupling degree analysis rules to quantify the influence weight of operations on fault propagation, and the specific process of generating root cause location instructions for negative coupling scenarios is as follows: Based on the association matrix between resource scheduling operations and the fault propagation chain, calculate the contribution weights of operations to the fault chain length and repair time, and generate initial coupling degree analysis rules; Set a dynamic decision threshold, and when the contribution weight of an operation to the fault propagation chain exceeds the threshold, mark it as a negative coupling operation; Trace back the nodes associated with negative coupling operations in the fault propagation graph, and filter out high-causal-strength nodes not covered by historical scheduling operations as candidate root causes; Conduct a reverse blocking test on the candidate root causes, and verify their interruption effect on the downstream fault chain by restricting their resource access or traffic distribution; Generate location instructions containing root cause node identifiers, influence paths, and repair suggestions, and push them to the operation and maintenance terminal.
7. An operation and maintenance dynamic optimization system for cloud monitoring services based on an AI agent according to claim 6, characterized in that: The specific process of analyzing the historical fault repair time and resource waste rate to fit the propagation cost gradient according to the fault propagation cost model is as follows: Extract the historical fault repair time data and the resource waste rate caused by resource scheduling operations, and construct an initial propagation cost function; Iteratively optimize the cost function parameters through the gradient descent method, and dynamically adjust the repair time weight and resource waste penalty factor; Calculate the current propagation cost gradient according to the real-time fault propagation path length and resource utilization rate changes; Continuously optimize the gradient parameters through policy execution feedback data to adapt to the dynamic changes of the hybrid cloud environment.
8. The dynamic optimization system for cloud monitoring service operation and maintenance based on an AI agent according to claim 7, characterized in that: The specific process of inputting the coupling degree index, real-time service level agreement default rate, and fault propagation cost gradient into the weight function, calculating the collaborative priorities of short-term suppression actions and long-term eradication actions through the dynamic weight function, and generating the execution sequences of the two types of actions through the asymmetric game strategy is as follows: By introducing a coupling degree penalty factor into the dynamic weight function, the priority of scheduling operations positively correlated with the fault propagation chain is suppressed; Based on the real-time service level agreement default rate, calculate the urgency weight of short-term suppression actions, and combine the propagation cost gradient to calculate the benefit weight of long-term eradication actions; Define short-term suppression actions and long-term eradication actions as asymmetric game participants, and construct a benefit function that quantifies their contributions to service availability improvement and fault propagation suppression; Solve the optimal collaborative strategy through the dynamic Nash equilibrium, prioritize the execution of short-term suppression actions to quickly stop losses, and asynchronously trigger long-term eradication actions.
9. The dynamic optimization system for cloud monitoring service operation and maintenance based on an AI agent according to claim 8, characterized in that: The strategy execution and feedback module specifically includes: The atomic execution engine disassembles the decision-making actions into independently executable atomic operations that call cloud platform APIs to trigger traffic limiting strategies or initiate storage volume migration tasks, and ensures the atomicity and consistency of cross-platform operations through the transaction lock mechanism; Monitor the changes in cross-layer metrics after execution, capture the suppression effect of actions on the fault propagation chain and the impact on resource utilization, adjust the weight parameters in the fault propagation cost model according to the feedback data, optimize the coupling degree analysis rules, and enhance the ability to early identify negative coupling scenarios.
10. A dynamic optimization method for cloud monitoring service operation and maintenance based on an AI agent, which applies the dynamic optimization system for cloud monitoring service operation and maintenance based on an AI agent described in any one of claims 1-9, characterized in that, It includes the following steps: S1. Deploy a containerized probe cluster to collect in real-time the fragmentation metrics of physical layer hardware resources, container lifecycle events in the virtual layer, and performance deviation data of microservice call chains in the application layer in the hybrid cloud environment. Through hierarchical tagging preprocessing and outlier filtering, output a standardized set of cross-layer metrics; S2. Connect to the cross-layer metric perception module, eliminate the time series deviation between the low-frequency data in the physical layer and the high-frequency data in the application layer through the incremental time series alignment algorithm, construct a fault propagation probability map based on the dynamic time window, and identify the hidden associated nodes with spatio-temporal coupling characteristics among cross-layer metrics; S3. Receive the fault propagation map output by the spatio-temporal association analysis module, verify the authenticity of cross-layer causal relationships through directional perturbation injection, fit the coupling degree index according to the association matrix between resource scheduling operations and the fault propagation chain, and construct coupling degree analysis rules to quantify the impact weight of operations on fault propagation, and generate root cause location instructions for negative coupling scenarios; S4. Analyze the historical fault repair time and resource waste rate according to the fault propagation cost model to fit the propagation cost gradient. Input the coupling degree index, real-time service level agreement default rate, and fault propagation cost gradient into the weight function, calculate the collaborative priorities of short-term suppression actions and long-term eradication actions through the dynamic weight function, and generate the execution sequences of the two types of actions through the asymmetric game strategy; S5. Call the cloud platform interface to atomically execute the decision-making actions, and collect the changed metric data after execution to dynamically update the fault propagation cost model and the coupling degree analysis rules.
Citation Information
Patent Citations
Microservice intelligent operation and maintenance system and method oriented to cloud native and application
CN117009119A
Service operation and maintenance method and system based on cloud service architecture
CN119383115A
Intelligent scheduling and real-time cooperative control method for HarmonyOS industrial equipment
CN119916752A
Operation and maintenance alarm processing method and system based on knowledge graph enhanced large model
CN119988154A
Platform for facilitating development of intelligence in an industrial internet of things system
EP3966695A1
Cited By
Fault route rapid positioning system based on AI
CN120474901A
Intelligent traffic jam real-time optimization method based on artificial intelligence
CN120510713A
Hyper-converged server multi-resource integration system and scheduling method
CN120561343A
Energy storage power supply operation supervision system based on artificial intelligence
CN120598218A
Intelligent operation and maintenance monitoring method and system for data center
CN120602308A