A cloud monitoring service operation and maintenance dynamic optimization system and method based on AI intelligent agent

By deploying containerized probe clusters in a hybrid cloud environment, collecting and processing heterogeneous data in real time, building a fault propagation map and optimizing action priorities, the problems of cross-layer indicator fragmentation and hidden fault identification in cloud monitoring are solved, and the stability and resource utilization of cloud services are improved.

CN120223501BActive Publication Date: 2025-08-08SICHUAN ZHIXING ZHICHENG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510698823.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-08-08
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

In the hybrid cloud environment, the existing cloud monitoring solutions have cross-layer indicator fragmentation and insufficient ability to identify implicit fault associations. The resource scheduling strategies and fault repair actions lack coordination, making it difficult to meet elastic needs, and the operation and maintenance strategies are poorly adaptable to actual scenarios.

Method used

By deploying containerized probe clusters to collect heterogeneous data from hybrid cloud environments in real time, standardized cross-layer indicators are built using hierarchical labeling and dynamic timing alignment technology, fault propagation maps are built with incremental correlation analysis and dynamic time windows, implicit correlation nodes are identified, and causal relationships are verified through directional perturbations, and coordinated execution strategies are generated by combining fault propagation cost model and asymmetric game strategy to optimize action priorities.

Benefits of technology

Cross-layer data fusion and precise fault location are realized, the stability and resource utilization of cloud services are improved, the cost of manual investigation is reduced, and the changes in the hybrid cloud environment are dynamically adapted to changes in the hybrid cloud environment, avoiding the conflict between resource scheduling and fault repair.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120223501B_ABST
    Figure CN120223501B_ABST
Patent Text Reader

Abstract

The present invention discloses a cloud monitoring service operation and maintenance dynamic optimization system and method based on AI intelligent body, which relates to the field of cloud computing intelligent operation and maintenance technology. It is used to solve the problems of cross-layer data fragmentation, hidden fault correlation failure and operation and maintenance action conflict in hybrid clouds. Heterogeneous data of the physical layer, virtual layer and application layer are collected through containerized probes, and standardized cross-layer indicators are constructed through layered labeling and dynamic time series alignment. Fault propagation maps are constructed through incremental correlation analysis and dynamic time windows to identify spatiotemporal coupling nodes. Causal relationships are verified based on directional disturbances, and coupling indexes are fitted in combination with resource scheduling and fault chain correlation matrices to generate root cause location instructions. Action priorities are optimized based on propagation cost gradients and asymmetric game strategies, and models and rules are updated through feedback closed loops. The present invention realizes cross-layer data fusion and precise fault location, improving cloud service stability and resource utilization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of cloud computing intelligent operation and maintenance technology, and specifically to a cloud monitoring service operation and maintenance dynamic optimization system and method based on AI intelligent body. Background Art

[0002] Against the backdrop of the rapid development of information technology, cloud computing has become a core infrastructure for enterprise digital transformation. Especially with the growing popularity of hybrid and multi-cloud environments, the types of services hosted by cloud platforms are becoming increasingly complex, service dependency chains are significantly lengthening, and system operations are becoming highly dynamic, highly concurrent, and multi-layered. To ensure the continuity and stability of critical business systems, cloud platform operations and maintenance management are gradually evolving from static monitoring to intelligent, automated, and dynamic optimization. There is an urgent need to proactively identify and respond to potential failure risks in complex systems through AI agents and other means, thereby improving overall service quality assurance.

[0003] However, existing cloud monitoring solutions often rely on static rule bases or single data source drivers, resulting in problems such as fragmented cross-layer metrics and insufficient ability to identify hidden fault correlations. For example, the causal link between physical layer resource fragmentation and application layer service performance deviations is difficult to effectively verify, resulting in a high false alarm rate and delayed root cause location. Resource scheduling strategies and fault repair actions lack coordination, making it easy to exacerbate the risk of cascading failures due to blind capacity expansion. Furthermore, the lack of a dynamic trade-off mechanism between short-term emergency response and long-term optimization goals results in poor adaptability of O&M strategies to actual scenarios, making it difficult to meet the elasticity requirements of hybrid cloud environments. Summary of the Invention

[0004] In response to the shortcomings of the existing technology, the present invention provides a cloud monitoring service operation and maintenance dynamic optimization system and method based on AI intelligent body, which solves the problems of the above-mentioned background technology.

[0005] To achieve the above objectives, the present invention is implemented through the following technical solutions: a cloud monitoring service operation and maintenance dynamic optimization system based on AI intelligent body, comprising the following steps: a cross-layer indicator perception module, a spatiotemporal correlation analysis module, a causal verification and coupling analysis module, a dynamic priority decision module, and a strategy execution and feedback module; the cross-layer indicator perception module is used to deploy a containerized probe cluster, collect real-time physical layer hardware resource fragmentation indicators, virtual layer container life cycle events and application layer microservice call chain performance offset data in a hybrid cloud environment, and output a standardized cross-layer indicator set through layered labeling preprocessing and outlier filtering; the spatiotemporal correlation analysis module is used to connect the cross-layer indicator perception module, eliminate the timing deviation of low-frequency data of the physical layer and high-frequency data of the application layer through an incremental timing alignment algorithm, construct a fault propagation map based on a dynamic time window, and identify implicit correlation nodes with spatiotemporal coupling characteristics between cross-layer indicators; the causal verification and coupling degree The analysis module is used to receive the fault propagation map output by the spatiotemporal correlation analysis module, verify the authenticity of the cross-layer causal relationship through directional disturbance injection, fit the coupling index according to the correlation matrix between the resource scheduling operation and the fault propagation chain, and construct the coupling analysis rules to quantify the impact weight of the operation on the fault propagation, and generate the root cause location instructions for the negative coupling scenario; the dynamic priority decision module is used to fit the propagation cost gradient according to the fault propagation cost model by analyzing the historical fault repair time and resource waste rate, input the coupling index, the real-time service level agreement default rate and the fault propagation cost gradient into the weight function, calculate the collaborative priority of the short-term suppression action and the long-term eradication action through the dynamic weight function, and generate the execution sequence of the two types of actions through the asymmetric game strategy; the strategy execution and feedback module is used to call the cloud platform interface to atomically execute the decision action, collect the indicator change data after execution, and dynamically update the fault propagation cost model and coupling analysis rules.

[0006] Furthermore, the cross-layer indicator perception module specifically includes: a containerized probe cluster is deployed on a hybrid cloud node with a microservice architecture, including physical layer probes to collect hardware resource fragmentation indicators, virtual layer probes to monitor container lifecycle events and resource contention behaviors, and application layer probes to track microservice call chain topology and performance deviations; the collection frequency is adjusted according to the dynamic change rate of the indicator, the physical layer adopts low-frequency trigger sampling, the application layer adopts event-driven high-frequency tracking, and the sliding window statistics are used to suppress transient noise; the layered labeling preprocessing unit adds environmental context labels of cloud platform type and service dependency layer to the original data, filters outliers based on the isolation forest algorithm, and outputs a standardized cross-layer indicator set.

[0007] Furthermore, an incremental timing alignment algorithm is used to eliminate the timing deviation between the low-frequency data of the physical layer and the high-frequency data of the application layer. The specific process of constructing a fault propagation map based on a dynamic time window is as follows: the low-frequency data of the physical layer and the high-frequency data of the application layer are dynamically interpolated, and a continuous timing sequence is generated based on the weighted data confidence; the potential phase difference between cross-layer indicators is detected through sliding correlation analysis, and the interpolation anchor point is dynamically adjusted to eliminate the timing offset; the conditional transition probability between cross-layer indicators is calculated based on the aligned timing data, the time window range is dynamically expanded, and the window is contracted when a sudden change in the indicator is detected, to generate a fault propagation map with weighted edges.

[0008] Furthermore, the identification logic for implicitly associated nodes with spatiotemporal coupling characteristics between cross-layer indicators is as follows: the propagation paths across the physical layer, virtual layer, and application layer are extracted from the fault propagation map, and candidate nodes whose transfer probabilities exceed the dynamic threshold are screened; mutual information entropy analysis is performed on the candidate nodes to quantify their dependence on upstream and downstream indicators and eliminate weakly correlated interference items; the directionality of the spatiotemporal causal relationship of the candidate paths is verified through the Granger causality test, and directional disturbances are injected into the high-probability paths to observe the response amplitude of downstream indicators and confirm the effectiveness of spatiotemporal coupling.

[0009] Furthermore, the authenticity of cross-layer causal relationships is verified through targeted disturbance injection, and the specific process of fitting the coupling index based on the correlation matrix of resource scheduling operations and fault propagation chains is as follows: a high-probability correlation path is selected in the fault propagation map, and a controllable disturbance that simulates a sudden increase in physical layer storage latency or limits the virtual layer container network bandwidth is injected into the path source node; the downstream indicator response is observed and the disturbance propagation path and amplitude are recorded, and the predicted path consistency is compared with the original map; the temporal relationship between historical resource scheduling operations and fault propagation chains is extracted, and the correlation matrix of resource scheduling operations and fault propagation chains is constructed. Based on the influence weight of resource scheduling operations on the fault chain length and repair time in the matrix, the coupling index is fitted by gradient descent method.

[0010] Furthermore, coupling analysis rules are constructed to quantify the impact weight of operations on fault propagation. The specific process of generating root cause location instructions for negative coupling scenarios is as follows: based on the association matrix between resource scheduling operations and fault propagation chains, the contribution weight of operations to the length of the fault chain and the repair time is calculated to generate initial coupling analysis rules; a dynamic judgment threshold is set. When the contribution weight of an operation to the fault propagation chain exceeds the threshold, it is marked as a negative coupling operation; nodes associated with negative coupling operations are traced back in the fault propagation graph, and high causal strength nodes that are not covered by historical scheduling operations are screened as candidate root causes; a reverse blocking test is performed on the candidate root causes to verify their interruption effect on the downstream fault chain by restricting their resource access or traffic distribution; a location instruction containing the root cause node identifier, the impact path, and repair suggestions is generated and pushed to the operation and maintenance terminal.

[0011] Furthermore, the specific process of fitting the propagation cost gradient by analyzing the historical fault repair time and resource waste rate according to the fault propagation cost model is as follows: extract the historical fault repair time data and the resource waste rate caused by resource scheduling operations to construct the initial propagation cost function; iteratively optimize the cost function parameters through the gradient descent method, and dynamically adjust the repair time weight and resource waste penalty factor; calculate the current propagation cost gradient based on the real-time fault propagation path length and resource utilization changes; continuously optimize the gradient parameters through policy execution feedback data to adapt to the dynamic changes of the hybrid cloud environment.

[0012] Furthermore, the coupling index, real-time service level agreement default rate, and fault propagation cost gradient are input into the weight function, and the collaborative priority of short-term suppression actions and long-term eradication actions is calculated through the dynamic weight function. The specific process of generating the execution sequence of the two types of actions through the asymmetric game strategy is as follows: a coupling penalty factor is introduced into the dynamic weight function to suppress the priority of scheduling operations that may aggravate fault propagation in high-coupling scenarios; the urgency weight of short-term suppression actions is calculated based on the real-time service level agreement default rate, and the benefit weight of long-term eradication actions is calculated based on the propagation cost gradient; short-term suppression actions and long-term eradication actions are defined as asymmetric game participants, and a benefit function is constructed to quantify their contribution to service availability improvement and fault propagation suppression; the optimal coordination strategy is solved through dynamic Nash equilibrium, short-term suppression actions are prioritized to quickly stop losses, and long-term eradication actions are triggered asynchronously.

[0013] Furthermore, the policy execution and feedback module specifically includes: the atomic execution engine decomposes decision-making actions into independently executable atomic operations such as calling the cloud platform API to trigger the current limiting policy or initiating the storage volume migration task, and ensures the atomicity and consistency of cross-platform operations through the transaction lock mechanism; monitors the changes in cross-layer indicators after execution, captures the inhibitory effect of the action on the fault propagation chain and the impact of resource utilization, adjusts the weight parameters in the fault propagation cost model according to the feedback data, optimizes the coupling analysis rules and enhances the early identification capability of negative coupling scenarios.

[0014] A dynamic optimization method for cloud monitoring service operation and maintenance based on AI intelligent agent, comprising the following steps: S1. Deploy containerized probe clusters to collect physical layer hardware resource fragmentation indicators, virtual layer container lifecycle events and application layer microservice call chain performance deviation data in real time in a hybrid cloud environment, and output a standardized cross-layer indicator set through layered labeling preprocessing and outlier filtering; S2. Connect the cross-layer indicator perception module, eliminate the timing deviation between the low-frequency data of the physical layer and the high-frequency data of the application layer through an incremental timing alignment algorithm, build a fault propagation map based on a dynamic time window, and identify implicit correlation nodes with spatiotemporal coupling characteristics between cross-layer indicators; S3. Receive the fault propagation map output by the spatiotemporal correlation analysis module, and verify the cross-layer causal relationship through directional disturbance injection. The coupling index is fitted according to the correlation matrix between resource scheduling operations and fault propagation chains, and coupling analysis rules are constructed to quantify the impact weight of operations on fault propagation, generating root cause location instructions for negative coupling scenarios; S4. According to the fault propagation cost model, the historical fault repair time and resource waste rate are analyzed to fit the propagation cost gradient, and the coupling index, real-time service level agreement breach rate and fault propagation cost gradient are input into the weight function. The collaborative priority of short-term suppression actions and long-term eradication actions is calculated through the dynamic weight function, and the execution sequence of the two types of actions is generated through the asymmetric game strategy; S5. The cloud platform interface is called to atomically execute decision actions, and the indicator change data after execution is collected to dynamically update the fault propagation cost model and coupling analysis rules.

[0015] The present invention has the following beneficial effects:

[0016] (1) A cloud monitoring service operation and maintenance dynamic optimization system based on AI intelligent agents. Through containerized probe clusters, it realizes the unified collection and standardized processing of heterogeneous data in the physical layer, virtual layer, and application layer, breaking through the single-layer data limitations of traditional monitoring tools and improving the cross-platform resource collaboration capabilities. Based on the incremental time alignment algorithm and dynamic time window, it constructs a fault propagation map, accurately identifies the spatiotemporal coupling characteristics between cross-layer indicators, and solves the problem of false association caused by static rules. Through the verification of causal authenticity through directed disturbance injection and coupling analysis rules, it reduces the cost of manual investigation and improves the positioning accuracy in complex fault scenarios. Combined with asymmetric game strategies, it dynamically balances the priority of short-term suppression actions and long-term eradication actions, avoids conflicts between resource scheduling and fault repair, and improves the overall resilience of the system. By dynamically updating models and rules through execution feedback data, it realizes the continuous iteration of operation and maintenance strategies and adapts to the dynamic changes of hybrid cloud environments.

[0017] (2) A dynamic optimization method for cloud monitoring service operation and maintenance based on AI agents, from data collection, time series alignment to graph construction, realizes the full life cycle management of cross-layer indicators, and eliminates the interference of data islands on operation and maintenance decisions. Through dynamic time window and spatiotemporal coupling characteristic analysis, potential cascading failure paths are identified in advance, and the fault warning capability is improved. The influence weight of resource scheduling on fault propagation is quantified by combining directed perturbation and correlation matrix, which enhances the reliability and explainability of root cause location. Based on the dynamic weight calculation of the fault propagation cost gradient and the real-time service level agreement default rate, a collaborative execution strategy adapted to complex scenarios is generated to reduce the risk of business interruption. By executing a feedback closed loop to drive the adaptive update of models and rules, it ensures that the operation and maintenance strategy always meets the actual environment requirements and improves long-term operation and maintenance efficiency.

[0018] Of course, any product implementing the present invention does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 This is a flow chart of the cloud monitoring service operation and maintenance dynamic optimization system based on AI intelligent body of the present invention.

[0020] Figure 2 This is a flow chart of a method for dynamic optimization of cloud monitoring service operation and maintenance based on AI intelligent agent in the present invention. DETAILED DESCRIPTION

[0021] The embodiments of the present application use a cloud monitoring service operation and maintenance dynamic optimization system and method based on an AI intelligent body to address the complex coupling relationship between cross-layer resources and services in a hybrid cloud environment. By integrating heterogeneous data from the physical layer, virtual layer, and application layer, and combining dynamic causal verification and game decision-making mechanisms, it solves problems such as data fragmentation, hidden fault correlation failures, and conflicts between resource scheduling and fault repair actions in traditional operation and maintenance solutions, thereby achieving closed-loop control of fault warning, root cause location, strategy optimization, and self-healing execution, and improving the stability and resource utilization efficiency of the cloud service system.

[0022] The overall idea of the solution in the embodiments of this application is as follows:

[0023] Cross-layer data integration and implicit association mining: Containerized probe clusters are used to collect heterogeneous data from physical layer hardware resources, virtual layer container clusters, and application layer microservice links in real time. Layered labeling and dynamic time series alignment technologies are used to eliminate differences in data semantics and collection frequency, and to build a unified cross-layer standardized indicator set. Based on dynamic time windows and incremental association analysis, fault propagation maps are generated to identify implicit fault paths with spatiotemporal coupling characteristics (such as the cascading effect of microservice call delays caused by storage performance degradation), breaking through the data fragmentation limitations of traditional single-layer monitoring.

[0024] Dynamic causal verification and coupling quantification: Inject targeted disturbances into the fault propagation map (such as simulating a sudden increase in physical layer I / O latency) to verify the authenticity of the causal relationship between cross-layer indicators. Combined with the correlation matrix of historical resource scheduling operations and the fault propagation chain, a coupling index is fitted and analysis rules are constructed to quantify the impact of operations on fault propagation (such as the probability of cascading fault recurrence caused by expansion operations). Root cause location instructions are generated for negative coupling scenarios to address the problems of low manual troubleshooting efficiency and high false alarm rates.

[0025] Multi-objective collaborative decision-making and closed-loop optimization: Based on a fault propagation cost model and real-time service level agreement (SLA) breach rates, a dynamic weighting function is used to calculate the collaborative priority between short-term suppression actions (throttling, degradation) and long-term eradication actions (storage migration, service reconfiguration). An asymmetric game strategy is introduced to balance resource conflicts between these two types of actions, generating an optimal execution sequence. After policy execution, the propagation cost model weights and coupling analysis rules are dynamically updated using feedback data, forming an adaptive closed-loop "perception-decision-execution-optimization" system to continuously improve fault recovery efficiency and resource utilization in hybrid cloud environments.

[0026] See also Figure 1, an embodiment of the present invention provides a technical solution: a cloud monitoring service operation and maintenance dynamic optimization system based on AI intelligent body, comprising the following steps: a cross-layer indicator perception module, a spatiotemporal correlation analysis module, a causal verification and coupling analysis module, a dynamic priority decision module, and a policy execution and feedback module; the cross-layer indicator perception module is used to deploy a containerized probe cluster, collect physical layer hardware resource fragmentation indicators, virtual layer container life cycle events and application layer microservice call chain performance offset data in real time in a hybrid cloud environment, and output a standardized cross-layer indicator set through layered labeling preprocessing and outlier filtering; the spatiotemporal correlation analysis module is used to connect the cross-layer indicator perception module, eliminate the timing deviation of low-frequency data of the physical layer and high-frequency data of the application layer through an incremental timing alignment algorithm, build a fault propagation map based on a dynamic time window, and identify implicit correlation nodes with spatiotemporal coupling characteristics between cross-layer indicators; the causal verification and coupling analysis module is used to The fault propagation map output by the receiving spatiotemporal correlation analysis module is used to verify the authenticity of the cross-layer causal relationship through targeted disturbance injection, and a coupling index is fitted according to the correlation matrix between resource scheduling operations and fault propagation chains. A coupling analysis rule is constructed to quantify the impact weight of operations on fault propagation, and generate root cause location instructions for negative coupling scenarios. The dynamic priority decision module is used to fit the propagation cost gradient based on the fault propagation cost model by analyzing the historical fault repair time and resource waste rate. The coupling index, real-time service level agreement default rate and fault propagation cost gradient are input into the weight function. The collaborative priority of short-term suppression actions and long-term eradication actions is calculated through the dynamic weight function, and the execution sequence of the two types of actions is generated through an asymmetric game strategy. The strategy execution and feedback module is used to call the cloud platform interface to atomically execute decision actions, collect indicator change data after execution, and dynamically update the fault propagation cost model and coupling analysis rules.

[0027] In this implementation, the cross-layer metric awareness module deploys a containerized probe cluster to continuously monitor key operational status indicators at different layers (physical, virtual, and application) in a hybrid cloud environment, preprocessing and standardizing them to output a comparable cross-layer metric set. A containerized probe cluster refers to a set of lightweight monitoring programs deployed based on container technology (such as Docker), which can flexibly adapt to multi-cloud environments and achieve highly scalable and low-intrusive data collection. Physical-layer hardware resource fragmentation indicators, such as CPU idle segments and discontinuous memory usage, reflect the fragmented nature of resource usage and can affect performance scheduling efficiency. Container lifecycle events include state change events such as container start, pause, migration, and destruction. Microservice call chain performance drift data refers to the dynamic changes in metrics such as response latency and error rate along the inter-service call path, used to analyze application-layer performance bottlenecks. Layered labeling preprocessing classifies and labels collected data based on source layer and service category to improve contextual matching in subsequent analysis. Outlier filtering eliminates invalid data caused by transient fluctuations, collection errors, or network latency to ensure the quality of input metrics. The spatiotemporal correlation analysis module eliminates timing errors between data at different levels and explores implicit correlations between indicators within dynamic time windows to construct fault propagation paths. The incremental time series alignment algorithm, based on a sliding window, dynamically adjusts the alignment of data points during continuous data updates to achieve synchronous matching between high-frequency and low-frequency indicators. The fault propagation graph is a graph structure based on a probabilistic model, with nodes representing indicators, edges representing potential fault transmission paths, and weights representing propagation probabilities. Implicit correlation nodes are indicator or event nodes that appear to have no direct correlation but exhibit highly correlated behavior under certain conditions or time periods. The causal verification and coupling analysis module verifies the existence of true causal relationships between cross-layer indicators and quantifies the impact of resource operations on fault propagation, thereby locating the root cause of the fault. Targeted perturbation injection injects small changes (such as resource upgrades or upgrades) into specific components or parameters without affecting system stability. The module observes the response to the overall indicator chain to identify causal relationships. Correlation matrix between resource scheduling operations and propagation chains: Construct an influence matrix between resource scheduling actions (such as container migration) and indicator changes to fit causal paths. Coupling index: A numerical indicator that quantifies the strength of mutual influence between indicators or operations. The higher the value, the stronger the coupling relationship. Negative coupling scenario: refers to certain operations or resource scheduling that actually aggravate system failures or performance degradation. Root cause location instructions: Based on the above analysis, generate instructions to accurately locate the core nodes or operations in the system that cause negative propagation. Dynamic priority decision module: Dynamically prioritize candidate actions based on multiple factors, balance short-term stop-loss and long-term governance, and formulate the optimal operation and maintenance strategy sequence. Fault propagation cost model: Analyze the time consumption and resource waste in past fault events to quantify the evolution cost of faults.Service Level Agreement (SLA) Default Rate: This refers to the percentage of services (such as latency and availability) that fail to meet contractual or platform requirements. Propagation Cost Gradient: This refers to the marginal contribution of different nodes or paths to the overall system cost increase. Dynamic Weight Function: This is used to dynamically adjust the weights of various input factors based on actual conditions (such as service level changes). Asymmetric Game Strategy: Given the asymmetric payoffs between different actions, a game model is used to solve the optimal response strategy for both parties, achieving globally optimal coordinated action. Policy Execution and Feedback Module: This module automatically executes policy actions by invoking the underlying cloud platform interface and collects feedback data for model iteration, achieving closed-loop optimization. Atomic Execution: This module breaks down policies into independently executable minimum units (such as restarting a service or migrating a container) to ensure system controllability and security. Post-execution Indicator Change Data: This includes the difference in indicators before and after an action is executed, as well as the rate of change, used to evaluate the effectiveness of the action. Dynamic Update of Cost Models and Rules: This module uses execution feedback to adjust decision models and causal rule weights in real time, enabling adaptive model evolution.

[0028] Specifically, the cross-layer indicator perception module includes: a containerized probe cluster deployed on a hybrid cloud node with a microservice architecture, including physical layer probes to collect hardware resource fragmentation indicators, virtual layer probes to monitor container lifecycle events and resource contention behaviors, and application layer probes to track microservice call chain topology and performance deviations; the collection frequency is adjusted according to the dynamic change rate of the indicator, the physical layer adopts low-frequency trigger sampling, the application layer adopts event-driven high-frequency tracking, and the sliding window statistics are used to suppress transient noise; the layered labeling preprocessing unit adds environmental context labels of cloud platform type and service dependency layer to the original data, filters outliers based on the isolation forest algorithm, and outputs a standardized cross-layer indicator set.

[0029] In this implementation, the cross-layer metric awareness module encapsulates probes in a containerized manner and deploys them on multiple nodes in a hybrid cloud environment according to a microservices architecture. This deployment approach allows probes to be flexibly scaled based on actual monitoring needs, adapting to different cloud platforms (such as private clouds, public clouds, or edge computing nodes), and independently updating and dynamically expanding their functionality. Physical-layer probes are primarily deployed on physical servers or bare metal environments to collect fine-grained information about the usage of underlying hardware resources, including but not limited to information about distributed idle CPU cores, the number of uncontiguous memory regions, and transient fluctuations in disk access. This information helps identify resource fragmentation and provides a foundation for efficient scheduling. Virtual-layer probes are deployed on virtualization platforms or container orchestration systems (such as Kubernetes) to monitor container lifecycle events (such as creation, startup, and termination) and resource contention with other containers or processes. By recording container scheduling failures and resource preemption, they help identify potential system bottlenecks or resource misallocation issues. Application-layer probes, integrated into the service gateway or middleware of each microservice, capture the call chain path between services and key performance indicators such as response time and latency variation for each call. This mechanism restores the service dependency graph and monitors performance drift between services, helping to quickly locate the root cause of faults or abnormal performance nodes. The cross-layer metric perception module adaptively adjusts the sampling frequency based on the changing trends of various metrics. For physical layer metrics with small fluctuations or slow-changing trends, a low-frequency sampling strategy is used to conserve computing resources. For application layer metrics with rapid fluctuations, such as response time, high-frequency sampling or an event-driven approach is used to ensure data timeliness and accuracy. The physical layer typically performs periodic sampling at fixed intervals, such as once per minute for resource utilization. The application layer uses an event-triggered approach. For example, microservice request response timeouts or spikes in error rates immediately trigger sampling and encrypted storage. This tiered sampling strategy balances system resource consumption with monitoring accuracy. To reduce misjudgments caused by transient data anomalies, the module introduces a sliding window mechanism for smoothing. This mechanism aggregates data over a continuous period (e.g., taking the average or median), effectively suppressing noise caused by short-term disturbances such as network jitter and load spikes, thereby improving the stability of anomaly identification. To achieve cross-platform and cross-level data fusion processing, the module embeds rich environmental context tags in the original monitoring data, including metadata such as the cloud platform type, the service layer to which the node belongs, the application name, and the geographical deployment location. This tag system facilitates subsequent clustering analysis, location analysis, and intra-system correlation analysis. During the data cleaning phase, the module uses the isolation forest algorithm to perform unsupervised removal of abnormal samples. This method can identify extreme or distorted data based on the ease with which the sample can be isolated without prior knowledge, further improving the overall quality and robustness of cross-layer indicators and preventing single-point anomalies from interfering with system judgment.Ultimately, the module outputs the processed cross-layer metric data in a unified format, forming a standardized set of metrics that can be used by downstream systems. This standardization includes unified metric names, data unit conversion, timestamp format alignment, and label structure unification, ensuring that heterogeneous data from multiple sources can be consistently parsed and utilized.

[0030] Specifically, an incremental timing alignment algorithm is used to eliminate the timing deviation between the low-frequency data of the physical layer and the high-frequency data of the application layer. The specific process of constructing a fault propagation map based on a dynamic time window is as follows: the low-frequency data of the physical layer and the high-frequency data of the application layer are dynamically interpolated, and a continuous timing sequence is generated based on the weighted data confidence; the potential phase difference between cross-layer indicators is detected through sliding correlation analysis, and the interpolation anchor point is dynamically adjusted to eliminate the timing offset; the conditional transition probability between cross-layer indicators is calculated based on the aligned timing data, the time window range is dynamically expanded, and the window is contracted when a sudden change in the indicator is detected, to generate a fault propagation map with weighted edges.

[0031] In this implementation, low-frequency data from the physical layer and high-frequency data from the application layer often exhibit timing deviations due to sampling frequency differences. To eliminate this deviation, dynamic interpolation of the low-frequency and high-frequency data is required to generate a continuous time series. Data confidence weighting: The confidence level of each data point is determined based on the reliability of the data source. For example, data from different probes may vary in accuracy, and data with higher confidence levels is given greater weight. During the interpolation process, these weights are used to adjust the influence of each data point. Interpolation: Interpolation typically uses linear interpolation or more complex interpolation methods (such as spline interpolation). Continuous time series data is generated based on weights and data confidence. The goal of this step is to fill gaps between data and ensure a smooth transition in the data series. Sliding Correlation Analysis to Detect Potential Phase Differences: Once continuous time series data is obtained through interpolation, the next step is to perform sliding correlation analysis to detect potential timing deviations (i.e., phase differences) between cross-layer indicators. The core idea of this step is to identify time differences between physical and application layer data and dynamically adjust data alignment. Sliding Correlation Analysis: This analysis method detects potential phase differences by calculating correlations between cross-layer indicator data. Specifically, a sliding window performs correlation calculations at different time points, gradually adjusting the data alignment point (i.e., the interpolation anchor point). If a large deviation in the indicator value within the time window is found, the interpolation anchor point is adjusted to reduce this deviation. Formula: ;in: and They are the indicator values of the physical layer and application layer respectively. is the time delay, indicating the possible phase difference. and yes and The mean of . It is the size of the time window that determines the range of the sliding window. The correlation calculated by the sliding window can identify and correct the phase difference between the data. Calculation of conditional transition probability and dynamic adjustment of the time window After completing the alignment of the time series data, the next step is to calculate the conditional transition probability between cross-layer indicators, that is, to predict the probability of the next state based on the state of the current indicator. Conditional transition probability: The conditional transition probability between cross-layer indicators describes the probability of the indicator transitioning from one state to another given the current state. During the calculation, the correlation at different time points is taken into account, and the transition probability model is fitted through historical data. Formula: ;in: is the state of the physical layer, It is the next state of the application layer. Is the physical layer status Application layer status The denominator is the number of times all possible The total number of values is used to normalize the probability. Dynamic time window expansion and contraction: The range of the time window will be dynamically adjusted according to the changes in the data. If there is a sudden change in cross-layer indicators (for example, a failure or performance anomaly), the time window will be narrowed to focus on analyzing recent data to ensure that the instantaneous changes when the failure occurs can be accurately captured. Conversely, when the system is stable, the time window will be expanded to cover data over a longer period of time to capture the overall trend of the system. Formula: ;in: is the size of the time window. and are the minimum and maximum sizes of the time window, respectively. Generate a fault propagation graph with weighted edges. By calculating the conditional transition probability and adjusting the dynamic time window, a fault propagation graph with weighted edges is finally generated. In this graph, nodes represent indicators at different levels, and edges represent the correlation and fault propagation probability between them. The weight represents the strength of fault propagation between different indicators. Graph with weighted edges: The size of the weight reflects the possibility of a fault propagating from one level to another. The larger the weight, the stronger the fault propagation between the two indicators, and the higher the probability of a fault occurring. Formula: ;in: It is a slave node To Node The weight of the edge. is a node To Node The conditional transition probability of . Is a reliability factor, indicating that the node and The connection quality between them can be evaluated based on factors such as system stability and device health. Through the above steps, an accurate fault propagation map is constructed, which helps to implement more accurate fault warning and repair strategies in cloud monitoring services.

[0032] Specifically, the identification logic for implicitly associated nodes with spatiotemporal coupling characteristics between cross-layer indicators is as follows: the propagation paths across the physical layer, virtual layer, and application layer are extracted from the fault propagation map, and candidate nodes whose transfer probability exceeds the dynamic threshold are screened; mutual information entropy analysis is performed on the candidate nodes to quantify their dependence on upstream and downstream indicators and eliminate weakly correlated interference items; the directionality of the spatiotemporal causal relationship of the candidate paths is verified through the Granger causality test, and directional disturbances are injected into the high-probability paths to observe the response amplitude of downstream indicators and confirm the effectiveness of spatiotemporal coupling.

[0033] In this implementation scheme, candidate nodes whose transfer probability exceeds the dynamic threshold are screened. In the fault propagation map, the propagation paths between the physical layer, virtual layer, and application layer across layers are regarded as potential correlation paths. First, it is necessary to screen out those candidate nodes whose transfer probability exceeds a certain dynamic threshold. Transfer probability: The transfer probability represents the probability of state change from one node to another. When the probability exceeds the set dynamic threshold, it means that the path may involve a strong correlation. Therefore, these nodes are considered candidate nodes. Dynamic threshold: The threshold setting is dynamic and is adjusted based on historical data and system status factors. The threshold is adjusted according to the fluctuation of different system loads or the fault history to improve the sensitivity of identification. Mutual information entropy analysis For the screened candidate nodes, the mutual information entropy analysis is then used to quantify the intensity of their dependence on upstream and downstream indicators. Mutual information entropy can help us evaluate the dependency between two variables and quantify the correlation between cross-layer indicators. Mutual information entropy: Mutual information entropy quantifies the common information between two variables and measures the degree of information sharing between them. The larger the mutual information entropy, the stronger the dependency between the two variables. Formula: ;in: and These are two variables between cross-layer indicators (indicators of the physical layer and the application layer). yes and The joint probability distribution of . and They are and The marginal probability distribution of . express and Mutual information between nodes. By calculating mutual information entropy, we can identify which nodes have strong correlations, and then eliminate interference items with weaker dependencies and retain important correlation paths. Granger Causality Test After mutual information entropy analysis, we use Granger Causality Test on candidate paths to verify the directionality of their spatiotemporal causal relationship. Granger Causality Test can help us determine whether one variable can predict the future changes of another variable. Granger Causality Test: Through the analysis of historical data, if a variable Historical data can significantly improve the understanding of variables The predictive power of the current value, we can say Granger causality This can help us understand the causal relationships between different levels in the system. The formula is as follows: ;in: is the target variable (such as performance indicators at the application layer). is the explanatory variable (such as resource consumption at the physical layer). is a constant term, and is the lag coefficient, which indicates the impact of historical data on the current value. is the number of lags, which represents the time range of the variable's history. is the error term, representing the portion of the model that cannot be explained. Granger causality tests can help us identify whether there is a temporal causal relationship between cross-layer indicators, that is, determine whether indicators at one level affect indicators at other levels. Directed perturbations and downstream indicator responses: Once we have confirmed the directionality of the causal relationship between candidate pathways, the next step is to inject directed perturbations into the high-probability pathways and observe the magnitude of the responses in downstream indicators. Directed perturbations: Directed perturbations involve artificially changing the state of certain nodes (for example, by simulating load changes or resource consumption) to test the system's response to these changes. The purpose of directed perturbations is to verify whether causal relationships have a real impact in the system. Downstream indicator response magnitude: We observe changes in downstream indicators after directed perturbations, specifically measured by the magnitude of the response. A larger magnitude of the response indicates stronger spatiotemporal coupling in the causal pathway. Significant changes in downstream indicators indicate a strong effectiveness of spatiotemporal coupling. Finally, after conducting directed perturbation and response testing, we confirm the effectiveness of spatiotemporal coupling. If the perturbation has a significant impact on downstream indicators, we can confirm that these pathways are valid spatiotemporal coupling pathways. Spatiotemporal coupling validity: Spatiotemporal coupling validity means that there are strong causal relationships between cross-layer indicators and that these relationships have a real impact on system performance. This relationship can be further used to optimize the system, detect potential failures, and allocate resources.

[0034] Specifically, the authenticity of cross-layer causal relationships is verified through targeted disturbance injection, and the specific process of fitting the coupling index based on the correlation matrix of resource scheduling operations and fault propagation chains is as follows: a high-probability correlation path is selected in the fault propagation map, and a controllable disturbance that simulates a sudden increase in physical layer storage latency or limits the virtual layer container network bandwidth is injected into the path source node; the downstream indicator response is observed and the disturbance propagation path and amplitude are recorded, and the predicted path consistency is compared with the original map; the temporal relationship between historical resource scheduling operations and fault propagation chains is extracted, and the correlation matrix of resource scheduling operations and fault propagation chains is constructed. Based on the influence weight of resource scheduling operations on the fault chain length and repair time in the matrix, the coupling index is fitted by gradient descent method.

[0035] In this implementation, we first select high-probability correlation paths within the constructed fault propagation graph. High-probability paths represent the paths most likely to have causal relationships across the physical, virtual, and application layers. After selecting these paths, we inject simulated perturbations into the source nodes of these paths to verify the authenticity of the causal relationships. Targeted perturbation injection: We inject controlled perturbations into the source nodes. Common perturbations include sudden increases in storage latency at the physical layer or limiting the network bandwidth of containers at the virtual layer. For example, the physical layer might inject a sudden increase in storage latency, while the virtual layer could limit the network bandwidth of containers to simulate resource bottlenecks or failures. Through targeted perturbations, we verify whether the impact of the source node on downstream metrics matches the predicted path derived from the fault propagation graph. If the downstream metrics respond as expected after the perturbation, we verify the existence of a causal relationship along this path. After injecting the perturbation into the source node, we observe the response of the downstream metrics and record the perturbation propagation path and response magnitude. Response observation: We monitor other system metrics (such as application layer performance and virtual layer resource usage) to observe changes after the perturbation injection. If the response amplitude of the downstream indicator is large, it indicates that the causal relationship along the path is valid and the impact of the disturbance is transmitted along the propagation path. Comparing the predicted path: The observed disturbance propagation path is compared with the predicted path in the original graph. If the two are consistent, the causal relationship path is valid. To further analyze the relationship between fault propagation and resource scheduling, we need to extract the temporal relationship between historical resource scheduling operations and fault propagation chains. This step aims to identify potential correlations between resource scheduling operations and fault propagation and understand how these operations affect the length and repair time of fault propagation chains. Temporal Relationship: By analyzing historical data, we construct temporal relationships between resource scheduling operations and fault propagation chains. For example, we can determine whether resource scheduling operations accelerate or delay the occurrence of fault propagation, or whether certain scheduling operations can shorten repair time. Based on the extracted historical data, we construct a correlation matrix between resource scheduling operations and fault propagation chains. This matrix contains the weight of each resource scheduling operation's impact on the length and repair time of the fault propagation chain. Correlation Matrix: In the correlation matrix, each element represents the impact of a resource scheduling operation on the fault propagation chain. Specifically, if an operation significantly shortens the length of the fault propagation chain or the repair time, the value of this element in the matrix is large. Based on the data in the correlation matrix, we use the gradient descent method to fit the coupling index to quantify the degree of coupling between resource scheduling operations and fault propagation chains. Definition of coupling index Coupling index ( ) reflects the relationship between a resource scheduling operation and the fault propagation chain, taking into account the impact of resource scheduling operations on the length of the fault propagation chain, repair time, and propagation path. To fit the coupling index, multiple factors can be introduced and calculated in combination with weights. Coupling index formula: ;in: : Represents the coupling index between the a-th resource scheduling operation and the fault propagation chain. : Total number of resource scheduling operations. : No. The weight of a resource scheduling operation indicates its impact proportion in the entire fault propagation chain. : No. Resource scheduling operations on the fault propagation chain The influence of length. : No. Resource scheduling operations on the fault propagation chain Impact of repair time. : No. Resource scheduling operations on the fault propagation chain The impact of the propagation path is usually the relationship between the path length and the nodes. :Represents the weight coefficients of chain length, repair time and propagation path respectively. Definition of each influencing factor, chain length influence : Indicates resource scheduling operations The impact on the length of the fault propagation chain is usually determined by the expansion or compression of nodes or chains caused by scheduling operations. For example, optimized resource scheduling may reduce the chain length and thus reduce the delay of the propagation chain. ; Repair time impact : Represents the impact of resource scheduling operations on the repair time of the fault propagation chain. Changes in scheduling operations may affect the timeliness of the recovery process and thus affect the repair time. ;Influence of propagation path : Indicates the impact of resource scheduling operations on fault propagation paths. Scheduling operations may activate new paths or affect the propagation effect of existing paths. ; Gradient descent method to fit the coupling index In order to optimize the coupling index, we can use the gradient descent method to tune the parameters to ensure that the model can better adapt to different resource scheduling and fault propagation situations. Gradient descent method to optimize the coupling index parameters as well as The formula is as follows: ;in: is the set of parameters to be optimized. is the learning rate. Is the loss function, which measures the error between the coupling index predicted by the model and the actual coupling index. Loss function definition loss function It is used to measure the error in the fitting process, which takes into account the difference between the actual observed coupling and the calculated coupling. It is defined as follows: ;in: It is a coupling index based on resource scheduling operations and model predictions. is the actual observed coupling index. By minimizing the loss function , we can find the optimal parameter combination, so that the coupling index between resource scheduling operations and fault propagation chains fits more accurately. The dynamic update of the model is updated through the iterative update of the gradient descent method. Each optimization will adjust the parameters , , and , making the calculated coupling index more consistent with the actual fault propagation chain characteristics. As the number of iterations increases, the coupling index will gradually optimize to the optimal configuration. Coupling Index: The coupling index indicates the degree of influence of resource scheduling operations on fault propagation chains. Specifically, it describes the extent to which resource scheduling operations can affect the length and repair time of fault propagation chains. A higher coupling index indicates a stronger correlation between resource scheduling operations and fault propagation chains.

[0036] Specifically, a coupling analysis rule is constructed to quantify the influence weight of the operation on fault propagation. The specific process of generating the root cause location instruction for the negative coupling scenario is as follows: Based on the association matrix between resource scheduling operations and fault propagation chains, the contribution weight of the operation to the length of the fault chain and the repair time is calculated to generate the initial coupling analysis rule; a dynamic judgment threshold is set. When the contribution weight of the operation to the fault propagation chain exceeds the threshold, it is marked as a negative coupling operation; the nodes associated with the negative coupling operation are traced back in the fault propagation graph, and the high causal strength nodes that are not covered by historical scheduling operations are screened as candidate root causes; a reverse blocking test is performed on the candidate root cause to verify its interruption effect on the downstream fault chain by restricting its resource access or traffic distribution; a location instruction containing the root cause node identifier, the impact path and the repair suggestion is generated and pushed to the operation and maintenance terminal.

[0037] In this implementation, an association matrix is first constructed based on the relationship between resource scheduling operations and fault propagation chains. This matrix reflects the impact of each resource scheduling operation on the length, repair time, and propagation path of the fault propagation chain. This matrix quantifies the contribution of each resource scheduling operation to the fault chain. Coupling analysis rules are generated: After the association matrix between resource scheduling operations and fault propagation chains is constructed, coupling analysis rules are generated. The primary goal of these rules is to quantify the impact of each resource scheduling operation on the fault propagation chain and identify operations with negative coupling. Contribution weight calculation: A weight is calculated for the relationship between each resource scheduling operation and the fault propagation chain. The weight reflects the degree of impact of the operation on the length, repair time, and propagation path of the entire fault chain. The higher the contribution weight, the more significant the impact of the resource scheduling operation on the fault chain. Dynamic decision threshold: To identify negatively coupled operations, a dynamic decision threshold is set. When the contribution weight of a resource scheduling operation to the fault propagation chain exceeds this threshold, it is considered a negatively coupled operation. Negatively coupled operations can exacerbate fault propagation and therefore require special attention and handling. Generating Initial Rules: Based on the above analysis, preliminary coupling analysis rules are formed. These rules define what constitutes a negative coupling operation, how to calculate contribution weights, and how to select resource scheduling operations that warrant attention. Identifying Negative Coupling Operations: After generating the coupling analysis rules, the nodes in the fault propagation graph are analyzed to identify those associated with negative coupling operations. Tracing Back Negative Coupling Operation-Associated Nodes: Using the marked negative coupling operations, trace back and analyze their associated nodes. These nodes are potential sources in the fault propagation graph and may be the most critical triggers in the fault propagation chain. Filtering Nodes with High Causal Strength: Among the traced nodes, select those with high causal strength. These nodes are those with the strongest correlation with negative coupling operations. Causal strength can be determined by measuring the influence of a node in the fault propagation process. Nodes Not Covered by Historical Scheduling Operations: Nodes not covered by historical scheduling operations require special attention, as these nodes may be key contributors to fault outbreaks. Root Cause Identification and Verification: Filtered candidate root cause nodes are further verified to determine whether they actually influence the fault propagation chain. Reverse blocking test: Perform a reverse blocking test on the candidate root cause node. By restricting its resource access or traffic distribution, observe whether the downstream fault chain is interrupted. If interrupted, it proves that the node is the root cause in the fault propagation chain. Verification method: During the reverse blocking test, restricting access to resources or traffic will change the state of the fault propagation chain. If the fault propagation path is effectively interrupted in this way, it means that the node is a valid root cause. Generate root cause location instructions: After verifying the candidate root cause, generate location instructions containing the root cause node identifier, the affected path, and repair suggestions. These instructions will be pushed to the operation and maintenance terminal for timely fault recovery and repair.Root cause location instructions: Root cause node identification: This includes the identification of key nodes in the fault propagation chain. Impact path: This identifies the fault propagation path for the root cause node and analyzes its impact on the overall system. Remediation recommendations: Based on the characteristics of the root cause node, remediation suggestions are provided, such as resource optimization, traffic restriction, and path adjustment. Finally, the root cause location instructions are pushed to the operation and maintenance terminal via the automated system for reference and subsequent action by the operation and maintenance personnel. The operation and maintenance personnel then take appropriate remediation measures based on the location instructions to ensure stable system operation.

[0038] Specifically, the specific process of fitting the propagation cost gradient by analyzing the historical fault repair time and resource waste rate according to the fault propagation cost model is as follows: extract the historical fault repair time data and the resource waste rate caused by resource scheduling operations to construct the initial propagation cost function; iteratively optimize the cost function parameters through the gradient descent method, and dynamically adjust the repair time weight and resource waste penalty factor; calculate the current propagation cost gradient based on the real-time fault propagation path length and resource utilization changes; continuously optimize the gradient parameters through policy execution feedback data to adapt to the dynamic changes of the hybrid cloud environment.

[0039] In this implementation plan, a propagation cost model is constructed: First, by extracting historical data, an initial propagation cost function is constructed, which mainly includes two key factors: historical fault repair time and resource waste rate. These factors will serve as the basis for propagation costs and help quantify the time cost and resource consumption in the fault repair process. Repair time and resource waste rate: Repair time: refers to the time consumed from the occurrence of the fault to complete repair, which is usually affected by factors such as system response, fault location, and repair methods. Resource waste rate: refers to the ineffective use or waste of system resources during the fault period, which may be caused by redundant calculations, idle resources, excessive resource allocation, etc. The expression of the initial propagation cost function can be as follows: ;in: : Initial propagation cost. : Historical fault repair time. : Resource waste rate. and : Weight coefficients, which respectively control the contribution of repair time and resource waste to the propagation cost. Use the gradient descent method to optimize the propagation cost function: After constructing the initial propagation cost function, use the gradient descent method to optimize the parameters of the propagation cost function, especially the adjustment of the repair time weight and resource waste penalty factor. This step is to reduce the total propagation cost by continuously iterating the calculation of the gradient and optimizing the weight. The optimization process aims to make the propagation cost function more reflective of the actual fault repair process, and dynamically adjust the weights of the repair time and resource waste to achieve more efficient resource utilization and shorter repair time. Implementation of the gradient descent method: The gradient descent method updates the parameters according to the gradient of the current propagation cost function, and optimizes the parameters. and The update rules are as follows: ; ;in: and are the repair time weight and resource waste penalty factor of the fth iteration respectively. η: learning rate, which controls the magnitude of each update. and They are the partial derivatives of the propagation cost function with respect to the weight parameters respectively. Calculate the current propagation cost gradient: Once the optimization process is completed and the propagation cost function parameters have been adjusted to appropriate values, real-time data can be used to calculate the current propagation cost gradient. At this point, the weights of the impact of repair time and resource waste have been dynamically adjusted, and the propagation cost function can also accurately reflect the actual situation. Real-time propagation path and resource utilization: Propagation path length: represents the span of fault propagation, which can be regarded as the number of links from the source of the fault to the end point. The longer the fault propagation path, the longer the time required for repair is usually, and the greater the resource consumption. Resource utilization: indicates the actual utilization of system resources during the fault repair process. High resource utilization can reduce resource waste, thereby reducing propagation costs. The calculation formula for the propagation cost gradient can be expressed as: ;in: : The gradient of the current propagation cost. : Real-time fault repair time variation. Real-time changes in resource waste rate. Policy execution feedback and gradient optimization: To adapt to dynamic changes in the hybrid cloud environment, gradient parameters must be continuously optimized using feedback data from policy execution. Feedback data can reflect improvements or deteriorations in system performance after policy adjustments, which in turn influences changes in propagation costs. Feedback data: Policy execution effectiveness: By monitoring the length of fault propagation paths and resource utilization in the system after policy execution, feedback on dynamic system changes is obtained. Gradient adjustment: The gradient parameters in the propagation cost function are adjusted based on feedback data to more accurately optimize repair time and resource waste during future fault remediation processes. Policy adjustment: The parameters in the gradient descent process are adjusted based on real-time feedback, allowing the model to adapt to changes in the hybrid cloud environment, ensuring efficient utilization of system resources and rapid fault remediation. Optimization process summary: By optimizing the propagation cost model using gradient descent, the system can dynamically adjust the impact weights of repair time and resource waste, thereby achieving the goal of real-time adaptation to changes in the hybrid cloud environment. During the continuous optimization process, the model not only accurately predicts fault remediation costs but also automatically adjusts to environmental changes, improving system efficiency and responsiveness.

[0040] Specifically, the coupling index, real-time service level agreement default rate, and fault propagation cost gradient are input into the weight function, and the collaborative priority of short-term suppression actions and long-term eradication actions is calculated through the dynamic weight function. The specific process of generating the execution sequence of the two types of actions through the asymmetric game strategy is as follows: a coupling penalty factor is introduced into the dynamic weight function to suppress the priority of scheduling operations that may aggravate fault propagation in high-coupling scenarios; the urgency weight of short-term suppression actions is calculated based on the real-time service level agreement default rate, and the benefit weight of long-term eradication actions is calculated based on the propagation cost gradient; short-term suppression actions and long-term eradication actions are defined as asymmetric game participants, and a benefit function is constructed to quantify their contribution to service availability improvement and fault propagation suppression; the optimal collaborative strategy is solved through dynamic Nash equilibrium, short-term suppression actions are prioritized to quickly stop losses, and long-term eradication actions are triggered asynchronously.

[0041] In this implementation scheme, a coupling penalty factor is introduced and the priority is adjusted: when calculating the collaborative priority of short-term suppression actions and long-term eradication actions, the coupling index needs to be considered first, which helps to identify operations that may exacerbate fault propagation. Scheduling operations in high-coupling scenarios may have a negative impact on the stability of the system and should therefore be given a lower priority. Coupling penalty factor: High-coupling scenarios are penalized by introducing a coupling penalty factor γcoupling. The coupling index ζ reflects the close relationship between indicators at all levels in fault propagation. A higher degree of coupling generally indicates a higher risk of fault propagation. The formula is: ;in: : The priority weight of the scheduling operation, taking into account the influence of the coupling penalty factor. : Basic weight coefficient, used to adjust the impact of priority. : Coupling penalty factor, which indicates the penalty intensity for high coupling scenarios, and its value range is between [0, 1]. : Coupling index, which reflects the strength of association between layers. The above formula can effectively suppress those scheduling operations that may aggravate the propagation of faults. Calculate the urgency weight of short-term suppression action: Short-term suppression action refers to measures that respond quickly and slow down the propagation of faults, and is usually used to intervene when the service default rate is high. Therefore, the urgency weight of short-term suppression action should be calculated based on the real-time service level agreement (SLA) default rate. SLA default rate: default rate It reflects the probability that the current system fails to meet service commitments on time. The higher the default rate, the greater the urgency of short-term suppression actions. The formula is: ;in: : Urgency weight of short-term inhibitory action. : Basic weight coefficient, used to adjust the urgency weight. : Real-time SLA breach rate. : The impact index of default rate on urgency weight, usually , so that the urgency weight increases significantly when the default rate increases. Through this formula, the execution urgency of the short-term suppression action can be dynamically adjusted according to the SLA default rate. Calculating the benefit weight of the long-term eradication action: The long-term eradication action aims to reduce the probability of future failures and improve the overall stability of the system by eliminating the root cause of the failure. Therefore, the benefit weight of the long-term eradication action should be associated with the current propagation cost gradient. Propagation cost gradient: The propagation cost gradient reflects the cost changes in the current fault propagation process. The higher the propagation cost, the greater the benefit of the long-term eradication action. The formula represents: ;in: : The payoff weight of long-term eradication actions. : Basic weight coefficient, used to adjust the benefit weight of long-term eradication actions. : Current propagation cost gradient, reflecting the economic cost of fault propagation. : The impact index of communication cost on benefit weight, usually , so that when the cost is higher, the priority and benefit of the eradication action are greater. Construct a game model and calculate the collaborative priority: In this step, short-term suppression actions and long-term eradication actions are regarded as two participants in the game model, and these two types of actions have different goals and benefits. Short-term suppression actions focus on quickly reducing the impact of faults, while long-term eradication actions focus on eliminating the root cause of the fault. Asymmetric game model: The contributions and benefits of short-term suppression actions and long-term eradication actions are different, so an asymmetric game strategy is needed to calculate the collaborative priority. Through the game model, the contribution of the two to fault propagation suppression and service availability can be quantified, and the optimal execution order can be finally determined. Formula representation: ;in: and They represent the execution strategies for short-term suppression actions and long-term eradication actions, respectively. and Denote the payoff functions for short-term suppression actions and long-term eradication actions, respectively, taking into account urgency weights, payoff weights, and priorities. Dynamic Nash equilibrium solves the optimal coordination strategy: By solving the Nash equilibrium of the game model, the optimal coordination strategy can be obtained, that is, how to choose short-term suppression actions and long-term eradication actions under given conditions to maximize the system's service availability improvement and fault propagation suppression effects. Execution order: Based on the Nash equilibrium solution of the game, the system will prioritize short-term suppression actions to quickly stop losses; then, asynchronously trigger long-term eradication actions to ensure a fundamental resolution of the fault. Through this method, the system's operating strategy can be dynamically adjusted according to the actual fault propagation situation, resource utilization, and service default rate to achieve efficient fault suppression and service assurance.

[0042] Specifically, the policy execution and feedback module includes: the atomic execution engine decomposes decision actions into independently executable atomic operations such as calling the cloud platform API to trigger the current limiting policy or initiate the storage volume migration task, and ensures the atomicity and consistency of cross-platform operations through the transaction lock mechanism; monitors the changes in cross-layer indicators after execution, captures the inhibitory effect of the action on the fault propagation chain and the impact of resource utilization, adjusts the weight parameters in the fault propagation cost model according to the feedback data, optimizes the coupling analysis rules and enhances the early identification capability of negative coupling scenarios.

[0043] In this implementation, this module consists of an atomic execution engine, a transaction control mechanism, a cross-layer indicator monitoring unit, and a feedback parameter tuning mechanism, designed to achieve reliable execution and dynamic optimization of scheduling policies. Atomic Execution Engine: The atomic execution engine is used to break down unstructured, high-level policy instructions into atomic operations at the cloud resource operation level. Each policy action is broken down into several independently executable atomic tasks that directly control the cloud platform resource management interface (API). Examples of atomic operations include: calling the rate limiting policy API to trigger service instance downgrade; initiating a volume migration task between different nodes to alleviate I / O bottlenecks or achieve load balancing; and ensuring atomicity and consistency. To avoid inconsistencies in intermediate states during cross-platform or cross-service execution, a distributed transaction lock mechanism is introduced to ensure the following two aspects: Operation atomicity: Any composite operation either succeeds in its entirety or rolls back in its entirety. Operation consistency: After the operation is completed, the resource states of each subsystem are in a consistent and controllable state. Monitoring content includes: the changing trends of key cross-layer indicators (such as system load, service latency, link traffic, instance failure rate, etc.) before and after policy execution; the status response of nodes in the fault propagation chain to quantify the interruption or mitigation effect of scheduling actions on potential propagation paths; and changes in resource utilization to evaluate the efficiency of action execution in allocating computing power, storage, or bandwidth resources. Capture method: Collect indicators at all layers through an integrated monitoring system; use a propagation chain analyzer to annotate the indicator response delay and slope changes between "controlled nodes" and "potential propagation paths." Model optimization goal: Based on monitoring feedback, dynamically adjust the fault propagation cost model and coupling analysis model used in the decision-making process to improve the targetedness and foresight of policy responses. Tuning content includes: Propagation cost weight adjustment: If a certain type of action has a significant effect on blocking the propagation chain in multiple execution rounds, its cost-effectiveness weight in the cost model can be increased (for example, increasing the propagation reduction benefit value corresponding to its unit execution cost). Coupling analysis rule revisions: Analyze the coordinated changes in cross-layer indicators before and after an action to identify "negative coupling scenarios" (i.e., operations on one layer causing reverse interference on another layer) that are not covered by the initial rules. This expands the rule base and improves the sensitivity of coupling discrimination. Early warning capabilities are enhanced: Key indicator change patterns identified in feedback are fed back into the training set to update early recognition models (such as the attention mechanism feature extraction network used for negative coupling identification).

[0044] See also Figure 2, a cloud monitoring service operation and maintenance dynamic optimization method based on AI intelligent agent, including the following steps: S1. Deploy containerized probe clusters to collect physical layer hardware resource fragmentation indicators, virtual layer container lifecycle events and application layer microservice call chain performance deviation data in real time in hybrid cloud environments, and output standardized cross-layer indicator sets through layered labeling preprocessing and outlier filtering; S2. Connect cross-layer indicator perception modules, eliminate the timing deviation between low-frequency data in the physical layer and high-frequency data in the application layer through incremental timing alignment algorithm, build fault propagation maps based on dynamic time windows, and identify implicit correlation nodes with spatiotemporal coupling characteristics between cross-layer indicators; S3. Receive the fault propagation maps output by the spatiotemporal correlation analysis module, and verify cross-layer causal relationships through targeted disturbance injection. Authenticity, and fit the coupling index according to the correlation matrix between resource scheduling operations and fault propagation chains, and construct coupling analysis rules to quantify the impact weight of operations on fault propagation, and generate root cause location instructions for negative coupling scenarios; S4. According to the fault propagation cost model, historical fault repair time and resource waste rate are analyzed to fit the propagation cost gradient, and the coupling index, real-time service level agreement breach rate and fault propagation cost gradient are input into the weight function. The collaborative priority of short-term suppression actions and long-term eradication actions is calculated through the dynamic weight function, and the execution sequence of the two types of actions is generated through the asymmetric game strategy; S5. Call the cloud platform interface to atomically execute decision actions, collect indicator change data after execution, and dynamically update the fault propagation cost model and coupling analysis rules.

[0045] In this implementation plan, S1. This step realizes the low-intrusive collection of multi-source heterogeneous data by deploying a "containerized probe cluster" and constructs a "cross-layer indicator structure system" consisting of the physical layer, virtual layer, and application layer. The physical layer collects content: server CPU usage fragmentation, memory allocation discreteness, hard disk IO fluctuation, etc.; the virtual layer collects content: container restart, migration, life cycle events; the application layer collects content: microservice link call delay, failure rate, throughput mutation, etc.; a "labeling preprocessing mechanism" is used to weight and classify indicators at different levels, and an outlier filter is integrated to improve the quality of training samples; the real-time extraction and normalization of cross-layer fault signs are realized, and a high-dimensional, low-coupling data feature space is constructed to provide structured input for subsequent propagation analysis and causal modeling. S2. An incremental time series alignment algorithm and a dynamic time window composition strategy are proposed to realize the unified modeling of cross-layer data flows under heterogeneous frequency sampling, and capture potential propagation paths and spatiotemporal coupling relationships. An incremental sliding window algorithm is used to synchronize data between the physical layer (e.g., 5-minute sampling) and the application layer (e.g., 5-second sampling). A propagation probability graph is constructed using the relationship between indicator change rates and delayed response. A "hidden associated node discovery mechanism" is introduced to uncover latent nodes that fail to manifest abnormalities in the early stages of propagation. This reveals "cross-layer causal diffusion chains" that traditional single-layer logs cannot capture, significantly enhancing the system's ability to interpret mixed fault paths in complex scenarios. S3. This step incorporates scheduling operations into causal chain modeling, proposes a "scheduling-propagation chain coupling index" and a "negative coupling behavior identification mechanism," and constructs a set of root cause location rules. Coupling Index Calculation: Builds a correlation matrix between scheduling operations and changes in the propagation chain, measuring the impact of scheduling operations on chain extension length and repair time. Negative Coupling Operation Determination: Introduces a dynamic contribution threshold; when an operation significantly amplifies the propagation chain or causes recovery delays, it is marked as a negative coupling source. Root Cause Screening: Traces back through nodes associated with negative operations, eliminating nodes with historical scheduling overlap and interference, retaining only nodes with high causal strength. Reverse Blockage Verification: Conducts resource blocking tests on candidate root cause nodes to verify their disruptive effect on downstream propagation. This effectively avoids traditional root cause location's overreliance on anomaly severity, achieving proactive, multi-path, and multi-source root cause inference from the perspective of causal structure and system behavior. S4. This step pioneers the integration of the coupling index, SLA default rate, and propagation cost gradient into a unified dynamic weighting function, and introduces an asymmetric game mechanism to achieve coordinated ranking of the two types of actions.Dynamic weight function design: Introducing a coupling penalty factor to prevent scheduling operations from generating negative feedback in high-coupling areas; Urgency weighting: Measuring the priority of short-term suppression actions by the degree of SLA breach; Long-term benefit weighting: Calculating the benefits of eradication actions through a propagation cost gradient function; Asymmetric game modeling: Treating both types of actions as game participants; Constructing a dual-objective benefit function based on service availability and propagation reduction capability; Determining an optimal action combination strategy that can evolve with the environment through a dynamic Nash equilibrium solution; Breaking through the rigidity of traditional rule-driven O&M, achieving coordinated control of short-term rapid stop-loss and long-term risk elimination, with extremely high O&M flexibility. S5. This step introduces the "cloud platform API-level atomic execution mechanism" and the "execution-monitoring-optimization feedback loop" to build a closed-loop decision-making system. Atomic operation decomposition: decompose policy actions into API-level tasks (such as bandwidth throttling, computing resource rescheduling, volume migration, etc.); adopt a distributed transaction lock mechanism to ensure execution consistency and uninterrupted operation in multi-cloud heterogeneous systems; feedback collection and model optimization: monitor the changes in key indicators after policy execution in real time; use feedback results to adjust the coupling weight factor and cost gradient function parameters; update negative coupling judgment rules and early identification policy models; realize a complete closed loop of strategy → execution → perception → feedback → re-optimization, supporting dynamic adaptation and rapid response in unstable resource environments.

[0046] In summary, this application has at least the following effects:

[0047] A system and method for dynamic optimization of cloud monitoring service operations and maintenance based on an AI agent uses a containerized probe cluster to achieve real-time collection and labeling preprocessing of multi-dimensional data from the physical, virtual, and application layers, forming a standardized cross-layer indicator set that provides a unified data foundation for subsequent analysis. An incremental time alignment and dynamic propagation map construction algorithm accurately captures cross-layer spatiotemporal coupling relationships, reveals hidden fault links, and improves the accuracy and foresight of operations and maintenance decisions. By fitting a coupling index and constructing an operation-propagation chain correlation matrix, it identifies and locates the root cause of negative scheduling behaviors, mitigating the system instability risks associated with traditional experience-driven scheduling. By integrating service level agreement (SLA) default rates, coupling indices, and fault propagation cost gradients, an asymmetric game strategy is used to output a collaboratively optimized action sequence, improving the overall return on investment and response efficiency of policy execution. A cloud platform API-level atomic execution engine and feedback tuning mechanism enable closed-loop control of the entire process, from policy generation to execution monitoring. This system possesses excellent system evolution capabilities and dynamic adaptability, significantly enhancing the intelligence and resilience of the cloud monitoring system.

[0048] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0049] The present invention is described with reference to flowcharts and / or block diagrams of systems, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0050] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0051] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0052] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0053] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.

Claims

1. A cloud monitoring service operation and maintenance dynamic optimization system based on AI agent, characterized by: The following steps are involved: Cross-layer indicator perception module, spatiotemporal correlation analysis module, causal verification and coupling analysis module, dynamic priority decision module, and strategy execution and feedback module; The cross-layer indicator perception module is used to deploy a containerized probe cluster to collect real-time physical layer hardware resource fragmentation indicators, virtual layer container lifecycle events, and application layer microservice call chain performance deviation data in a hybrid cloud environment. Through layered labeling preprocessing and outlier filtering, it outputs a standardized cross-layer indicator set; The spatiotemporal correlation analysis module is used to connect to the cross-layer indicator perception module, eliminate the timing deviation between the low-frequency data of the physical layer and the high-frequency data of the application layer through an incremental timing alignment algorithm, build a fault propagation map based on a dynamic time window, and identify implicit correlation nodes with spatiotemporal coupling characteristics between cross-layer indicators; The causal verification and coupling analysis module is used to receive the fault propagation map output by the spatiotemporal correlation analysis module, verify the authenticity of cross-layer causal relationships through targeted disturbance injection, fit a coupling index based on the correlation matrix between resource scheduling operations and fault propagation chains, and construct coupling analysis rules to quantify the impact of operations on fault propagation, thereby generating root cause location instructions for negative coupling scenarios. The dynamic priority decision module is used to analyze the historical fault repair time and resource waste rate according to the fault propagation cost model to fit the propagation cost gradient, input the coupling index, the real-time service level agreement breach rate and the fault propagation cost gradient into the weight function, calculate the collaborative priority of the short-term suppression action and the long-term eradication action through the dynamic weight function, and generate the execution sequence of the two types of actions through an asymmetric game strategy; The strategy execution and feedback module is used to call the cloud platform interface to atomically execute decision actions, collect indicator change data after execution, and dynamically update the fault propagation cost model and coupling analysis rules; The incremental timing alignment algorithm is used to dynamically adjust the alignment of data points during continuous data updates to achieve synchronous matching between high-frequency and low-frequency indicators; the negative coupling scenario refers to a scenario where resource scheduling operations exacerbate fault propagation; the propagation cost gradient refers to the marginal contribution rate of different nodes or paths to the overall system cost growth; the fault propagation graph is a graph structure based on a probability model, in which nodes represent indicator items, edges represent potential fault transmission paths, and weights represent propagation probabilities; the directed disturbance injection refers to injecting slight changes into specific components to verify causal relationships; the coupling index is a numerical indicator used to quantify the intensity of mutual influence between indicators or operations; the real-time service level agreement default rate refers to the proportion of service quality that is not provided in accordance with contract or platform requirements; the short-term suppression action is used to quickly curb the impact of the fault; the long-term eradication action is used to fundamentally eliminate the fault factor; the atomic execution decision action is used to decompose the strategy into the smallest operation unit that can be executed independently.

2. The cloud monitoring service operation and maintenance dynamic optimization system based on AI agent according to claim 1 is characterized by: The cross-layer indicator perception module specifically includes: The containerized probe cluster is deployed on hybrid cloud nodes using a microservices architecture. It includes physical-layer probes to collect hardware resource fragmentation indicators, virtual-layer probes to monitor container lifecycle events and resource contention, and application-layer probes to track microservice call chain topology and performance deviations. The acquisition frequency is adjusted according to the dynamic change rate of the indicator. The physical layer uses low-frequency trigger sampling, the application layer uses event-driven high-frequency tracking, and the sliding window statistics suppress instantaneous noise. The hierarchical labeling preprocessing unit adds cloud platform type and service dependency layer environmental context labels to the original data, filters outliers based on the isolation forest algorithm, and outputs a standardized cross-layer indicator set.

3. The cloud monitoring service operation and maintenance dynamic optimization system based on AI agent according to claim 2 is characterized by: The incremental timing alignment algorithm is used to eliminate the timing deviation between low-frequency data at the physical layer and high-frequency data at the application layer. The specific process of constructing a fault propagation map based on a dynamic time window is as follows: Dynamically interpolate low-frequency data at the physical layer and high-frequency data at the application layer, and generate a continuous time series based on data confidence weighting; Detect potential phase differences between cross-layer indicators through sliding correlation analysis and dynamically adjust interpolation anchor points to eliminate timing offsets; The conditional transition probability between cross-layer indicators is calculated based on the aligned time series data, the time window range is dynamically expanded and the window is contracted when a sudden change in the indicator is detected, and a fault propagation graph with weighted edges is generated.

4. The cloud monitoring service operation and maintenance dynamic optimization system based on AI agent according to claim 3 is characterized by: The identification logic of implicit association nodes with spatiotemporal coupling characteristics between cross-layer indicators is as follows: Extract the propagation paths across the physical layer, virtual layer, and application layer from the fault propagation graph, and screen candidate nodes whose transfer probabilities exceed dynamic thresholds; Perform mutual information entropy analysis on candidate nodes to quantify their dependence on upstream and downstream indicators and eliminate weakly correlated interference items; The directionality of the spatiotemporal causal relationship of the candidate paths is verified by Granger causality test, and directional disturbances are injected into high-probability paths to observe the response amplitude of downstream indicators and confirm the effectiveness of spatiotemporal coupling.

5. The cloud monitoring service operation and maintenance dynamic optimization system based on AI agent according to claim 4 is characterized by: The specific process of verifying the authenticity of cross-layer causality through targeted disturbance injection and fitting the coupling index based on the correlation matrix of resource scheduling operations and fault propagation chains is as follows: A high-probability associated path is selected in the fault propagation graph, and a controllable disturbance is injected into the source node of the path to simulate a sudden increase in physical layer storage latency or limit the network bandwidth of the virtual layer container. Observe the response of downstream indicators and record the disturbance propagation path and amplitude, and compare the predicted path consistency with the original map; The temporal relationship between historical resource scheduling operations and fault propagation chains is extracted, and a correlation matrix between resource scheduling operations and fault propagation chains is constructed. Based on the influence weights of resource scheduling operations on the length and repair time of the fault chain in the matrix, the coupling index is fitted using the gradient descent method.

6. The cloud monitoring service operation and maintenance dynamic optimization system based on AI agent according to claim 5 is characterized by: The specific process of constructing coupling analysis rules to quantify the impact of operations on fault propagation and generating root cause location instructions for negative coupling scenarios is as follows: Based on the correlation matrix between resource scheduling operations and fault propagation chains, the contribution weight of the operations to the length of the fault chain and the repair time is calculated to generate the initial coupling analysis rules; Set a dynamic judgment threshold. When the contribution weight of an operation to the fault propagation chain exceeds the threshold, it is marked as a negative coupling operation. In the fault propagation graph, trace back the nodes associated with negative coupling operations and select high causal strength nodes that are not covered by historical scheduling operations as candidate root causes; Perform reverse blocking tests on candidate root causes to verify their interruption effect on the downstream fault chain by restricting their resource access or traffic distribution. Generates a positioning instruction containing the root cause node identifier, impact path, and repair suggestions, and pushes it to the operation and maintenance terminal.

7. The cloud monitoring service operation and maintenance dynamic optimization system based on AI agent according to claim 6 is characterized by: The specific process of fitting the propagation cost gradient based on the fault propagation cost model by analyzing the historical fault repair time and resource waste rate is as follows: Extract historical fault repair time data and resource waste rate caused by resource scheduling operations to construct the initial propagation cost function; Iteratively optimize the cost function parameters through the gradient descent method, and dynamically adjust the repair time weight and resource waste penalty factor; Calculate the current propagation cost gradient based on the real-time fault propagation path length and resource utilization changes; Continuously optimize gradient parameters through policy execution feedback data to adapt to dynamic changes in the hybrid cloud environment.

8. The cloud monitoring service operation and maintenance dynamic optimization system based on AI agent according to claim 7 is characterized by: The coupling index, real-time service level agreement default rate, and fault propagation cost gradient are input into the weight function. The dynamic weight function is used to calculate the collaborative priority of short-term suppression actions and long-term eradication actions. The specific process of generating the execution sequence of the two types of actions through an asymmetric game strategy is as follows: By introducing a coupling penalty factor into the dynamic weight function, the scheduling operation priority that is positively correlated with the fault propagation chain is suppressed; The urgency weight of short-term suppression actions is calculated based on the real-time service level agreement default rate, and the benefit weight of long-term eradication actions is calculated based on the propagation cost gradient; Define short-term suppression actions and long-term eradication actions as participants in an asymmetric game, and construct a payoff function that quantifies their contribution to improving service availability and suppressing fault propagation. The optimal collaborative strategy is solved through dynamic Nash equilibrium, prioritizing short-term suppression actions to quickly stop losses and asynchronously triggering long-term eradication actions.

9. The cloud monitoring service operation and maintenance dynamic optimization system based on AI agent according to claim 8, characterized in that: The strategy execution and feedback module specifically includes: The atomic execution engine breaks down decision-making actions into independently executable atomic operations, such as calling cloud platform APIs to trigger rate limiting policies or initiate storage volume migration tasks. It uses a transaction lock mechanism to ensure the atomicity and consistency of cross-platform operations. Monitor changes in cross-layer indicators after execution, capture the inhibitory effect of actions on the fault propagation chain and the impact of resource utilization, adjust the weight parameters in the fault propagation cost model based on feedback data, optimize the coupling analysis rules and enhance the early identification capability of negative coupling scenarios.

10. A method for dynamic optimization of cloud monitoring service operation and maintenance based on AI agent, applying a system for dynamic optimization of cloud monitoring service operation and maintenance based on AI agent according to any one of claims 1 to 9, characterized in that: The following steps are involved: S1. Deploy a containerized probe cluster to collect real-time data on physical layer hardware resource fragmentation, virtual layer container lifecycle events, and application layer microservice call chain performance deviations in the hybrid cloud environment. Through layered labeling preprocessing and outlier filtering, a standardized cross-layer metric set is output. S2. Connect the cross-layer indicator perception module and use an incremental timing alignment algorithm to eliminate the timing deviation between low-frequency data at the physical layer and high-frequency data at the application layer. This algorithm constructs a fault propagation map based on a dynamic time window and identifies implicit correlation nodes with spatiotemporal coupling characteristics between cross-layer indicators. S3. Receive the fault propagation map output by the spatiotemporal correlation analysis module, verify the authenticity of cross-layer causal relationships through targeted disturbance injection, fit a coupling index based on the correlation matrix between resource scheduling operations and fault propagation chains, and construct coupling analysis rules to quantify the impact of operations on fault propagation. This generates root cause location instructions for negative coupling scenarios. S4. Analyze historical fault repair times and resource waste rates based on the fault propagation cost model to fit a propagation cost gradient. Input the coupling index, real-time service level agreement (SLA) breach rate, and fault propagation cost gradient into a weighting function. This dynamic weighting function calculates the collaborative priorities of short-term suppression actions and long-term eradication actions. An asymmetric game strategy is then used to generate execution sequences for the two types of actions. S5. Call the cloud platform interface to atomically execute decision actions, collect indicator change data after execution, and dynamically update the fault propagation cost model and coupling analysis rules.

Citation Information

Patent Citations

  • Microservice intelligent operation and maintenance system and method oriented to cloud native and application

    CN117009119A

  • Service operation and maintenance method and system based on cloud service architecture

    CN119383115A