Distributed cloud native application computing method for intelligent operation and maintenance

By constructing a dynamic multi-layer dependency map DMLDG and combining GNN and RL agents, the insufficient modeling of emergent faults in distributed cloud-native applications is solved, accurate prediction and security avoidance of faults are achieved, and the stability and resilience of the system are improved.

CN120276805AInactive Publication Date: 2025-07-08BEIJING YINGHUAN TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510447149.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing technology lacks the ability to model emerging failures in distributed cloud-native applications, resulting in blind and unsafe decision-making operations that may aggravate the deterioration of system status, and existing monitoring and operation and maintenance methods are difficult to accurately predict and avoid cascading risks.

Method used

A dynamic multi-layer dependency graph DMLDG is constructed, combined with the graph neural network GNN model and reinforcement learning RL agent, evaluate the side effects of operation and maintenance operations through causal inference CI model, generate and verify safe and efficient operation and maintenance strategies, and form an adaptive cycle of prediction-decision-execution-feedback.

Benefits of technology

It realizes accurate prediction and security avoidance of potential failures in distributed cloud-native applications, significantly reducing the average repair time and operation and maintenance costs, and improving system stability and resilience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
Patent Text Reader

Abstract

The invention discloses an intelligent operation and maintenance distributed cloud native application computing method, and relates to the technical field of cloud computing, and the method comprises the steps: collecting heterogeneous data of an application layer, a middleware layer, a container layer and a service governance layer, constructing a dynamic multilayer dependency graph, and explicitly modeling a cross-layer dependency relationship and implicit dependency caused by resource competition; based on the spatio-temporal features of the spatio-temporal diagram neural network learning map, predicting a cascade fault risk and outputting a high-risk sub-graph; candidate operation and maintenance operations are generated through reinforcement learning, indirect negative influences of the operations on non-target components are evaluated in combination with a causal inference model, and risk filtering operations are verified through a multi-dimensional safety threshold value; and after the safety operation is executed, continuously optimizing the model and the atlas through a feedback closed loop. According to the method, deep insight, accurate risk prediction and self-adaptive decision making of side effect minimization can be carried out on a complex dependency relationship, and the system stability and the operation and maintenance efficiency are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cloud computing, and particularly relates to a distributed cloud native application computing method for intelligent operation and maintenance. Background Art

[0002] Distributed cloud native applications, such as microservice clusters built based on the Go language, Redis in-memory databases, Kubernetes containerized deployment, and service mesh governance architectures, have become the mainstream mode of modern application development. Although such architectures have high availability and scalability, their large number of components, complex dependencies, and dynamically changing states make traditional operation and maintenance methods and existing AIOps solutions still insufficient in dealing with the following deep problems:

[0003] Problem 1: Insufficient depth of observation and prediction: Existing monitoring technologies rely on distributed tracing and time series data analysis, but cannot effectively capture implicit dependency relationships across microservices, cache instances, container layers, and service mesh control planes, making it difficult to predict emergent failures caused by dynamic interactions. For example, the performance jitter of a certain Go service may indirectly cause a sharp increase in the API latency of other non-related services through shared Redis resource competition, and such cascading risks cannot be detected in time due to the lack of system-level health state modeling capabilities.

[0004] Problem 2: Insufficient accuracy and security of decision-making and execution: Existing automated operation and maintenance strategies are mostly based on simple rules and cannot evaluate the global impact of operations on complex coupled systems. For example, scaling out a Go service may exhaust node resources or increase the pressure on Redis, and adjusting service mesh routing may cause user requests to fail. Existing technologies are difficult to generate accurate operation and maintenance operation sequences that can solve local anomalies and minimize chain reactions.

[0005] The above problems are deeply coupled: The lack of the ability to model emergent failures leads to blindness in decision-making, and unsafe operations may exacerbate the deterioration of the system state. Existing technologies have not effectively coordinated to solve the operation and maintenance problems under such complex architectures. Summary of the Invention

[0006] The present invention provides a distributed cloud native application computing method for intelligent operation and maintenance to solve the problems in the prior art that the lack of the ability to model emergent failures leads to blindness in decision-making, and unsafe operations may exacerbate the deterioration of the system state.

[0007] The present invention provides a distributed cloud native application computing method for intelligent operation and maintenance, including the following steps:

[0008] Collect and fuse real-time data from multiple heterogeneous data sources within the distributed cloud native application architecture;

[0009] Based on the fused data, construct and dynamically maintain a dependency map DMLDG that represents entities within the application architecture and their dynamic multi-level dependency relationships;

[0010] Take the dynamically maintained dependency map as input, apply a graph neural network GNN model to learn the spatio-temporal features of the map, and predict potential emergent fault risks in the application architecture, outputting risk indication information;

[0011] Based on the risk indication information, use a reinforcement learning RL agent to generate candidate operation and maintenance operations;

[0012] Before executing the candidate operation and maintenance operations, use a causal inference CI model to evaluate the possible side effects of the candidate operation and maintenance operations on the application architecture;

[0013] According to the side effect evaluation results, verify or adjust the candidate operation and maintenance operations to obtain the final operation and maintenance operations to be executed;

[0014] Execute the final operation and maintenance operations to avoid risks or recover from faults.

[0015] Preferably, the steps of constructing and dynamically maintaining the dependency map DMLDG specifically include: by analyzing the shared resource access patterns extracted from the interaction data between application layer services and middleware layer performance data, explicitly identifying and modeling the implicit dependency relationships generated by resource competition as specific types of edges in the map; and, by correlating the data of the container layer, node layer, and service governance layer, establishing and dynamically updating the cross-layer dependency edges that span the application layer, middleware layer, container layer, node layer, and service governance layer in the map.

[0016] Preferably, the steps of using the graph neural network GNN model to predict emergent fault risks specifically include: using the GNN model to learn the multi-level spatio-temporal embedding representations of nodes and edges in the dependency map DMLDG, and the embedding representations capture the transfer characteristics of dependency relationships between different levels; and, based on the learned embedding representations, identifying and outputting a key path subgraph with a high probability of anomaly as the risk indication information, and the key path subgraph represents a potential cascading fault mode.

[0017] Preferably, the steps of using the causal inference CI model to evaluate side effects specifically include: representing the candidate operation and maintenance operations as interventions on target nodes or edges in the dependency map DMLDG; using the CI model to perform causal path tracing on the dependency map structure to quantitatively evaluate the possible indirect negative causal effects of the intervention on other key nodes or subgraphs in the map that are not targets.

[0018] Preferably, the step of validating or adjusting the candidate operation and maintenance operations according to the side effect evaluation result specifically includes: comparing the indirect negative causal effect value quantified by the CI model with a preset multi-dimensional security risk threshold; if any threshold is exceeded, preferentially selecting to discard the candidate operation or triggering the RL agent to regenerate a new candidate operation considering the side effect constraint; and using the quantified negative causal effect value as a negative reward signal to feedback to the RL agent to shape its strategy for generating safer operations.

[0019] Preferably, in the step of the reinforcement learning RL agent generating candidate operation and maintenance operations, the action space of the agent is dynamically generated, and it selectively includes the most relevant subset of atomic operation and maintenance operations based on the node types and states of the high-risk subgraphs associated with the current risk indication information.

[0020] Preferably, the graph neural network GNN model is a spatio-temporal graph neural network STGNN or other GNN model variants capable of capturing the evolution pattern of dynamic graph structures.

[0021] Preferably, it further includes a feedback step: feeding back and updating the actual system response and the occurrence of side effects after executing the final operation and maintenance operation to be executed to the dependency graph DMLDG, and using it to continuously optimize at least one of the GNN model, the CI model, and the strategy of the RL agent.

[0022] Preferably, the method further includes: using the prediction uncertainty information output by the GNN model or the weak causal confidence region identified by the CI model to guide the RL agent to plan and execute controlled chaos engineering experiments; performing side effect evaluation using the CI model before execution; and using the experimental observation results to specifically optimize at least one of the dependency graph, the GNN model, and the CI model.

[0023] The technical solution provided by this application has at least the following technical effects or advantages:

[0024] By constructing and analyzing the Dynamic Multi - layer Dependence Graph (DMLDG) and applying Graph Neural Networks (GNN), it is possible to deeply understand the implicit, cross - level dependence relationships within such complex architectures, and accurately predict the emergent fault risks and their propagation paths caused by complex interactions between components, which are difficult to discover by traditional methods. Combining Causal Inference and Reinforcement Learning (CI - RL), a series of context - aware and minimum - side - effect operation and maintenance operations can be automatically generated and verified, realizing precise, safe, and efficient automated avoidance of the predicted risks or rapid recovery of the occurred faults, significantly reducing the mean time to repair and operation and maintenance labor costs. The accurate prediction of GNN provides high - quality input for CI - RL, while the safe decision - making and execution results of CI - RL provide valuable feedback data for the continuous optimization of GNN and CI models, forming a virtuous adaptive cycle of prediction - decision - making - execution - feedback. By introducing model - based introspective chaos engineering, the system can actively learn and adapt to unknown or rare risk scenarios, continuously improving the overall stability and resilience. Description of the Drawings

[0025] Figure 1 It is a flowchart of a distributed cloud - native application computing method for intelligent operation and maintenance of the present invention. Detailed Embodiments

[0026] The present invention relates to a distributed cloud - native application computing method for intelligent operation and maintenance to solve the technical problems in the prior art that the lack of the ability to model emergent faults leads to blindness in decision - making, and unsafe operations may exacerbate the deterioration of the system state.

[0027] The above - mentioned technical solutions will be described in detail below in combination with the drawings in the specification and specific embodiments to better understand the above - mentioned technical solutions. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments of the present invention. It should be understood that the present invention is not limited to the example embodiments only used to explain the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention. In addition, it should be noted that, for the sake of description, only the parts related to the present invention are shown in the drawings rather than all of them. Embodiment 1

[0028] As Figure 1 shown in a flowchart of a distributed cloud - native application computing method for intelligent operation and maintenance, the method includes the following steps:

[0029] Collect and fuse real - time data from multiple heterogeneous data sources within the distributed cloud - native application architecture;

[0030] Based on the fused data, construct and dynamically maintain a dependence graph DMLDG that represents the entities within the application architecture and their dynamic multi - layer dependence relationships;

[0031] Taking the dynamically maintained dependency graph as input, applying the graph neural network GNN model to learn the spatio-temporal features of the graph, and predicting the potential emergent fault risks in the application architecture, and outputting risk indication information;

[0032] Based on the risk indication information, using a reinforcement learning RL agent to generate candidate operation and maintenance operations;

[0033] Before executing the candidate operation and maintenance operations, using a causal inference CI model to evaluate the possible side effects of the candidate operation and maintenance operations on the application architecture;

[0034] According to the side effect evaluation results, verifying or adjusting the candidate operation and maintenance operations to obtain the final operation and maintenance operations to be executed;

[0035] Executing the final operation and maintenance operations to be executed to avoid risks or recover from faults.

[0036] Specifically, this embodiment provides a specific implementation of an intelligent operation and maintenance calculation method, which is applied to manage a typical and complex distributed cloud native application system, and the system architecture features are as follows:

[0037] Front-end application: WeChat Mini Program, which processes user interactions and calls back-end APIs.

[0038] Back-end microservices: Built based on the Go language and running in the form of microservices (such as user services, order services, etc.), communicating through RESTful APIs or gRPC.

[0039] Cache / data storage: Widely use a Redis cluster (for example, deployed through Redis Cluster or Sentinel) as a distributed cache and a possible simple message queue / distributed lock.

[0040] Deployment environment: All back-end services and Redis instances are containerized (Docker) and deployed in a Kubernetes (K8s) cluster.

[0041] Service governance: (Optional but typical) Adopt a service mesh (such as Istio) for traffic management, policy control, and enhanced observability.

[0042] Basic operation and maintenance facilities: Already have Prometheus monitoring, Jaeger / OpenTelemetry tracing, log aggregation (such as EFK / Loki), CI / CD, etc.

[0043] The specific implementation steps of the intelligent operation and maintenance calculation method of the present invention under this architecture are as follows:

[0044] Step 1: Multi-source heterogeneous data collection and fusion: To achieve a comprehensive perception of the target complex architecture, this step requires real-time collection and fusion of data distributed at different levels. Specific data sources and collection methods may include:

[0045] 1. Application layer data:

[0046] Go microservice metrics: By integrating the Prometheus client library (such as prometheus / client_golang) into the Go code, expose the key performance metrics (QPS, latency distribution, error rate), business metrics (such as the number of order creations), and Go runtime metrics (number of goroutines, GC activity) of the service. Regularly scraped by the PrometheusServer.

[0047] Inter-service call information: Use a distributed tracing client (compliant with the OpenTelemetry specification) integrated in the service framework or service mesh (such as Istiosidecar) to generate and report Trace / Span data to Jaeger or a similar tracing system. This data includes service call relationships, elapsed times, and key attributes (such as HTTP method / status code, gRPC method / status code).

[0048] Interaction metrics with Redis: Inject monitoring logic into the client library for the Go service to access Redis (or use a library that supports monitoring) to record the operation types (GET, SET, DEL) on a specific Redis instance / cluster, the target key pattern, hit rate, operation latency, and error messages. Can also be exposed through Prometheus or directly pushed to a time series database.

[0049] 2. Middleware layer data:

[0050] Redis performance metrics: Use RedisExporter (such as oliver006 / redis_exporter) to connect to each Redis instance (Master, Slave, ClusterNode) and collect detailed metrics provided by the INFO command (memory usage, CPU, network I / O, number of connections, key space statistics, slow query log count, cluster status, etc.). Scraped by the PrometheusServer.

[0051] Redis topology: Periodically query the Redis cluster (CLUSTERNODES) or Sentinel (SENTINEL master / slaves) to obtain the current topology and master-slave status information.

[0052] 3. Container layer data:

[0053] Pod Status and Resources: Obtain the running status (Pending, Running, Failed, etc.), restart count, IP address, affiliated Node, labels, etc. of Pods in all relevant namespaces through the Kubernetes API Server. Obtain the actual CPU and memory usage of Pods through the Metrics Server API.

[0054] Container Events: Listen to the event stream of the Kubernetes API Server and capture events such as Pod scheduling, startup, failure, resource pressure eviction, etc.

[0055] 4. Node-level Data:

[0056] Node Status and Resources: Obtain the status of the Node (Ready, NotReady), total resources (CPU, memory), and allocated resources through the Kubernetes API Server. Use Node Exporter deployed on each Node to collect detailed operating system-level metrics (CPU usage breakdown, detailed memory usage, network traffic, disk I / O, etc.). Scraped by the Prometheus Server.

[0057] 5. Service Governance Layer Data (if using Istio):

[0058] Traffic Topology and Metrics: Service-to-service traffic metrics (request volume, latency, success rate, TCP connection information) generated by Istio Mixer (or the new Telemetry API), which can be scraped by Prometheus. Lower-level metrics exposed by the Envoy proxy can also be collected.

[0059] Policy Configuration: Obtain the configuration resources of Istio, such as VirtualService (routing rules), DestinationRule (traffic policies), Gateway, etc. through the Kubernetes API (CRDs).

[0060] 6. Configuration and Event Stream Data:

[0061] Deployment Configuration: Obtain deployment event records (which version of which service is deployed to which environment), rollback events from the CI / CD system (such as GitLab CI, Jenkins). Obtain the configuration (number of replicas, resource requests / limits, environment variables, etc.) of resources such as Deployment / StatefulSet through the Kubernetes API.

[0062] Redis Configuration: Regularly obtain the content of the Redis instance's configuration file or obtain key runtime configurations through the CONFIG GET command.

[0063] Data Fusion: The collected data is sent to a unified data processing platform. This can be a system that supports the storage of multiple data types, such as:

[0064] Time-series data (Metrics): Stored in a Prometheus-compatible time-series database (such as VictoriaMetrics, Thanos, M3DB).

[0065] Trace data (Traces): Stored in Jaeger, Tempo, or OpenSearch / Elasticsearch.

[0066] Log data (Logs): Stored in Loki, OpenSearch / Elasticsearch.

[0067] Event / configuration data: Can be stored in a relational database, a document database, or directly consumed by the processing engine.

[0068] The key is that data from different sources and different levels need to be associated and time-aligned through shared identifiers (such as Pod Name, Service Name, Trace ID, timestamp) to provide a basis for constructing the DMLDG later.

[0069] Step 2: Construction and Maintenance of the Dynamic Multi-level Dependence Graph (DMLDG): By analyzing the shared resource access patterns extracted from the interaction data between application layer services and the performance data of the middleware layer, explicitly identify and model the implicit dependence relationships generated by resource competition as specific types of edges in the graph; and, by associating the data of the container layer, node layer, and service governance layer, establish and dynamically update the cross-layer dependence edges that span the application layer, middleware layer, container layer, node layer, and service governance layer in the graph.

[0070] Use the GNN model to learn the multi-level spatio-temporal embedding representations of the nodes and edges in the dependence graph DMLDG. The embedding representations capture the transfer characteristics of dependence relationships between different levels; and, based on the learned embedding representations, identify and output the critical path subgraphs with a high probability of anomaly as risk indication information. The critical path subgraphs characterize potential cascading failure modes:

[0071] 1. Node Representation: Identify the entities in the system based on the collected data and create graph nodes of the corresponding types. For example: GoServiceInstance: Represents a specific Go service Pod;

[0072] GoService: Represents a Go service;

[0073] RedisInstance: Represents a Redis process;

[0074] RedisKeyPattern: Represents a type of frequently accessed RedisKey;

[0075] K8sPod: Represents a Kubernetes Pod (the same carrier as GoServiceInstance or RedisInstance);

[0076] K8sNode: Represents a Kubernetes worker node;

[0077] IstioVirtualService: Represents an Istio routing rule;

[0078] ConfigParameter: Represents a key configuration item;

[0079] Dynamic update of node attributes: Attach the real - time status and performance metrics of the entity (such as CPU usage, latency, number of connections) as node attributes, which change with data updates.

[0080] 2. Edge representation and dependency modeling:

[0081] 1. Explicit intra - layer relationships:

[0082] Service call: Create CALLS edges between GoServiceInstance nodes according to Trace data, and the attributes can include QPS, average latency, and error rate.

[0083] Cross - layer dependency relationships:

[0084] Instance -> Pod: Create a RUNS_ON edge from GoServiceInstance to the K8sPod on which it runs.

[0085] Pod -> Node: Create a SCHEDULED_ON edge from K8sPod to the K8sNode where it is located according to K8s scheduling information.

[0086] Service -> Redis instance / Key: Create DEPENDS_ON or ACCESSES edges from GoServiceInstance to RedisInstance or RedisKeyPattern according to Go service interaction metrics or configurations.

[0087] Policy -> Service: Create an APPLIES_TO edge from IstioVirtualService to the GoService it manages according to Istio configurations.

[0088] The attributes of the edge can carry information such as interaction metrics (e.g., Redis access latency), configuration status (e.g., routing weight), and resource ownership.

[0089] 2. Implicit Dependencies:

[0090] Identify and model resource competition: Continuously analyze the performance metrics (e.g., reported operation latency) of all GoServiceInstances accessing the same RedisInstance and the performance metrics (CPU, memory, number of connections) of the RedisInstance itself; when it is detected that multiple GoServiceInstances access the same RedisInstance concurrently and the RedisInstance experiences performance bottlenecks (e.g., CPU saturation, generally increased latency), create an edge of type IMPLICIT_CONTENTION_REDIS between these GoServiceInstance nodes (or through the shared RedisInstance node). The weight of the edge can be dynamically adjusted based on factors such as the access volume of the service and the observed degree of performance impact.

[0091] Identify and model policy linkage impacts: Listen for Istio configurations (e.g., changes in connection pooling, circuit breaking, and load balancing policies in DestinationRule); when a policy change affects multiple GoServices simultaneously (e.g., setting global connection pool parameters for the same target service), create an IMPLICIT_POLICY_IMPACT edge between these affected GoService nodes, indicating that they may have performance linkages due to this policy.

[0092] 3. Dynamic Maintenance: Use a stream processing engine (e.g., Flink, SparkStreaming) or an event-based update mechanism to update the node attributes and the existence / attributes of edges in the DMLDG in real time according to newly incoming data. For example, remove the corresponding node and its associated edges when a Pod is destroyed; update the GoServiceInstance node and adjust the relevant edges when a new version is deployed; a database that supports dynamic graphs (e.g., Neo4j combined with stream processing, or a dedicated dynamic graph database) can be used to store and maintain the DMLDG.

[0093] Step 3: Emergent Fault Risk Prediction Based on GNN:

[0094] Selection and Input of the GNN Model: Select the STGNN model, the Spatio-Temporal Graph Neural Network STGNN, or other GNN model variants that can capture the evolution pattern of the dynamic graph structure to simultaneously learn the spatial structure (dependency) of the graph and the temporal dynamics of node / edge features.

[0095] Take the snapshots (node feature vectors, adjacency matrices / edge lists, and edge attributes) of DMLDG at each time step (e.g., every minute) as the input sequence of the GNN. The node feature vectors can include normalized performance metrics, resource utilization rates, status encodings, etc.

[0096] Model training and prediction objectives: Multi-level spatio-temporal embedding learning: The GNN learns the embedding representations of each node in time and space (graph structure) through graph convolutional layers and time convolutional or RNN / Transformer layers. These embeddings can capture how states propagate through dependencies (across layers, implicitly). For example, Node resource tension may affect the performance of the Pods on it, which in turn affects the latency of the service instances that call the Pods. This propagation characteristic will be captured by the embeddings.

[0097] Critical path subgraph risk prediction: Define the critical path: Identify critical user processes (such as "user login", "product browsing", "order submission") according to business logic. Map the service calls and resource dependencies involved in these processes onto the DMLDG as critical path subgraphs.

[0098] Training objective: Train the GNN to predict the probability that the comprehensive anomaly score (which can be a health score based on multiple metrics) of the nodes on the critical path subgraph is lower than the threshold within the next T time (e.g., 5 minutes), or the probability that the overall end-to-end latency of the path exceeds the SLA.

[0099] Output risk indication information: When the prediction probability exceeds the warning threshold, output the critical path subgraph (including nodes and their predicted anomaly scores) that constitutes the high-risk prediction as risk indication information. This directly reflects the potential cascading failure modes that spread along the graph spectrum.

[0100] Step 4: RL-based candidate operation generation:

[0101] 1. RL agent: Adopt an RL algorithm suitable for handling complex states and a potentially large action space, such as PPO (Proximal Policy Optimization); the action space of the agent is dynamically generated, which selectively includes the most relevant subset of atomic operation and maintenance operations based on the node types and states of the high-risk subgraph associated with the current risk indication information.

[0102] 2. State input: Encode the high-risk subgraph information (involved nodes, edges, their current detailed attributes, and predicted risk values) output by the GNN as the state (Observation) of the RL agent.

[0103] 3. Dynamic action space (supporting claim 6):

[0104] Implementation: Dynamically construct the currently available action subset based on the node types in the input state (high-risk subgraph) and the problem characteristics.

[0105] Example: If the risk subgraph indicates that a GoServiceInstance (PodA) has high CPU usage, and it runs on a K8sNode (NodeX), the overall CPU of NodeX is also high, and the RedisInstance (RedisB) that the service depends on responds slowly. The dynamic action space may preferentially include:

[0106] ScaleUpGoService(serviceA)

[0107] MigrateGoPod(PodA) (if K8s supports it and there are available nodes)

[0108] AdjustRedisConfig(RedisB,slowlog-log-slower-than,...)

[0109] RestartGoPod(PodA) (as a last resort)

[0110] It does not include unrelated operations, such as ScaleDownOtherService(serviceY).

[0111] 4. Action generation: The RL policy network outputs one or a series of candidate operations (Action) based on the current state, which may be an action selection with probability distribution or a planned sequence.

[0112] Step 5: Based on the side effect evaluation of CI, the candidate operation is represented as an intervention on the target node or edge in the dependency graph DMLDG. Using the CI model, causal path tracing is performed on the dependency graph structure to quantitatively evaluate the indirect negative causal effects that the intervention may have on other non-target key nodes or subgraphs in the graph, including:

[0113] CI model construction: Use historical operation and maintenance data (operation records, system status changes) and the structural information provided by DMLDG to build a causal model. You can use PC-based algorithms, LiNGAM, or deep learning methods (such as Gumbel-MaxSCM) to learn the causal graph structure, or directly use DMLDG as the basic causal graph framework and estimate the connection strength (causal effect). Domain knowledge (such as knowing that expanding service A will inevitably increase requests to RedisB) can be used as a priori constraints.

[0114] Side Effect Evaluation Process:

[0115] Intervention representation: Represent the candidate operation as an intervention on the corresponding node in the DMLDG.

[0116] Causal path tracing: Starting from the intervention point, propagate along the directed edges on the DMLDG (as a causal graph).

[0117] Quantifying indirect negative effects: The CI model uses the learned causal effect function to estimate the indirect (not the direct operation target) negative effects of the intervention on non-target key nodes (such as the performance metrics of key service C that has nothing to do with service A but shares NodeX, or the latency of other service D that shares RedisB, or the overall CPU usage of NodeX, Y). The output can be the expectation of the impact value, probability distribution, or risk score.

[0118] Step 6: Verification and adjustment based on side effect assessment: Compare the indirect negative causal effect values quantified by the CI model with the preset multi-dimensional security risk thresholds; if any threshold is exceeded, preferentially discard the candidate operation or trigger the RL agent to regenerate a new candidate operation considering this side effect constraint; and, feed the quantified negative causal effect value back to the RL agent as a negative reward signal to shape its strategy for generating safer operations, specifically including:

[0119] 1. Safety risk threshold comparison: Pre-define multi-dimensional safety thresholds, for example:

[0120] Shared Node CPU increment < 10%

[0121] Shared Redis instance P99 latency increment < 5ms

[0122] Probability of impact on other key service (non-target) SLA < 1%

[0123] Compare the quantified side effect values output by the CI model with these thresholds.

[0124] 2. Decision-making and feedback: If the side effect assessment value in any dimension exceeds the threshold, the candidate operation is considered unsafe and is preferentially discarded.

[0125] Trigger regeneration / adjustment: After discarding, notify the RL agent (possibly add the violated constraints to the state of its next-round decision-making) to make it generate a new candidate operation. Or, the system attempts to adjust the original operation parameters (such as reducing the expansion quantity) and then re-evaluate. Only when all side effect assessment values are within the acceptable range, the operation is verified and becomes the final operation to be executed.

[0126] Step 7: Execution and Feedback Loop: Feed back and update the actual system response and side effect occurrences after performing the final operation to be performed on the operation and maintenance to the Dependency Map DMLDG, and use it to continuously optimize at least one of the GNN model, CI model, and RL agent's strategy:

[0127] 1. Execution: Execute the verified operation and maintenance operations by calling the corresponding APIs (Kubernetes API, Redis CLI / API, Istio API).

[0128] 2. Monitoring and Feedback: Continuously monitor the changes in the system state after the operation is executed, and stream the updated data back to data collection and fusion to update the DMLDG in real time.

[0129] Record the actual effects of the operation (Is the problem solved? Have the metrics recovered?) and the actual side effects that occurred (Have any negative impacts predicted by the CI model or other unanticipated negative impacts been observed?).

[0130] 3. Model Optimization: RL Reward and Policy Update: Combine the actual effects of the operation (degree of achievement of the main goal) and the quantified side effects (from CI evaluation or actual observation) to calculate the final reward value, and feedback it to the RL agent to update its policy network, making it tend to generate more effective and less side-effect operations in the future.

[0131] GNN / CI Model Optimization: Use the execution results (success / failure, side effect situation) as new labeled data or experience to periodically retrain or fine-tune the GNN prediction model and CI causal model to improve their accuracy.

[0132] Step 8, optionally, Model-based Self-introspective Chaos Engineering: Use the predicted uncertainty information output by the GNN model or the weak causal confidence region identified by the CI model to guide the RL agent to plan and execute controlled chaos engineering experiments; use the CI model to evaluate side effects before execution; and use the experimental observation results to specifically optimize at least one of the dependency map, GNN model, and CI model:

[0133] 1. Triggering and Planning: When the GNN model has a high uncertainty in its risk prediction for a certain subgraph (such as the output probability is close to 0.5 or the variance is large), or the CI model identifies a potential strong causal association between two components but the historical data is insufficient resulting in low confidence, the system identifies this as a "knowledge blind spot" that needs to be actively explored.

[0134] The RL agent (or a dedicated chaos engineering module) is triggered, and based on this information, it plans a small-scale, targeted chaos experiment. For example, inject a small perturbation (such as a brief increase in network latency, simulate a brief Redis key failure) on the identified low-confidence causal link.

[0135] 2. Safety assessment and execution: Input the planned chaos experiment (regarded as a special operation) into the CI model in step 5 for rigorous side effect assessment.

[0136] Only when the experimental risk is assessed to be controllable (the side effects are within a very small range), is the experiment allowed to be performed within a predetermined time window (such as a business off-peak period).

[0137] 3. Learning and optimization: Accurately observe the actual response of the system during the experiment (which indicators have changed? How much has changed? Is it in line with expectations?).

[0138] Use these observations as high-quality annotated data:

[0139] Used to verify or correct the existence or weight of the corresponding edge in DMLDG.

[0140] Used to fine-tune the GNN model to improve its prediction accuracy in this scenario.

[0141] Used to enhance the CI model's understanding and confidence in the causal relationship of this link.

[0142] This will enable model-driven, safe, controllable, and targeted system resilience enhancement and model self-improvement.

[0143] This specific implementation describes in detail how to apply the intelligent operation and maintenance computing method of the present invention to a typical Go+Redis+K8s+Mesh+WeChat applet complex distributed system. By combining multiple layers of heterogeneous data to build a dynamic dependency graph (DMLDG), using GNN for deep emergent risk prediction, and innovatively integrating causal inference (CI) to perform pre-side effect evaluation and safety verification on the operation and maintenance operations generated by RL, a closed-loop system is finally formed that can make accurate, safe, and adaptive operation and maintenance decisions and executions, and has (optional) intelligent chaos engineering capabilities. This implementation fully demonstrates the feasibility and advancement of the technical solution of the present invention and its significant advantages in solving the operation and maintenance problems of complex cloud native systems.

[0144] It should be understood that the embodiments disclosed in the present invention and the above description can enable those skilled in the art to use the present invention to implement the present invention. At the same time, the present invention is not limited to the above-mentioned embodiments. It should be understood that those skilled in the art can still modify the technical solutions recorded in the above-mentioned embodiments, or replace some of the technical features therein by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention.

Claims

1. A distributed cloud-native application computing method for intelligent operation and maintenance, characterized in that, It includes the following steps: Collect and fuse the real-time data of multiple heterogeneous data sources within the distributed cloud-native application architecture; Based on the fused data, construct and dynamically maintain a dependency map DMLDG that represents the entities within the application architecture and their dynamic multi-layer dependency relationships; Taking the dynamically maintained dependency map as input, apply a graph neural network GNN model to learn the spatio-temporal features of the map and predict the potential emergent failure risks in the application architecture, and output risk indication information; Based on the risk indication information, use a reinforcement learning RL agent to generate candidate operation and maintenance operations; Before executing the candidate operation and maintenance operations, use a causal inference CI model to evaluate the possible side effects of the candidate operation and maintenance operations on the application architecture; According to the side effect evaluation results, verify or adjust the candidate operation and maintenance operations to obtain the final operation and maintenance operations to be executed; Execute the final operation and maintenance operations to be executed to avoid risks or recover from faults.

2. The distributed cloud native application computing method for intelligent operation and maintenance according to claim 1, wherein, The step of constructing and dynamically maintaining the dependency map DMLDG specifically includes: by analyzing the shared resource access patterns extracted from the interaction data between application layer services and middleware layer performance data, explicitly identifying and modeling the implicit dependency relationships generated by resource competition as specific types of edges in the map; and by associating the data of the container layer, node layer, and service governance layer, establishing and dynamically updating the cross-layer dependency edges that span the application layer, middleware layer, container layer, node layer, and service governance layer in the map.

3. The distributed cloud native application computing method for intelligent operation and maintenance according to claim 1, characterized in that, The step of applying the graph neural network GNN model to predict emergent failure risks specifically includes: using the GNN model to learn the multi-level spatio-temporal embedding representations of the nodes and edges in the dependency map DMLDG, and the embedding representations capture the transfer characteristics of the dependency relationships between different levels; and based on the learned embedding representations, identifying and outputting the critical path subgraphs with high anomaly probabilities as the risk indication information, and the critical path subgraphs represent potential cascading failure modes.

4. The distributed cloud native application computing method for intelligent operation and maintenance according to claim 1, characterized in that, The step of using the causal inference CI model to evaluate side effects specifically includes: representing the candidate operation and maintenance operations as interventions on the target nodes or edges in the dependency map DMLDG; using the CI model to perform causal path tracing on the dependency map structure and quantitatively evaluating the possible indirect negative causal effects of the intervention on other critical nodes or subgraphs in the map that are not the target.

5. A distributed cloud-native application computing method for intelligent operation and maintenance according to claim 1, characterized in that, The step of verifying or adjusting the candidate operation and maintenance operations according to the side effect evaluation results specifically includes: comparing the indirect negative causal effect values quantitatively evaluated by the CI model with the preset multi-dimensional security risk thresholds; if any threshold is exceeded, preferentially choose to discard the candidate operation or trigger the RL agent to regenerate new candidate operations considering the side effect constraints; and using the quantitatively evaluated negative causal effect values as negative reward signals to feedback to the RL agent to shape its strategy of generating safer operations.

6. The distributed cloud native application computing method for intelligent operation and maintenance according to claim 1, wherein In the steps where the reinforcement learning (RL) agent generates candidate operation and maintenance operations, the action space of the agent is dynamically generated, which selectively includes the most relevant subset of atomic operation and maintenance operations based on the node types and states of the high-risk subgraph associated with the current risk indication information.

7. The distributed cloud native application computing method for intelligent operation and maintenance according to claim 1, characterized in that The graph neural network (GNN) model is a spatio-temporal graph neural network (STGNN) or other variants of the GNN model that can capture the evolution pattern of the dynamic graph structure.

8. A distributed cloud-native application computing method for intelligent operation and maintenance according to claim 1, characterized in that, It further includes a feedback step: feeding back and updating the actual system response and the occurrence of side effects after executing the final operation and maintenance operation to be executed to the dependency map DMLDG, and using it to continuously optimize at least one of the GNN model, the CI model, and the policy of the RL agent.

9. The distributed cloud-native application computing method for intelligent operation and maintenance according to claim 1, wherein, The method further includes: using the prediction uncertainty information output by the GNN model or the weak causal confidence region identified by the CI model to guide the RL agent to plan and execute controlled chaos engineering experiments; using the CI model to evaluate side effects before execution; and using the experimental observation results to specifically optimize at least one of the dependency map, the GNN model, and the CI model.

Citation Information

Patent Citations

  • Private domain traffic scheduling and content distribution method and system based on group behavior analysis

    CN119377498A

  • Intelligent dependency graph construction method based on multi-dimensional analysis and adaptive optimization

    CN119474227A

  • Methods and systems for autonomous cloud application operations

    US20200371857A1

Cited By

  • Computing power resource optimization method and system based on deep learning

    CN120560859A

  • Interface service state analysis system and method based on distributed testing and monitoring

    CN121210248A

  • Interface service state analysis system and method based on distributed testing and monitoring

    CN121210248B

  • Software performance analysis method and system in virtualization environment with cross-layer data association

    CN121349825A

  • Dynamic fault perception and risk propagation prediction method for agent workflow

    CN121637492A