Big data analysis and automatic decision scheme generation method for operation and maintenance risk of cloud platform

By constructing probabilistic invariant models and causal graphs, the problem of insufficient causal transmission models in cloud platform operation and maintenance is solved, enabling a deep understanding and forward-looking prediction of system behavior, and improving the accuracy of fault root cause localization and the effectiveness of decision response.

CN120973580APending Publication Date: 2025-11-18BEIJING YINTAI TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511111668.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing cloud platform operation and maintenance technologies struggle to establish causal transmission models within the system, leading to inaccurate root cause localization of faults and delayed decision-making responses, as well as a lack of deep understanding of system behavior and forward-looking predictive capabilities.

Method used

Construct a probabilistic invariant model and a causal graph. By converting multimodal operation and maintenance data into a structured behavioral event flow, establish the logical boundary of system health behavior, and use the causal graph for risk inference and decision generation. Dynamically update the model to adapt to system changes.

Benefits of technology

It improves the accuracy of risk identification and the effectiveness of decision-making solutions, ensuring that decision-making solutions are causally oriented and forward-looking, and enhancing the success rate and efficiency of automated intervention in operation and maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120973580A_ABST
    Figure CN120973580A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of cloud computing, and discloses a cloud platform operation and maintenance risk big data analysis and automatic decision scheme generation method comprising the following steps: obtaining multi-modal operation and maintenance data, and converting the multi-modal operation and maintenance data into a structured behavior event stream; constructing a probabilistic invariant model based on the structured behavior event stream; constructing a causal atlas for representing the causal conduction relationship among the internal factors of the system; the causal atlas is corrected by comparing the deduction result of the causal atlas with the probabilistic invariant model, so that the accuracy of the causal atlas is improved; deducing system state evolution caused by a real-time behavior event by using the corrected causal atlas, and identifying a risk trajectory; and generating an automatic decision scheme according to the risk track. According to the method, conversion from passive response to active prediction can be realized, and the intelligent level of large-scale cloud platform operation and maintenance and the system reliability are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of cloud computing technology, specifically to a method for generating big data analysis and automated decision-making solutions for cloud platform operation and maintenance risks. Background Technology

[0002] With the rapid development and widespread application of cloud computing technology, the scale and complexity of cloud platforms are growing at an unprecedented rate. While cloud-native technologies, represented by containerization and microservice architecture, bring agility and elasticity, they also lead to a surge in the number of internal system components, intricate dependencies between services, and rapidly changing system states. This highly dynamic and distributed nature presents severe challenges to the operations and maintenance of cloud platforms.

[0003] Existing cloud platform operation and maintenance technologies largely rely on the analysis of massive amounts of monitoring data. However, these methods often have inherent limitations. Traditional monitoring and alarm systems are typically based on static thresholds, making it difficult to adapt to dynamic changes in business load. This often leads to alarm storms or missed critical anomalies, leaving operations personnel overwhelmed by massive amounts of low-value information. When a failure actually occurs, operations personnel usually need to rely on personal experience to locate the root cause by checking logs, metrics, and tracing information. This process is not only time-consuming and labor-intensive but also struggles to handle complex failures triggered by multiple cascading factors.

[0004] More importantly, existing technologies generally lack a deep understanding of system behavior and the ability to predict it proactively. They struggle to reveal the deep causal chains between system configuration, resource status, and application performance. Therefore, when faced with proactive operations such as version releases and configuration changes, operations teams cannot accurately predict the potential risks these operations may trigger, often only able to respond and remedy passively after problems occur, resulting in a persistently high risk of service interruptions. Simultaneously, when formulating fault recovery strategies, there is a lack of a systematic method to evaluate the potential effects and costs of different intervention measures; the decision-making process often relies on intuition rather than data-driven deduction, making it difficult to guarantee the optimality of the solution. Therefore, there is an urgent need for an intelligent operations technology capable of proactively predicting risks and automatically generating optimal decision-making solutions to address the challenges of the complexity of modern cloud platforms. Summary of the Invention

[0005] The technical problem that this invention aims to solve is that existing technologies in cloud platform operation and maintenance rely heavily on correlation analysis between performance indicators, making it difficult to establish a causal transmission model within the system, resulting in inaccurate root cause location of faults and delayed decision-making response.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: The first aspect of this invention provides a method for analyzing and generating decision-making solutions for cloud platform operation and maintenance risks, the method comprising the following steps: Acquire multimodal operation and maintenance data and convert it into a structured stream of behavioral events; Based on the structured behavioral event flow, a probabilistic invariant model is constructed, which is used to characterize the probabilistic rules that the system's healthy behavior should follow in a specific context. Construct a causal graph to characterize the causal transmission relationships between factors within the system; The causal graph is corrected by comparing the inference results of the causal graph with the probabilistic invariant model, thereby improving the accuracy of the causal graph. Using the modified causal graph, the evolution of system state triggered by real-time behavioral events is deduced, and risk trajectories are identified. Based on the risk trajectory, an automated decision-making scheme is generated.

[0007] Through the above methods, this invention defines the healthy behavior boundary of a system by constructing a probabilistic invariant model, and uses this model to continuously verify and correct the causal graph, ensuring that the graph accurately reflects the causal transmission logic of the system. Based on this accurate graph, risk inference and decision generation are performed, thereby improving the accuracy of risk identification and the effectiveness of decision-making schemes.

[0008] In one specific embodiment, the step of constructing the probabilistic invariant model includes: Define the invariant as a logical assertion, and formally define the logical assertion as follows: ; in, Represents an invariant. It represents a relational expression between one or more attributes within the system. Represents the current context vector of the system. In context Relationship The conditional probability of it being true. This is the preset confidence threshold.

[0009] This step also includes, in the specific context When changes occur, the probabilistic rules in the probabilistic invariant model related to the context are dynamically updated. Specifically, the conditional probabilities of the relevant rules are recalculated, and the rules are retained or eliminated based on whether they meet the confidence threshold.

[0010] Preferably, the step of correcting the causal map includes: The system state is simulated and deduced using the causal graph to obtain the predicted state; The consistency between the predicted state and the probabilistic invariant model is evaluated, and this consistency is quantified by calculating the combined probability that the predicted state satisfies all relevant rules in the probabilistic invariant model; Based on the consistency evaluation results, the parameters of the causal graph are adjusted to minimize the logical conflict between the inference results of the causal graph and the probabilistic invariant model.

[0011] In one specific embodiment, the step of adjusting the parameters of the causal graph is performed using a gradient-based optimization method. Specifically, a logical conflict function is defined. To quantify the spectrum With invariant model Inconsistencies between them: ; in, This is the current state. The Simulate function is a graph derivation function for simulating action sequences. To predict the validity probability of a state under an invariant model, the edge weights in the causal graph are iteratively adjusted using the following formula. : ; in, This is the learning rate.

[0012] Preferably, the step of identifying risk trajectories includes: Real-time behavioral events are used as initial perturbations input into the causal graph. By performing probabilistic propagation on the causal graph, a probability distribution sequence of future states is deduced, forming a probabilistic trajectory; Define one or more key invariants. When the probability of the probabilistic trajectory leading to a state that violates any of the key invariants exceeds a preset risk threshold, the trajectory is identified as a risk trajectory.

[0013] In one specific embodiment, the step of generating an automated decision-making scheme first includes, before generating the decision-making scheme: On the causal graph, the invariant node that is about to be violated in the risk trajectory is traced backward, and all causal paths pointing to the node are traversed to filter out candidate operations that have a causal impact on the invariant node, forming a set of candidate operations.

[0014] In one specific embodiment, the step of generating an automated decision-making scheme further includes: Within the decision space comprised of the candidate operation set, a constrained optimization problem is solved through counterfactual decision deduction. This problem is formalized as: Searching for: ; The constraints are: ; in, This is the optimal operation sequence. Let be the set of candidate operation sequences. To execute the operation sequence The cost function, In the current state Execution sequence The probability that the system will subsequently enter any risk state. This is a preset safety threshold.

[0015] In one specific embodiment, the nodes of the causal graph include: Configuration nodes are used to represent configurable parameters, resource status nodes are used to represent system performance indicators or resource status, and invariant nodes represent each probabilistic rule in the probabilistic invariant model.

[0016] Preferably, the step of converting multimodal operation and maintenance data into a structured behavioral event stream involves uniformly converting configuration data, observation data, and operation data into a structured event sequence that includes declaration events, execution events, invocation events, and status attribute events.

[0017] A second aspect of this invention provides a system for analyzing and generating decision-making solutions for cloud platform operation and maintenance risks, the system comprising: The data transformation module is used to acquire multimodal operation and maintenance data and convert it into a structured stream of behavioral events. The model building module is used to construct a probabilistic invariant model based on the structured behavioral event flow to characterize the probabilistic rules that the system's healthy behavior should follow in a specific context. The graph construction and correction module is used to construct a causal graph to characterize the causal transmission relationship between factors within the system, and to correct the causal graph by comparing the inference results of the causal graph with the probabilistic invariant model, so as to improve the accuracy of the causal graph. The risk simulation module is used to use the modified causal graph to simulate the evolution of the system state caused by real-time behavioral events and identify risk trajectories. The decision generation module is used to generate automated decision-making schemes based on the risk trajectory.

[0018] This invention provides a method for big data analysis and automated decision-making generation of cloud platform operation and maintenance risks. It has the following beneficial effects: 1. This invention constructs a probabilistic invariant model to objectively characterize the logical boundaries of a system's healthy behavior, and uses this model to continuously revise the causal graph representing the transmission relationships of internal factors. This mechanism, which combines and verifies historical data statistical patterns with causal structure inference, ensures that the causal graph accurately reflects the actual causal transmission logic of the system, thereby avoiding erroneous inferences that may result from relying solely on correlation analysis and improving the accuracy of identifying potential risk trajectories.

[0019] 2. The probabilistic invariant model of this invention has dynamic update capabilities, enabling it to automatically recalibrate, eliminate, or generate probabilistic rules in the model when the system context (e.g., application version, deployment environment) changes. Because this model directly participates in the correction process of the causal graph, the entire analysis and decision-making model can adaptively adjust with system evolution, maintaining its analytical effectiveness in a continuously changing cloud platform environment.

[0020] 3. When generating decision-making schemes, this invention first uses reverse tracing of causal graphs to screen candidate operations that have a direct causal impact on risk, narrowing the decision search space and avoiding ineffective attempts at irrelevant operations. Secondly, through counterfactual decision deduction, it seeks the optimal sequence of operations that satisfies safety constraints and has the best operational cost among the screened candidate operations. This approach ensures that the final generated decision-making scheme not only has a clear causal orientation but also possesses foresight, improving the success rate and efficiency of automated intervention. Attached Figure Description

[0021] Figure 1 This is a structural block diagram of a cloud platform operation and maintenance risk analysis and decision-making scheme generation system according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating a method for analyzing and generating decision-making schemes for cloud platform operation and maintenance risks, according to an embodiment of the present invention.

[0022] The module consists of: 10. Data conversion module; 20. Model building module; 30. Map building and correction module; 40. Risk simulation module; and 50. Decision generation module. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0024] See attached document Figure 1 , Figure 1 This is a structural block diagram of a cloud platform operation and maintenance risk analysis and decision-making scheme generation system according to an embodiment of the present invention. The present invention provides a cloud platform operation and maintenance risk analysis and decision-making scheme generation system, which may include: a data conversion module 10, a model building module 20, a graph building and correction module 30, a risk inference module 40, and a decision generation module 50.

[0025] The functionality of these modules can be achieved by executing corresponding computer program instructions on processors running on one or more servers.

[0026] The data conversion module 10 is used to acquire multimodal operation and maintenance data from various sources and in various formats from the cloud platform, and convert the multimodal operation and maintenance data into a structured behavioral event stream.

[0027] In a specific implementation, multimodal operation and maintenance data includes configuration data, observation data, and operational data. A structured behavioral event stream is a sequence of events containing timestamps, event types, and data payloads. Event types include declaration events, execution events, invocation events, and status attribute events.

[0028] The output of the data conversion module 10 is connected to the input of the model building module 20 and the map building and correction module 30.

[0029] Model building module 20 receives a structured stream of behavioral events. Based on the received event stream, this module constructs a probabilistic invariant model. The probabilistic invariant model is used to characterize the probabilistic rules that the healthy behavior of the system should follow in a specific context.

[0030] In one specific implementation, the module defines the invariant as a logical assertion and has the ability to dynamically update the probabilistic rules in the model that are related to the system context when the system context changes.

[0031] The output of the model building module 20 is connected to the input of the graph building and correction module 30 to provide the completed probabilistic invariant model.

[0032] The graph construction and correction module 30 is used to construct causal graphs to characterize the causal transmission relationships between factors within the system.

[0033] In one specific implementation, the nodes of the causal graph include: configuration nodes, resource state nodes, and invariant nodes representing the probabilistic rules in the probabilistic invariant model. Simultaneously, this module receives the probabilistic invariant model provided by the model building module 20, and corrects the causal graph by comparing the inference results of the causal graph with the probabilistic invariant model to improve its accuracy. This correction process includes performing simulations using the graph, evaluating the consistency between the simulation results and the model, and adjusting the graph parameters based on the evaluation results. The output of this module is connected to the inputs of the risk inference module 40 and the decision generation module 50.

[0034] The risk projection module 40 receives a modified causal graph. Upon receiving a real-time behavioral event, this module uses the modified causal graph to project the system state evolution triggered by the event and identify risk trajectories. This projection process is accomplished through probabilistic propagation on the graph, and the existence of a risk trajectory is ultimately determined by comparing the projection results with a preset risk threshold. The output of the risk projection module 40 is connected to the input of the decision generation module 50 to provide information on the identified risk trajectories.

[0035] The decision generation module 50 is used to generate automated decision-making schemes based on the risk trajectories identified by the risk simulation module 40 and the corrected causal graphs provided by the graph construction and correction module 30. Its input is connected to the output of the risk simulation module 40 and the graph construction and correction module 30.

[0036] In one specific implementation, the module first performs reverse tracing on the causal graph to filter out a set of candidate operations. Then, through counterfactual decision deduction, it solves for an optimal sequence of operations with operation cost as the optimization objective and the probability of the system entering a risky state in the future being lower than a preset safety threshold as the constraint. This sequence serves as the final automated decision-making scheme.

[0037] See attached document Figure 2 , Figure 2 This is a flowchart of a method for analyzing and generating decision-making solutions for cloud platform operation and maintenance risks according to an embodiment of the present invention. The method provided in this embodiment can be executed by the aforementioned system embodiment. The main process of this method may include the following steps: S100: Acquire multimodal operation and maintenance data and convert it into a structured behavioral event stream.

[0038] S200: Based on structured behavioral event flows, construct a probabilistic invariant model.

[0039] S300. Construct a causal graph to characterize the causal transmission relationship between factors within the system.

[0040] S400. By comparing the derivation results of the causal graph with the probabilistic invariant model, the causal graph is corrected to improve its accuracy.

[0041] S500 uses a modified causal graph to extrapolate the evolution of system states triggered by real-time behavioral events and identify risk trajectories.

[0042] S600 generates automated decision-making solutions based on risk trajectories.

[0043] The steps in this method will be explained in detail below.

[0044] In step S100, the system performs data acquisition and structured event stream transformation. The purpose of this step is to process multimodal operation and maintenance data collected from various parts of the cloud platform, which come from different sources and have different formats, into a structured behavioral event stream with a unified format required for subsequent analysis steps.

[0045] First, the system obtains multimodal operation and maintenance data from the cloud platform environment. In practice, this data mainly includes three categories: The first category is configuration data, which describes the static deployment information and expected state of the system, such as deployment description files (e.g., YAML files), service configurations, network policies, or resource quota definitions.

[0046] The second category is observation data, which reflects the dynamic behavior and state of the system during operation. Examples include numerical performance indicators (such as CPU utilization, memory consumption, and network throughput) collected by monitoring systems, text logs generated by services or components, and call chain data that records the call relationships and time consumption between services in a distributed system.

[0047] The third category is operational data, which records actions performed by users or automated systems that are intended to change the state of the system, such as API call records, command line execution history, or operation orders in the change management system.

[0048] The system then parses and transforms the acquired multimodal operation and maintenance data to generate a structured behavioral event stream. A structured behavioral event is defined as a data unit containing at least three fields: A timestamp is used to record the precise time when an event occurred. Event type, used to identify the semantic category of an event; A data payload is a collection of key-value pairs used to store specific attribute information related to an event.

[0049] This transformation process maps data from different sources to a unified event type. In a specific implementation, event types include: declarative events, execution events, invocation events, and state attribute events.

[0050] For example, when a change in configuration data is detected, such as when a new deployment file is applied, the system generates a "declaration event" whose data payload contains configuration information such as the application name, version number, and number of replicas.

[0051] When a specific execution action is parsed from the operation data, such as when a user executes an expansion command, the system will generate an "execution event" whose data payload records the operation subject, target object, and execution parameters.

[0052] When a service request is parsed from the call chain data, the system generates a "call event" whose data payload includes the source service, the target service, the request time, and the return status code.

[0053] When a snapshot of a system's state attribute is obtained from performance metrics or structured logs, such as CPU utilization at a certain moment, the system generates a "state attribute event," whose data payload records the name and value of that attribute.

[0054] Through this step, all heterogeneous raw operation and maintenance data are uniformly transformed into a structured event sequence ordered by timestamps, containing declaration events, execution events, invocation events, and status attribute events. This event sequence provides standardized data input for constructing the probabilistic invariant model in subsequent step S200 and constructing the causal graph in step S300.

[0055] In step S200, the system constructs a probabilistic invariant model based on the structured behavioral event stream generated in step S100. This model consists of a set of probabilistic invariants, and its overall function is to provide a quantifiable, probability-based set of rules for the healthy behavior of the system under different operating contexts.

[0056] In one specific embodiment, each probabilistic invariant is defined as a logical assertion. This logical assertion specifies the conditional probability that a relationship between one or more attributes within the system holds true in a particular context, and this probability must satisfy a preset confidence threshold. The logical assertion is defined by the following formal expression: ; In this expression, the symbols are defined as follows: It represents an independent probabilistic invariant.

[0057] This represents a relational expression. This expression is a logical or mathematical expression that can be judged as true or false, and its variables are derived from the data payload of events in the structured behavioral event stream.

[0058] For example, It can be a threshold judgment for a single state attribute, such as "StateAttr.CPU_Usage < 0.8"; it can also be a correlation between multiple attributes, such as "Declare.Replicas == StateAttr.RunningPods"; or it can be an event aggregation statistics within a time window, such as "within any one-minute time window, the number of events of type call event with the Status field of its data payload being Failed is 0".

[0059] This represents a context vector. This vector consists of one or more attributes extracted from the event data payload that identify the current environment or state of the system, such as the application name, deployment environment identifier (e.g., production, testing), or software version number. Context Vector Its function is to limit the relational expression. The specific context in which it was established.

[0060] This represents the conditional probability. It indicates the probability given the system's context vector. In the specific scenario described, relational expressions The probability of its validity. This probability value is obtained through statistical analysis of the behavioral event stream under historical health conditions.

[0061] This represents a preset confidence threshold. It's a constant close to 1, such as 0.999. It serves as a decision boundary for determining whether a probabilistic rule can be considered an invariant. Only when the calculated conditional probability... Greater than or equal to the threshold Only when the corresponding rule is confirmed as a valid probabilistic invariant and incorporated into the probabilistic invariant model can it be included.

[0062] By mining all probabilistic invariants that satisfy the above formal definition The system can construct a complete probabilistic invariant model. Specifically, the model learning and mining steps include: Starting with the generation of candidate invariants, the system uses a set of predefined relation templates to systematically explore all possible relations.

[0063] In a specific implementation, these templates include: Single attribute value constraint template (e.g., comparing CPU_Usage in a state attribute event with a constant to form candidate relationships where CPU_Usage < 0.8); Multi-attribute association templates (e.g., comparing Replicas in the declaration event with RunningPods in the status attribute event to form a candidate relationship of Replicas == RunningPods); Time window aggregation template (e.g., counting the number of specific types of events occurring within a one-minute time window to form a candidate relationship of COUNT(EventType='Failed') == 0).

[0064] The system applies these relation templates to various attributes in the behavior event stream and combines them with preset context vectors (such as application name and version number) to generate a large number of candidate invariants.

[0065] For each generated candidate invariant, the system performs statistical validation to calculate its conditional probability. .

[0066] Specifically, the system first considers the context vector of the candidate invariants. The relevant subset of events is selected from the historical health behavior event stream.

[0067] Then, within this subset of events, the system statistical relational expression The total number of times that can be evaluated is counted as .

[0068] Meanwhile, the system statistics show that in these evaluations, the relational expressions... The number of times the judgment result is true is counted as . The conditional probability of this candidate invariant is calculated using the following formula: ; Finally, the system will calculate the conditional probability. Compared with the preset confidence threshold The comparison is performed. If the calculated probability value is greater than or equal to the confidence threshold, the candidate invariant is confirmed as a valid probabilistic invariant and formally added to the probabilistic invariant model. Conversely, if the probability value is less than the threshold, the candidate invariant is discarded.

[0069] By repeating the above statistical verification and screening process on all candidate invariants, the system finally constructs a model consisting of a series of verified and valid probabilistic invariants, which provides a benchmark for the healthy behavior of the system for subsequent steps.

[0070] To ensure that the probabilistic invariant model accurately reflects the behavior of the continuously evolving cloud platform system, the model in this embodiment of the invention possesses a dynamic evolution mechanism. This mechanism enables the model to be dynamically updated to adapt to changes in system behavior patterns caused by factors such as version iterations, configuration changes, or environment migrations.

[0071] The dynamic evolution mechanism is triggered by a change in the system's specific context. For example, when the system receives a declaration event containing a new software version number or an operation event indicating a change in the deployment environment through step S100, the system determines that the context has changed and initiates the evolution process. In its specific implementation, this dynamic update process includes the following operations: recalibration, elimination, and regeneration.

[0072] Recalibration is performed on probabilistic invariants that already existed in the old context. For each old invariant associated with the change, the system uses the behavioral event flow corresponding to the new context to re-execute the aforementioned statistical validation process and calculate its conditional probability in the new context. If the newly calculated probability value still meets the preset confidence threshold, the system updates the probability value of that invariant stored in the model.

[0073] Elimination is a result of recalibration. During recalibration, if the conditional probability calculated for an old invariant in the new context no longer meets the preset confidence threshold, it indicates that the rule has failed under the new system behavior pattern. The system then removes this rule from the probabilistic invariant model for the new context.

[0074] "Emergence" refers to the discovery of entirely new invariants in a new context. After triggering the evolution mechanism, the system not only recalibrates the old rules but also completely re-executes the aforementioned model learning and mining process for all behavioral event streams collected in the new context. This process can discover and establish entirely new probabilistic invariants that do not exist or do not meet the threshold conditions in the old context but are stable in the new context, and add them to the model.

[0075] Through a series of operations including recalibration, elimination, and regeneration, the probabilistic invariant model completes adaptive dynamic updates to specific contextual changes, ensuring its continued accuracy in describing the healthy behavior of the system.

[0076] In step S300, the system constructs a causal graph to characterize the causal transmission relationships between factors within the system. This graph is formally defined as a directed weighted graph. ,in It is a set of nodes. It is a set of edges.

[0077] Node set Used to represent various key factors in a system, in a specific implementation, a set of nodes. It includes the following three types of nodes: configuration nodes, resource status nodes, and invariant nodes.

[0078] A configuration node represents a parameter that can be set by the user or the system. Its value is usually derived from the declared event obtained in step S100, such as the number of replicas of a service, the number of CPU requests, or the memory limit value.

[0079] Resource status nodes represent an observable system runtime metric or state, and their values ​​are usually derived from status attribute events or call events, such as actual CPU utilization, average service response latency, or network error rate.

[0080] An invariant node is a node in the graph that directly maps each probabilistic rule in the probabilistic invariant model constructed in step S200. The state of the node (e.g., satisfied or violated) depends on the state of the other nodes it is associated with.

[0081] set of edges Used to represent direct causal relationships between nodes, a node in the graph is a branch node. Pointing to node directed edges , representing a node The factors represented are those that lead to the node A direct cause of the change in the factors it represents. Each edge Each is associated with a weight This weight is used to quantify the strength and direction of this causal effect. For example, a positive weight represents a positive correlation, while a negative weight represents a negative correlation.

[0082] In the initial construction phase of the graph, the system needs to assign initial weights to these edges. The initial weights can be estimated based on prior domain knowledge or by analyzing historical data.

[0083] For example, known dependencies in the system architecture, such as increasing the number of replicas of a service (configuration nodes) leading to an increase in total memory consumption (resource status nodes), can be pre-defined as an edge with a positive initial weight. Furthermore, a preliminary causal discovery algorithm can be applied to the historical behavioral event stream generated in step S100 to estimate the temporal dependencies and correlation strengths between variables, serving as initial weights. This initially constructed graph provides the foundation for the self-consistency correction in subsequent step S400.

[0084] In step S400, the system corrects the causal graph initially constructed in step S300. The purpose of this step is to use a closed-loop feedback mechanism, employing the probabilistic invariant model established in step S200 as a benchmark, to verify and optimize the structural parameters of the causal graph, thereby improving its accuracy. This correction process includes simulation, consistency assessment, and parameter adjustment.

[0085] First, the system performs a simulation, selecting a system state from the historical stream of behavioral events. and one or more subsequent actions As input, utilizing the current causal graph. Simulate and extrapolate the system state, that is, calculate the actions based on the node relationships and edge weights defined in the graph. State The influence of each node in the system is analyzed to obtain a predicted state. This process can be viewed as a positive simulation process, the purpose of which is to test the predictive ability of the causal graph for the system's behavior.

[0086] Next, the system performs a consistency evaluation on the predicted state, comparing the predicted state obtained in the previous step with the probabilistic invariant model. A comparison is then performed. Specifically, the extent to which the system evaluates and predicts the state satisfies the probabilistic invariant model is determined. All relevant rules defined in the model. This degree of consistency is quantified as a probability value, representing the probability that the predicted state is deemed "valid" according to the invariant model.

[0087] The system further transforms the results of the above consistency assessment into a quantifiable logical conflict value. In a specific implementation, this logical conflict is defined as a logical conflict function. Its goal is to quantify causal maps. The deduction results and the probabilistic invariant model The inconsistencies between them. The formal expression of this function is as follows: ; in: This indicates that the expected value is calculated for multiple state and action pairs sampled from historical data; It uses causal graphs From state Start simulating the execution of action sequences The predicted state obtained afterwards; Does the predicted state satisfy a probabilistic invariant model? The combined probability of all applicable rules. Logical conflict function. The larger the value, the greater the deviation between the inference results of the causal graph and the health behaviors defined by the invariant model.

[0088] Finally, the system adjusts the parameters of the causal graph based on the consistency evaluation results. The goal of this process is to minimize the quantification of logical conflicts.

[0089] In one specific implementation, this parameter adjustment is performed using a gradient-based optimization method. The system computes a logical conflict function. Regarding the weight of each edge in the causal graph The gradient is calculated, and the weights are updated in the opposite direction of the gradient. This update process is performed using the following formula: ; in: It is a causal graph from the node To the node The weight of the edge; It is a preset learning rate used to control the step size of each update; It is a logical conflict function for weights The partial derivative of the weight indicates the direction for adjusting the weight to reduce logical conflicts.

[0090] By repeatedly executing the closed-loop process of simulation, consistency assessment, conflict quantification, and parameter adjustment, the edge weights of the causal graph are continuously optimized, enabling the causal structure of the graph to more accurately reproduce the system health behavior patterns defined by the probabilistic invariant model.

[0091] In step S500, the system uses the causal graph corrected in step S400 to deduce the system state evolution caused by real-time behavioral events.

[0092] When the system receives a real-time behavioral event, such as a change in a configuration parameter or a sudden change in a key performance indicator, this event is first used as an initial perturbation input to the causal graph. Specifically, the system parses the data payload of the event and identifies the node in the graph that directly corresponds to the event. For example, a configuration change event that updates the number of service replicas will be mapped to a change in the value of the corresponding "configuration node" in the graph; a status attribute event indicating a sharp increase in service response latency will be mapped to a change in the value of the corresponding "resource status node". This initial value change occurring on one or more nodes constitutes the starting point of the inference process.

[0093] After an initial perturbation is injected, the system infers the continuous impact of this perturbation on the system's future state through probabilistic propagation along the causal graph. This propagation process proceeds layer by layer along the directed edges in the graph, from the changed node to all its downstream adjacent nodes. In a specific implementation, the state of any node at the next time step is calculated as the result of a weighted function of the current states of all its upstream parent nodes, where the weights are the weights of the connecting edges. .

[0094] Because this propagation process is based on probability, its final result is not a definite future state, but a sequence of probability distributions of future states. This means that at each time step in the propagation process, the state of each affected node is represented as a probability distribution, rather than a single numerical value. This sequence depicts the probability distribution of all possible states of the system under the influence of the initial disturbance over multiple future time steps, thus forming one or more probabilistic trajectories, providing a data foundation for subsequent identification of risk trajectories.

[0095] Furthermore, after generating one or more probabilistic trajectories, the system evaluates these trajectories to identify risky trajectories. The identification of a risky trajectory is based on the probability that it leads to a violation of the system's health status rules.

[0096] In one specific implementation, the system first predefines one or more key invariants from the complete probabilistic invariant model constructed in step S200. These key invariants are a subset of the model and typically correspond to rules that ensure the stability of the system's core functions or meet the Service Level Agreement (SLA).

[0097] For each probabilistic trajectory generated above, the system calculates the probability that the trajectory will lead to a state that violates preset key invariants during its future evolution. Specifically, since the probabilistic trajectory is a probability distribution sequence of future states, the system analyzes this sequence to calculate the total probability that the system state will enter an "unhealthy" state space defined by one or more violated key invariants at a certain point in time or time period in the future.

[0098] The system then compares this calculated probability value with a preset risk threshold. This risk threshold is a numerical limit used to determine whether a risk is significant.

[0099] When the probability that a probabilistic trajectory leads to a state that violates preset key invariants exceeds a certain risk threshold, the system identifies the probabilistic trajectory as a risk trajectory. The identified risk trajectory indicates the most likely evolutionary path of the system under the current disturbance, leading to failure or performance degradation, and serves as direct input for the subsequent step S600 to generate an automated decision-making scheme.

[0100] In step S600, the system generates an automated decision-making scheme based on the risk trajectory identified in step S500. Before generating the final decision-making scheme, the system first performs a screening step to narrow down the decision-making scope.

[0101] The goal of this step is to construct a relevant and finite set of candidate operations based on the causal source of the risk. The system receives the risk trajectory information provided in step S500 and locates the critical invariant node in the trajectory that is about to be violated. This node is the starting point for subsequent operations.

[0102] The system traces backwards from the invariant nodes in the risk trajectory that are about to be violated on the corrected causal graph. This backward tracing process involves starting from the invariant node and traversing in the opposite direction of all directed edges pointing to it in the causal graph, continuing to the source nodes of these edges, and repeating this process with these source nodes until no further tracing is possible. This process will traverse one or more causal chains composed of nodes in the graph.

[0103] Through this reverse tracing, the system can identify all upstream nodes that have a direct or indirect causal impact on the invariant node that is about to be violated. Among these identified upstream nodes, the system further filters out nodes of type "configuration nodes." These configuration nodes represent parameters in the system that can be directly modified through external intervention, such as the number of service replicas, resource allocation quotas, or feature on / off states.

[0104] Finally, based on these selected configuration nodes, the system forms a set of candidate operations. Each candidate operation corresponds to a single modification of the value or state of an identified configuration node. Since each operation in this set is located on a causal chain leading to risk, they all have the potential to alter the course of that risk. This step significantly reduces the number of operations that need to be evaluated in subsequent decision-making steps by retaining only operations causally related to the risk, thus providing a foundation for rapidly solving for the optimal operation sequence.

[0105] After generating a set of candidate operations through causal pruning, the system performs counterfactual decision inference to determine and generate the final automated decision scheme from this set. This process formalizes the decision problem as a constrained optimization problem to find the lowest-cost sequence of operations while satisfying preset safety conditions.

[0106] The goal of this constrained optimization problem is to find an optimal sequence of operations. This sequence can guide the system away from the identified risk trajectory with minimal operating cost. The optimization objective is defined by the following mathematical expression: ; in: It represents a sequence of one or more operations selected from a set of candidate operations; It is a cost function used to quantify the sequence of operations to be performed. The resulting costs. In a specific implementation, this cost can be a weighted sum of multiple factors, such as the amount of computing resources required to perform the operation, the duration of service interruption that may result from the operation, and the complexity of the change.

[0107] At the same time, the selected operation sequence A constraint must be met: after executing this sequence of operations, the probability that the system will enter a risky state in future evolution must be lower than a preset safety threshold. This constraint is defined by the following mathematical expression: ; in: This represents the current state of the system before a decision is made. This is the candidate operation sequence being evaluated; It is the key invariant that is identified as about to be violated in the risk trajectory of step S500; Represents key invariants The incident that was violated; It is a conditional probability, representing the probability under the current state. Execution sequence Then, the probability that the key invariant is violated. This probability is calculated by simulating the sequence of operations on a causal graph. The new state trajectory is obtained by performing the same probabilistic propagation and risk assessment as in step S500.

[0108] It is a preset safety threshold, which is a constant close to 0, such as 0.01, representing an acceptable level of residual risk.

[0109] To solve this constrained optimization problem, the system considers each sequence of operations in the candidate operation set. The evaluation process begins by checking whether the sequence satisfies the aforementioned constraints. For all sequences that satisfy the constraints, the system calculates their corresponding operation costs. Ultimately, the system selects the sequence with the lowest cost as the optimal operation sequence. In a specific implementation, if the set of candidate operations is small, the system can solve the problem by traversing all possible sequences. If the set is large, a heuristic search algorithm or a planning algorithm can be used to find a solution that meets the conditions.

[0110] By solving this constrained optimization problem, the system ultimately generates an optimal sequence of operations that combines low cost and high security. This sequence is the final output of the automated decision-making scheme.

[0111] In summary, this invention constructs and synergistically utilizes two core models: a probabilistic invariant model and a causal graph. The probabilistic invariant model learns from historical health data, providing a data-driven, quantifiable benchmark for the system's normal operating status. The causal graph characterizes the causal transmission mechanism between various factors within the system. This invention establishes a closed-loop, self-consistent correction mechanism between models, using the probabilistic invariant model as a "fact standard" to continuously verify and optimize the accuracy of the causal graph. Based on this, real-time risk extrapolation using the corrected, high-precision causal graph can transform real-time events into predictions of future risk trajectories, and accurately locate relevant actionable variables through reverse tracing. Finally, by formalizing the decision-making process as an optimization problem with the goal of minimizing cost and constrained by a risk probability below a safety threshold, the system can automatically generate an optimal intervention strategy that is both economical and safe. This method achieves a leap from passive response to proactive prediction and automated decision-making, significantly improving the intelligence and reliability of complex system operation and maintenance.

[0112] The following is a specific embodiment to help understand the operation of the method disclosed in this invention in a real-world scenario.

[0113] Example: This embodiment uses an "order service" microservice deployed on a Kubernetes cluster as an example to illustrate how the present invention can proactively analyze risks and generate decisions when the microservice is being updated.

[0114] Scenario: The current stable version of the "Order Service" is v2.0. The operations team plans to release a new version, v2.1, which includes new features. One of the core Service Level Agreements (SLAs) for this service is that 99% of request processing latency must be below 300 milliseconds.

[0115] Steps 1 and 2 (S100-S200): Learning health behavior benchmarks; During the stable operation of version v2.0, the system of this invention continuously performs data acquisition and model learning.

[0116] Data Acquisition (S100): The system collects configuration data (5 replicas, CPU limited to 1 core), observation data (CPU utilization, memory consumption, request latency), and operation data (scaling up and down records) from the Order Service-v2.0. This data is converted into a unified structured event stream.

[0117] Invariant Mining (S200): Based on the aforementioned historical health data, the system mines a series of probabilistic invariants for context C (app=order service, version=v2.0). This includes a rule designated as a key invariant, directly corresponding to its SLA: I_SLA: P(StateAttr.latency_p99 < 300ms|C) >= 0.999. Additionally, other rules are included, such as I_CPU: P(StateAttr.cpu_usage < 0.85|C) >= 0.99.

[0118] Steps 3 and 4 (S300-S400): Construct and revise the causal graph; Initial Construction (S300): The system constructs an initial causal graph containing various nodes. For example, the graph includes configuration nodes "replica count" and "CPU limit", resource status nodes "CPU utilization", "memory utilization", and "request latency", and invariant nodes "I_SLA" and "I_CPU". The system establishes initial edges based on domain knowledge; for example, an increase in "CPU utilization" will "positively" affect "request latency".

[0119] Self-consistency correction (S400): The system initiates a closed-loop correction mechanism. For example, it simulates the operation of "reducing CPU limits" on the graph, and the graph predicts that "CPU utilization" will increase, and "request latency" will also increase accordingly. The system compares this prediction with the invariant model and finds that the predicted latency value increases the probability of I_SLA being violated. By minimizing this "logical conflict" between the prediction and the invariant model, the system repeatedly adjusts the weights of each edge in the graph (such as the influence strength of "CPU utilization" on "request latency") until the graph can accurately reproduce the healthy behavior pattern of version 2.0.

[0120] Step 5 (S500): Risk simulation after the release of the new version; The operations team performed a release operation, upgrading the "Order Service" to v2.1.

[0121] Perturbation Injection: The system captures the "Configuration Change" event and switches the context to C(app=Order Service, version=v2.1). Shortly after the release, the monitoring system reports a new status attribute event: the "Memory Usage" of the Order Service continues to climb, reaching 80%. This "Continuously Climbing Memory Usage" event is used as the initial perturbation and injected into the corresponding "Memory Usage" node in the corrected causal graph.

[0122] Risk Trajectory Identification: The system performs probabilistic propagation on the graph: High memory usage may trigger more frequent garbage collection (GC), leading to intermittent spikes in CPU usage, which in turn increases request latency. The probabilistic trajectory deduced by the system shows that there is a 75% probability that the 99th percentile value of request latency will exceed 300ms within the next 10 minutes. This probability (75%) far exceeds the preset risk threshold (e.g., 50%), so the system identifies this trajectory as a "risk trajectory" and determines that the key invariant I_SLA is about to be violated.

[0123] Step 6 (S600): Automated decision-making process generation; Once the system identifies a risk, it immediately initiates the decision-making process.

[0124] Causal pruning: The system traces backwards from the "I_SLA" invariant node in the graph (whose state depends on the "Request Delay" node). The tracing path is: "Request Delay" ← "CPU Utilization" ← "Memory Utilization". The system continues tracing from the "Memory Utilization" node and finds operable configuration nodes upstream, such as "Memory Limit" and "Replica Count". Meanwhile, the root cause of this problem is the code behavior of version v2.1 itself. Therefore, the system selects the following candidate operations: {"Increase Memory Limit", "Increase Replica Count", "Rollback to v2.0"}.

[0125] Counterfactual decision deduction: The system formalizes the decision problem into an optimization problem with the objective of minimizing "operational costs" and the constraint that "the probability of future SLA violations is less than 1%".

[0126] Evaluation of "Increasing memory limits": Counterfactual analysis shows that even with increased memory, the memory leak patterns in the new code version will quickly exhaust it again, failing to meet the constraints. This proposal was rejected.

[0127] Evaluation of "Increasing the number of replicas": The simulation shows that increasing the number of replicas can distribute traffic, temporarily alleviate the latency problem, and meet the constraints. However, its cost is "moderate" (requiring additional resources).

[0128] Evaluation of "Rollback to v2.0": The simulation shows that rollback can completely solve the problem and meet the constraints. However, its cost is "high" (it interrupts the release process and requires manual intervention for review).

[0129] The system compared all solutions that met the constraints and determined that "increasing the number of replicas" was less costly than "rolling back to v2.0". Therefore, the final decision generated by the system was: "Execute automated operation: expand the number of replicas for 'Order Service' from 5 to 8 to address the anticipated high latency risk." Simultaneously, the system generated an alert, alerting the operations team that version v2.1 had a potential memory issue and recommending offline investigation.

[0130] As can be seen from this embodiment, the present invention can automatically complete the entire process from sensing changes and predicting risks to decision-making and handling without human intervention, which significantly improves the predictability and automation level of operation and maintenance.

[0131] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for generating automated decision-making solutions for cloud platform operation and maintenance risks using big data analysis, characterized in that: Includes the following steps: Acquire multimodal operation and maintenance data and convert it into a structured stream of behavioral events; Based on the structured behavioral event flow, a probabilistic invariant model is constructed, which is used to characterize the probabilistic rules that the system's healthy behavior should follow in a specific context. Construct a causal graph to characterize the causal transmission relationships between factors within the system; The causal graph is corrected by comparing the inference results of the causal graph with the probabilistic invariant model, thereby improving the accuracy of the causal graph. Using the modified causal graph, the evolution of system state triggered by real-time behavioral events is deduced, and risk trajectories are identified. Based on the risk trajectory, an automated decision-making scheme is generated.

2. The method for generating big data analysis and automated decision-making solutions for cloud platform operation and maintenance risks according to claim 1, characterized in that, The steps for constructing the probabilistic invariant model include: An invariant is defined as a logical assertion, wherein the conditional probability of a relationship between one or more attributes within the system being true in a specific context satisfies a preset confidence threshold; and, When the specific context changes, the probabilistic rules in the probabilistic invariant model related to that context are dynamically updated.

3. The method for generating big data analysis and automated decision-making solutions for cloud platform operation and maintenance risks according to claim 1, characterized in that, The step of correcting the causal map includes: The system state is simulated and deduced using the causal graph to obtain the predicted state; Evaluate the consistency between the predicted state and the probabilistic invariant model; and, Based on the consistent evaluation results, the parameters of the causal graph are adjusted to minimize the logical conflict between the inference results of the causal graph and the probabilistic invariant model.

4. The method for generating big data analysis and automated decision-making solutions for cloud platform operation and maintenance risks according to claim 3, characterized in that, The step of adjusting the parameters of the causal graph is performed using a gradient-based optimization method with the goal of minimizing the quantization value of the logical conflict.

5. The method for generating big data analysis and automated decision-making solutions for cloud platform operation and maintenance risks according to claim 1, characterized in that, The steps for identifying risk trajectories include: Real-time behavioral events are used as initial perturbations input into the causal graph. By probabilistically propagating along the causal graph, a sequence of probability distributions for future states is deduced, forming a probabilistic trajectory; and, When the probability of a probabilistic trajectory leading to a state that violates a preset key invariant exceeds a risk threshold, it is identified as a risk trajectory.

6. The method for generating big data analysis and automated decision-making solutions for cloud platform operation and maintenance risks according to claim 1, characterized in that, The step of generating an automated decision-making scheme further includes, before generating the decision-making scheme: On the causal graph, the invariant nodes that are about to be violated in the risk trajectory are traced back in reverse to screen out candidate operations that have a causal impact on the invariant nodes, forming a set of candidate operations.

7. The method for generating big data analysis and automated decision-making solutions for cloud platform operation and maintenance risks according to claim 6, characterized in that, The steps for generating automated decision-making solutions also include: Within the decision space formed by the set of candidate operations, an optimal sequence of operations is obtained through counterfactual decision deduction, with the optimization objective of operation cost and the constraint that the probability of the system entering a risky state in the future is lower than a preset safety threshold. This sequence serves as the automated decision-making scheme.

8. The method for generating big data analysis and automated decision-making solutions for cloud platform operation and maintenance risks according to claim 1, characterized in that, The nodes of the causal graph include: Configuration nodes, resource status nodes, and invariant nodes representing the probabilistic rules in the probabilistic invariant model.

9. The method for generating big data analysis and automated decision-making solutions for cloud platform operation and maintenance risks according to claim 1, characterized in that, The step of acquiring multimodal operation and maintenance data and converting it into a structured behavioral event stream involves uniformly converting configuration data, observation data, and operation data into a structured event sequence that includes declaration events, execution events, invocation events, and status attribute events.

10. A cloud platform operation and maintenance risk analysis and decision-making scheme generation system, used to execute the method as described in any one of claims 1-9, characterized in that, include: The data transformation module is used to acquire multimodal operation and maintenance data and convert it into a structured stream of behavioral events. The model building module is used to construct a probabilistic invariant model based on the structured behavioral event flow to characterize the probabilistic rules that the system's healthy behavior should follow in a specific context. The graph construction and correction module is used to construct a causal graph to characterize the causal transmission relationship between factors within the system, and to correct the causal graph by comparing the inference results of the causal graph with the probabilistic invariant model, so as to improve the accuracy of the causal graph. The risk simulation module is used to use the modified causal graph to simulate the evolution of the system state caused by real-time behavioral events and identify risk trajectories. The decision generation module is used to generate automated decision-making schemes based on the risk trajectory.