Operation and maintenance strategy evaluation method and device, equipment, medium and product
By acquiring anomaly datasets from business systems for root cause analysis, generating executable operation and maintenance programs, and executing them in a sandbox isolation environment, the problem of insufficient intelligent reasoning capabilities in traditional business systems is solved, enabling effective operation and maintenance strategy evaluation for complex business scenarios.
Patent Information
- Application Number
- CN202511857927.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-10
- Publication Date
- 2026-02-27
AI Technical Summary
Traditional business system operation and maintenance methods suffer from insufficient intelligent reasoning capabilities and poor adaptability, making them unable to effectively cope with complex business scenarios.
By acquiring the abnormal dataset of the business system, we perform root cause analysis, generate executable operation and maintenance programs, execute them in a sandbox isolation environment, record status data, and generate operation and maintenance strategy evaluation reports.
It enables data-driven automated operation and maintenance strategy evaluation in complex business scenarios, providing effective and reliable operation and maintenance decision-making references.
Smart Images

Figure CN121579262A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of operation and maintenance, and in particular to an operation and maintenance strategy evaluation method, device, equipment, medium and product. BACKGROUND
[0002] Intelligent operation and maintenance is an intelligent management system that deeply integrates artificial intelligence, machine learning and big data technology into traditional IT operation and maintenance, aiming to realize real-time monitoring, anomaly detection, root cause analysis, predictive maintenance and self-healing repair through automation and intelligent means, thereby improving operation and maintenance efficiency and reducing the need for manual intervention.
[0003] Traditional business systems are mainly operated and maintained in the following ways: threshold-based alarm, rule engine based on expert system, and repair based on single fault. These methods have core defects such as insufficient intelligent reasoning ability and poor adaptability, and cannot effectively cope with complex business scenarios. SUMMARY
[0004] The present application provides an operation and maintenance strategy evaluation method, device, equipment, medium and product to solve the problem of insufficient intelligent reasoning ability and poor adaptability of the operation and maintenance method of traditional business systems, which cannot effectively cope with complex business scenarios.
[0005] In a first aspect, an operation and maintenance strategy evaluation method is provided, comprising:
[0006] Obtaining an abnormal data set of a business system; the abnormal data set includes abnormal alarm data and monitoring data;
[0007] Performing root cause analysis on the abnormal data set to obtain a root cause analysis result;
[0008] Generating an executable operation and maintenance program according to the abnormal data set, the root cause analysis result and a target program template;
[0009] Determining a resource configuration strategy corresponding to the executable operation and maintenance program according to the root cause analysis result;
[0010] Executing the executable operation and maintenance program in a sandbox isolation environment according to the resource configuration strategy, and recording state data of the executable operation and maintenance program in the execution process;
[0011] Generating an operation and maintenance strategy evaluation report according to the root cause analysis result, the executable operation and maintenance program and the state data.
[0012] In a second aspect, an operation and maintenance strategy evaluation device is provided, comprising:
[0013] The data set acquisition module is configured to acquire an abnormal data set of a business system, wherein the abnormal data set comprises abnormal alarm data and monitoring data.
[0014] The root cause analysis module is configured to perform root cause analysis on the abnormal data set to obtain a root cause analysis result.
[0015] The program generation module is configured to generate an executable operation and maintenance program according to the abnormal data set, the root cause analysis result and a program template.
[0016] The resource configuration module is configured to determine a resource configuration strategy corresponding to the executable operation and maintenance program according to the root cause analysis result.
[0017] The program execution module is configured to execute the executable operation and maintenance program in a sandbox isolation environment according to the resource configuration strategy, and record state data of the executable operation and maintenance program in an execution process.
[0018] The operation and maintenance evaluation module is configured to generate an operation and maintenance strategy evaluation report according to the root cause analysis result, the executable operation and maintenance program and the state data.
[0019] In a third aspect, an electronic device is provided, and the electronic device comprises:
[0020] at least one processor;
[0021] and a memory in communication connection with the at least one processor;
[0022] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the operation and maintenance strategy evaluation method described in any embodiment of the present application.
[0023] In a fourth aspect, a computer readable storage medium is provided, and the computer readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the operation and maintenance strategy evaluation method described in any embodiment of the present application when executed.
[0024] In a fifth aspect, a computer program product is provided, and the computer program product comprises a computer program, and the computer program is used to enable a processor to implement the operation and maintenance strategy evaluation method described in any embodiment of the present application when executed.
[0025] The technical scheme of the embodiment of the application comprises the following steps: obtaining an abnormal data set of a business system; the abnormal data set comprises abnormal alarm data and monitoring data; performing root cause analysis on the abnormal data set to obtain a root cause analysis result; generating an executable operation and maintenance program according to the abnormal data set, the root cause analysis result and a target program template; determining a resource configuration strategy corresponding to the executable operation and maintenance program according to the root cause analysis result; executing the executable operation and maintenance program in a sandbox isolation environment according to the resource configuration strategy, and recording state data of the executable operation and maintenance program in the execution process; and generating an operation and maintenance strategy evaluation report according to the root cause analysis result, the executable operation and maintenance program and the state data. The root cause of abnormal data is analyzed by root cause analysis, and an executable operation and maintenance program is generated. The executable operation and maintenance program is safely executed in a sandbox isolation environment to evaluate the operation and maintenance strategy. Through the closed-loop operation and maintenance strategy of analysis-decision-execution-evaluation, an effective and reliable operation and maintenance strategy can be automatically given based on data driving in the face of complex business scenarios. The problems of insufficient intelligent reasoning ability, poor self-adaptability and inability to effectively respond to complex business scenarios in the traditional operation and maintenance mode of a business system are solved, and important reference value is provided for operation and maintenance decision-making.
[0026] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the application, nor is it intended to limit the scope of the application. Other features of the application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0027] In order to more clearly illustrate the technical solutions in the embodiments of the application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0028] Figure 1 A flowchart of an operation and maintenance strategy evaluation method provided for the first embodiment of the application;
[0029] Figure 2 A flowchart of an operation and maintenance strategy evaluation method provided for the second embodiment of the application;
[0030] Figure 3 A structural schematic diagram of an operation and maintenance strategy evaluation device provided for the third embodiment of the application;
[0031] Figure 4 A structural schematic diagram of an electronic device for implementing the operation and maintenance strategy evaluation method of the embodiment of the application. DETAILED DESCRIPTION
[0032] In the following, the technical solutions in the embodiments of the present application will be described clearly and completely with reference to the drawings in the embodiments of the present application in order to make the technical personnel in the technical field better understand the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by the ordinary skilled in the art without creative work should belong to the scope of protection of the present application.
[0033] It should be noted that the terms "first", "second", and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product, or device that includes a series of steps or units does not have to be limited to only those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products, or devices.
[0034] Embodiment one
[0035] Figure 1 A flowchart of a method for evaluating an operation and maintenance strategy provided by the first embodiment of the present application, the present embodiment can be applicable to the generation and evaluation of executable operation and maintenance procedures for abnormal business systems. The method can be executed by an operation and maintenance strategy evaluation device, which can be realized in the form of hardware and / or software, and can be configured in an electronic device. As shown in the figure, the method comprises: Figure 1
[0036] S110, acquiring an abnormal data set of a business system; the abnormal data set comprises abnormal alarm data and monitoring data.
[0037] The abnormal data set can be considered as a data set obtained under the condition of abnormal business system. The abnormal alarm data can be considered as alarm data for prompting the abnormality of the business system, and the monitoring data can be considered as data monitored during the operation of the business system.
[0038] In this embodiment, the abnormal alarm data can be triggered automatically by a rule engine, threshold judgment or abnormal detection logic built in the business system, and can generally be a structured message (such as JSON), which can include, for example, information such as the type of abnormality, the occurrence time of the abnormal event, and the business entity of the abnormality. The monitoring data can include log files (such as application running logs, system logs, interface call logs), image data (such as network topology diagrams and traffic fluctuation diagrams), link tracking data (such as path data of cross-service calls in a distributed system: recording the complete call path of a request in the system, service time consumption and dependency relationship), and numerical data. The abnormal data set can have multiple modalities, and the embodiment does not limit the format and specific content of the abnormal data set. In this embodiment, the image data can be subjected to OCR or visual analysis to extract key information. The collected abnormal data set can also be subjected to data cleaning and normalization processing.
[0039] In this embodiment, the abnormal data set can be used as a fact basis for analysis of the abnormal event of the business system and participate in the generation of the operation program.
[0040] S120, performing root cause analysis on the abnormal data set to obtain a root cause analysis result.
[0041] The root cause analysis can be considered as a process of analyzing the root cause of the abnormal event. The root cause analysis result can be considered as a conclusion obtained by analysis. In this embodiment, the root cause analysis result can include an abnormal root cause, an evidence chain set and an abnormal type. The abnormal root cause can be considered as a cause and object for describing the occurrence of the abnormal event. The evidence chain can be considered as an ordered set of all data and fact bases supporting the abnormal root cause.
[0042] In this embodiment, the root cause analysis on the abnormal data can be performed by a reasoning large model or a root cause analysis algorithm. In the root cause analysis process, the domain knowledge graph and / or application rules of the business system can also be combined for reasoning.
[0043] S130, generating an executable operation program according to the abnormal data set, the root cause analysis result and a target program template.
[0044] The executable operation program can be understood as an executable program for maintaining the abnormal event. The target program template can be considered as a standard file for generating the executable operation program.
[0045] In this embodiment, a suitable target program template is selected according to the abnormal data set and the root cause analysis result, target parameters are determined according to the abnormal data set and the root cause analysis result, and the executable operation program is generated according to the target parameters and the target program template.
[0046] In this embodiment, a program template library can be constructed to store program templates for various anomaly types. A program template can include a business logic framework, configurable parameters, metadata tags, and adaptation points. The business logic framework is an analysis process defined using pseudocode or real code with placeholders (Python / SQL, etc.); configurable parameters are the parameters that need to be input; metadata tags are tags describing the scenario in which the template is used, the data type of the input, performance characteristics, etc.; adaptation points are predefined extension interfaces that support adjustments based on root cause conclusions. For example, the template types can include: time-series detection templates, log mining templates, and trajectory analysis templates. Time-series detection templates encapsulate the detection logic for metrics such as CPU, memory, and latency, integrate multiple algorithms, and automatically select based on data characteristics, ultimately outputting a judgment of whether an anomaly is present and its confidence level; log mining templates encapsulate the anomaly pattern process extracted from historical records, employing natural language processing technology and unsupervised clustering algorithms to automatically discover error log clusters and anomaly patterns. Trajectory analysis templates are used to construct flowcharts from time logs based on process mining algorithms, automatically identifying bottleneck nodes, abnormal paths, and frequent fault combinations.
[0047] S140. Determine the resource configuration strategy corresponding to the executable operation and maintenance program based on the root cause analysis results.
[0048] Resource allocation strategy can be considered as the strategy for allocating hardware or software resources for executing executable operation and maintenance programs.
[0049] In this embodiment, executing the executable operation and maintenance program requires necessary resources, such as the CPU, memory, and network I / O of the execution node. Therefore, it is also necessary to determine a resource configuration strategy for the executable operation and maintenance program.
[0050] S150. Execute the executable operation and maintenance program in the sandbox isolation environment according to the resource configuration policy, and record the status data of the executable operation and maintenance program during the execution process.
[0051] A sandbox, or sandbox for short, is a security mechanism used to run unverified programs or code in a restricted environment, isolating them from the host system, network, and other critical resources. By minimizing the impact of untrusted code, it significantly improves the overall security and stability of the system.
[0052] In this embodiment, after determining the resource configuration strategy and the executable operation and maintenance program, a sandbox isolation environment is constructed according to the resource configuration strategy, and the executable operation and maintenance program is executed in the sandbox isolation environment. During the execution of the executable operation and maintenance program, status data is continuously recorded, such as the status data of key nodes, execution results, execution trajectory, and performance status data, to achieve real-time monitoring of the executable operation and maintenance program.
[0053] S160, generating an operation and maintenance strategy evaluation report according to the root cause analysis result, the executable operation and maintenance program and the state data.
[0054] In this embodiment, the influence range and program of the exception, the feasibility of the executable operation and maintenance program and the required consumed resources are comprehensively analyzed according to the obtained root cause analysis result, the generated executable operation and maintenance program and the state data generated by executing the executable operation and maintenance program.
[0055] The technical scheme of the embodiment of the application comprises the following steps: obtaining an exception data set of a business system; the exception data set comprises exception alarm data and monitoring data; performing root cause analysis on the exception data set to obtain a root cause analysis result; generating an executable operation and maintenance program according to the exception data set, the root cause analysis result and a target program template; determining a resource configuration strategy corresponding to the executable operation and maintenance program according to the root cause analysis result; executing the executable operation and maintenance program in a sandbox isolation environment according to the resource configuration strategy, and recording state data of the executable operation and maintenance program in the execution process; and generating an operation and maintenance strategy evaluation report according to the root cause analysis result, the executable operation and maintenance program and the state data. The root cause of the exception data is analyzed and the executable operation and maintenance program is generated through root cause analysis, the operation and maintenance strategy is evaluated by safely executing the executable operation and maintenance program in the sandbox isolation environment, and the closed-loop operation and maintenance strategy of analysis-decision-execution-evaluation can automatically provide an effective and reliable operation and maintenance strategy based on data driving to face complex business scenarios, thereby providing an important reference value for operation and maintenance decision-making.
[0056] Embodiment two
[0057] Figure 2 A flowchart of an operation and maintenance strategy evaluation method provided by the second embodiment of the application is shown in the figure, and the embodiment further refines S120 of the above-mentioned embodiment. Specifically, the root cause analysis on the exception data set is performed to obtain a root cause analysis result, which comprises the following steps: obtaining an analysis target by performing semantic analysis and target mapping on the exception data set in combination with a domain knowledge graph; performing exception reasoning based on a forward chain and a backward chain combined reasoning mechanism according to the analysis target and the exception data set to obtain a root cause analysis result; and the root cause analysis result comprises an exception root cause and a correlation evidence chain.
[0058] As shown in the figure, the method comprises the following steps: Figure 2
[0059] S210, obtaining an exception data set of a business system; the exception data set comprises exception alarm data and monitoring data.
[0060] S220, obtaining an analysis target by performing semantic analysis and target mapping on the exception data set in combination with a domain knowledge graph.
[0061] In this embodiment, the dependency relationship between words in a sentence can be identified by a large model driven based on the historical data set of the business system, and a domain knowledge graph (such as the relationship between device views) is constructed. The semantics of entities and actions in the abnormal data set are parsed with the assistance of the domain knowledge graph, and the description of the fault entity is mapped to a structured target for analysis.
[0062] For example, the abnormal level is parsed from the abnormal alarm data, and the entity keywords such as the subject and object of the abnormal event are parsed from the monitoring data. In a complex scenario, a “coarse parsing-fine parsing” method is used, that is, the core element logic is first extracted, and then the abnormal level and entity keywords are refined from the core element logic. According to the mapping of the abnormal registration and the entity keywords to the target in the domain knowledge graph, the target for analysis is obtained.
[0063] S230, according to the target for analysis and the abnormal data set, based on the reasoning mechanism combining the forward chain and the backward chain, the abnormal reasoning is performed to obtain the root cause analysis result; the root cause analysis result includes the abnormal root cause and the associated evidence chain.
[0064] In this embodiment, the reasoning of the forward chain is to start from the known data / facts, apply rules, and deduce new conclusions until no new facts are generated or the target is reached. The reasoning of the backward chain is to start from a hypothetical target / conclusion, and find rules and evidence supporting the hypothesis. If sufficient evidence cannot be found, another hypothesis is tried.
[0065] As an implementation manner of this embodiment, according to the target for analysis and the abnormal data set, based on the reasoning mechanism combining the forward chain and the backward chain, the abnormal reasoning is performed to obtain the root cause analysis result, including: starting from the known facts in the abnormal data set and the target for analysis, the rules in the rule library are applied to deduce until the abnormal root cause is obtained; starting from the abnormal root cause, the evidence chain supporting the abnormal root cause is found from the fact library in reverse.
[0066] In this embodiment, high-frequency rules can be extracted from historical operation and maintenance data, and low-frequency rules can be eliminated to construct a rule library. The specific construction method is as follows: the frequency of rules successfully applied in historical reasoning is counted, new rules are generated based on high-frequency rules, abnormal patterns are analyzed by clustering, new rules are added to the rule library, and low-frequency rules are eliminated.
[0067] In the embodiment, the reasoning of the forward chain is based on data-driven rule discovery, starting from known facts in the abnormal data set and the target to be analyzed, and gradually deducing the conclusion by applying rules in the rule base; when the facts meet the premise conditions of the rules, the rules are triggered and new facts are generated, until the target is reached or the deduction cannot continue. The specific steps are: extracting known facts in the abnormal data set, traversing the rule base, matching the content of the fact occurrence with the rules in the rule base, finding the matching rule items, the rule base contains multiple "if-then" rules; executing the matched rules to determine the possible hypothesis facts or conclusions; repeating the above steps until no new facts are generated or the number of preset reasoning is reached, generating an initial possible hypothesis abnormal root cause.
[0068] The reasoning of the backward chain is based on hypothesis abnormal root cause driven hypothesis deduction, starting from the hypothesis abnormal root cause obtained in the above step, and finding supporting evidence in the reverse direction to prove that the hypothesis abnormal root cause is correct, and checking whether there is evidence supporting the hypothesis abnormal root cause in the fact base. The specific steps are: finding all rules in the rule base whose conclusion part can match the hypothesis abnormal root cause to obtain multiple sub-targets that need to be proved. Check if the sub-target is in the fact base. If the corresponding data of the sub-target is found in the fact base, the sub-target is verified. If the sub-target is not in the fact base, it is regarded as a new hypothesis target, and the rules in the rule base whose conclusion part can match the hypothesis abnormal root cause are returned to obtain multiple sub-targets that need to be proved, and the matching and decomposition rules are continued until all sub-targets are decomposed into verifiable atomic facts. If all sub-targets corresponding to the hypothesis abnormal root cause are verified, the hypothesis abnormal root cause is confirmed to be correct, that is, the hypothesis abnormal root cause is determined as the abnormal root cause; otherwise, the proved root cause hypothesis is not correct, backtrack to find other matching root cause rules, and repeatedly verify the sub-targets until the target is verified. If all possible targets cannot be verified, the root cause hypothesis is not correct.
[0069] The abnormal root cause and the associated evidence chain are obtained by comprehensively reasoning the forward chain and the backward chain.
[0070] S240, generating an executable operation and maintenance program according to the abnormal data set, the root cause analysis result, and the target program template.
[0071] As an implementation manner of the embodiment, the executable operation and maintenance program is generated according to the abnormal data set, the root cause analysis result, and the target program template, including: determining a target program template from the program template library according to the abnormal type of the abnormal root cause based on multi-dimensional features and reinforcement learning strategies; extracting target parameters from the abnormal data set and the evidence chain; inputting the target parameters into the target program template to generate initial operation and maintenance code; and performing syntax checking on the initial operation and maintenance code, and determining the initial operation and maintenance code that passes the syntax checking as executable operation and maintenance code.
[0072] In this embodiment, based on multi-dimensional features and reinforcement learning strategy, a program template with the same abnormal type as the abnormal root cause is matched from the program template library as a target program template. Based on the feature extraction algorithm, target parameters such as abnormal time, abnormal entity, and abnormal log path are extracted from the abnormal data set and the evidence chain. The extracted target parameters are input into the target program template, and an initial operation and maintenance code is generated by using a dynamic code generation technology. The initial operation and maintenance code can be in the format of Python, SQL query, Shell command, etc.
[0073] In this embodiment, if the target program template cannot be determined from the program template library according to the abnormal type of the abnormal root cause, a large language model can be used to convert the system fault description, operation and maintenance scene, and abnormal feature information into a structured code template framework through natural language, thereby realizing dynamic expansion of the template library. For example, an LLM model is called to convert the natural language description input by the user into a structured template, including: detection strategy (selecting what model and rule), usage scene description, script / code logic, and metadata (such as input parameter, applicable index, and historical success rate). The generated template is automatically stored in the dynamic template library and labeled as a new scene tag, supporting manual review and optimization.
[0074] In this embodiment, the multi-dimensional feature and reinforcement learning strategy can be represented by a multi-dimensional weighted scoring model matching algorithm as follows: ; represents the weight of the i-th scoring dimension, represents the weight of the i-th scoring dimension, and the scoring dimensions can include execution efficiency (such as execution time), resource occupancy rate, historical success rate, semantic matching degree, and scene matching degree. The initial values of the weights of the scoring dimensions can be preset, or the weights can be adjusted according to the actual execution effect (such as success rate and efficiency) to adjust the rewards, so that the matching strategy is more accurate. The weight adjustment process can be: executing the executable code generated by the template, recording the execution result of the executable code, and observing the execution effect (such as success rate and delay); calculating the reward value according to the execution result. Using a policy gradient algorithm, the weight is adjusted in reverse according to the reward value, so that in the future in a similar scene, a template that can obtain a higher reward can be selected.
[0075] For example, the policy gradient (Policy Gradient) algorithm can be as follows: key elements include: state (State / s): composed of business parameters and system state. Action (Action / a): represents the adjustment amount of the weight w in the scoring formula. Reward (Reward / r): a scalar value calculated according to the execution effect. The policy gradient formula is:
[0076] ;
[0077] where, is the parameter vector of the policy network, containing all trainable weights and biases; is the policy value function, measuring the overall performance of the policy, with a larger value indicating a better policy; is the "state-action-reward" sequence; : the gradient of the policy parameter , indicating the direction of parameter adjustment to maximize ; is the parameterized policy, determined by the parameter ; E represents the expectation operator, representing the average over all possible trajectories; is the weight generation policy, determined by the parameter ; is the action value function, representing the expected cumulative return that can be obtained by following the policy after performing action in state .
[0078] In order to enable the policy gradient to "perceive" the execution effect that the business cares about, the reward function needs to include target indicators (such as success rate, efficiency). In the form of multi-objective linear weighting, it is flexible and can intuitively adjust the priority of different indicators. The reward R(s,a) needs to include target indicators (such as success rate, efficiency), which can be represented as
[0079]
[0080] where, are two hyperparameters that need to be tuned in combination with the business scenario: if "success rate" is more important, increase ; if delay (the opposite of efficiency) has a greater impact on the business, increase .
[0081] Use REINFORCE to update parameters, REINFORCE is a Monte Carlo implementation of policy gradient, which uses "the discounted return of the entire trajectory" to approximate the action value function , thereby avoiding modeling the value function and directly updating the policy parameters through sampling
[0082]
[0083] where N is the total number of sampled trajectories, superscript (i) represents "the i-th trajectory"; a is the learning rate, which controls the parameter update step size; T is the length of a single trajectory (i.e. how many steps of state-action-reward interaction does the trajectory contain); The weight generation strategy for the i-th trajectory at step t (determined by parameters) Decision, in state Lower output weight adjustment (probability) This is the discounted reward for the i-th trajectory and the t-th step: if no discount is used (i.e., immediate feedback is valued), then γ=1. ; It is the instant reward obtained at the k-th step of the i-th trajectory.
[0084] By following the complete chain of "policy gradient formula - REINFORCE implementation - reward function design", the dynamic adjustment goal of "automatically optimizing weights based on execution effect under unlabeled data" can be achieved. This allows the template to be dynamically optimized and a better template to be selected.
[0085] S250. Determine the resource configuration strategy corresponding to the executable operation and maintenance program based on the root cause analysis results.
[0086] As one implementation of this embodiment, determining the resource configuration strategy of the executable operation and maintenance program based on the root cause analysis results includes: determining resource requirement information based on the anomaly type of the abnormal root cause; predicting the load status of multiple execution nodes executing the executable operation and maintenance program based on the resource requirement information; and determining the target execution node based on the load status of the multiple execution nodes.
[0087] In this embodiment, the resource allocation strategy may include resource requirement information and target execution nodes. The corresponding resource requirement information is determined based on the anomaly type of the root cause of the anomaly. Resource requirement information may include CPU, memory, network bandwidth, and network throughput, etc. Historical resource data from multiple execution nodes is collected based on the resource requirement information. An LSTM time series prediction model is used to analyze the historical resource data of each node to predict the load status of executable operation and maintenance programs. The prediction results are used to optimize the resource allocation strategy, such as avoiding assigning new tasks to nodes that are about to reach saturation, thus avoiding local overload; or triggering expansion, contraction, or scheduling strategy optimization in advance to avoid system overload.
[0088] For example, a Long Short-Term Memory (LSTM) network is used to achieve resource balance between short-term memory and long-term state through a core gating structure: input gate, forget gate, cell state, and output gate. The specific steps are as follows: Step 1: Historical data acquisition and preprocessing. S1 acquires and preprocesses historical resource data (e.g., timestamps, task types, resource utilization, error codes, etc.), sorts out the time series dataset, and each time step contains multi-dimensional features. The data features are vectorized for use as model input.
[0089] Step 2: Forget Gate Processing (Controlling the Degree of Retention of Historical Information). This step determines which historical execution patterns need to be forgotten (e.g., if a certain period in the past has no impact on the current task, its weight is reduced). ; The forget gate activation value at time t is represented, with a value range of [0,1], indicating the degree of retention of historical information; The weight matrix representing the forget gate is learned through training; σ represents the bias term of the forget gate; σ represents the Sigmoid activation function, which compresses the value to the interval between 0 and 1. This indicates the hidden state at the previous moment.
[0090] Step 3: Input Gate and Candidate Cell States. The input gate controls the inflow of new information. This step determines which new information needs to be stored in the cell state at the current moment: the input gate calculation can determine which parts are updated. For example: The input gate activation value at time t determines the degree of new information inflow and takes the value [0, 1]. This represents the weights corresponding to the input gates; This represents the bias term of the input gate; This represents the input feature vector at the current time t. Candidate cell state calculation can provide candidates for new values, such as: in: This indicates the candidate cell state and represents the new memory content that may result from the current input. This represents the weight corresponding to the candidate state; This represents the bias corresponding to the candidate state; : Activation function that maps values to [-1, 1] and is used to generate new memory content.
[0091] Step 4: Cell State Update - Integrating Old and New Information: Cell state update (long-term memory update) integrates old and new information to form a comprehensive understanding of the task execution status. ; The cell state at time t is represented, storing long-term memory information. C represents the cell state and is the core variable responsible for long-term memory in LSTM. ⊙ represents element-wise multiplication (Hadamard product). This represents the cell state at time t-1.
[0092] Step 5: Output gate and hidden state generation control the output level of the current state, controlling how much information from the current cell state needs to be output to the hidden layer. This is used for the final resource requirement prediction and bottleneck state determination in this step.
[0093] Output gate calculation: ;in: The output gate activation value at time t determines how much information is output for the current state. Hidden State Output: ;in, This represents the hidden state at the current moment, which is used for the final prediction output (e.g., resource requirements, bottleneck probability). This indicates that the final hidden state is obtained by multiplying the cell state by the output gate after performing a nonlinear transformation on the cell state.
[0094] S260. Execute executable operation and maintenance programs in a sandbox isolation environment according to resource configuration policies, and record the status data of the executable operation and maintenance programs during the execution process.
[0095] In one implementation of this embodiment, resources are scheduled and the executable operation and maintenance program is executed in a sandbox isolation environment according to the resource configuration strategy, and the status data of the executable operation and maintenance program during execution is recorded. This includes: constructing a sandbox isolation environment according to the resource configuration strategy, wherein the sandbox isolation environment includes resource isolation and process isolation; executing the executable operation and maintenance program in the sandbox isolation environment; and recording the status data during program execution, wherein the status data includes status data and performance status data, wherein the status data includes variable data of key nodes, execution results, and execution trajectory; and the performance status data includes resource utilization and response time.
[0096] In this embodiment, the resource usage of each task instance in the executable operation and maintenance program is limited by containerization technology according to the resource configuration strategy, ensuring that each task instance runs in an independent resource environment and avoiding resource competition and interference between different task instances. The kernel's namespace technology is utilized to allow task instances to execute in independent process spaces, network stacks, and process trees, avoiding inter-process interference or escape. Specifically, this may include: allocating an independent process space for each task instance, limiting its process tree range to prevent signal interference or handle leakage; and supporting rapid isolation and reclamation of related resources when process anomalies (such as crashes, deadlocks, or resource overruns) are detected, preventing impact on other task instances.
[0097] In this embodiment, during the execution of the executable operation and maintenance program, status data and performance status data across steps are maintained to ensure the continuity and recoverability of the analysis. Status data includes variable data, execution results, and execution trajectories for key nodes. Saving variable data for key nodes is used for error recovery; when anomalies occur during program execution, these intermediate results can be used for troubleshooting and recovery, reducing the cost of re-executing tasks. Detailed execution trajectories are recorded for debugging and optimization; by analyzing the execution trajectory, the program's execution process can be understood, potential problems can be identified, and code and task execution strategies can be optimized. Performance status data includes resource utilization and response time. By monitoring resource utilization, timely alerts and handling can be initiated when resource consumption is too high; an alarm is triggered when resource usage exceeds a safe threshold.
[0098] S270. Generate an operation and maintenance strategy evaluation report based on the root cause analysis results, executable operation and maintenance procedures, and status data.
[0099] The technical solution of this invention involves acquiring an anomaly dataset from a business system; the anomaly dataset includes anomaly alarm data and monitoring data; performing root cause analysis on the anomaly dataset to obtain root cause analysis results; generating an executable operation and maintenance program based on the anomaly dataset, root cause analysis results, and a target program template; determining the resource configuration strategy corresponding to the executable operation and maintenance program based on the root cause analysis results; executing the executable operation and maintenance program in a sandbox isolation environment according to the resource configuration strategy, and recording the status data of the executable operation and maintenance program during execution; and generating an operation and maintenance strategy evaluation report based on the root cause analysis results, the executable operation and maintenance program, and the status data. By analyzing the root causes of anomalies through root cause analysis and generating executable operation and maintenance programs, and securely executing these programs in a sandbox isolation environment to evaluate operation and maintenance strategies, this closed-loop operation and maintenance strategy of analysis-decision-execution-evaluation can automatically provide effective and reliable operation and maintenance strategies based on data-driven approaches in complex business scenarios, providing important reference value for operation and maintenance decisions.
[0100] Example 3
[0101] Figure 3 This is a schematic diagram of the structure of an operation and maintenance strategy evaluation device provided in Embodiment 3 of the present invention. Figure 3 As shown, the device includes: a dataset acquisition module 310, a root cause analysis module 320, a resource configuration module 330, a program generation module 340, a program execution module 350, and an operation and maintenance evaluation module 360; wherein,
[0102] The dataset acquisition module 310 is used to acquire the abnormal dataset of the business system; the abnormal dataset includes: abnormal alarm data and monitoring data.
[0103] Root cause analysis module 320 is used to perform root cause analysis on the abnormal dataset and obtain root cause analysis results;
[0104] Resource configuration module 330 is used to determine the resource configuration strategy corresponding to the executable operation and maintenance program based on the root cause analysis results;
[0105] Program generation module 340 is used to generate an executable operation and maintenance program based on the abnormal dataset, the root cause analysis results and the program template;
[0106] The program execution module 350 is used to execute the executable operation and maintenance program in a sandbox isolation environment according to the resource configuration policy, and record the status data of the executable operation and maintenance program during the execution process;
[0107] The operation and maintenance assessment module 360 is used to generate an operation and maintenance strategy assessment report based on the root cause analysis results, the executable operation and maintenance program, and the status data.
[0108] Optional, the root cause analysis module 320 is specifically used for:
[0109] The target determination unit is used to perform semantic parsing and target mapping on the abnormal dataset by combining the domain knowledge graph to obtain the target to be analyzed.
[0110] The reasoning unit is used to perform anomaly reasoning based on the target to be analyzed and the abnormal dataset, using a reasoning mechanism that combines forward and backward chains, to obtain root cause analysis results; the root cause analysis results include anomaly root causes and associated evidence chains.
[0111] Optionally, the inference unit is specifically used for:
[0112] Starting from the known facts in the abnormal dataset and the target to be analyzed, the root cause of the anomaly is derived by applying rules in the rule base.
[0113] Starting from the root cause of the anomaly, we search backwards from the fact base for a chain of evidence supporting the root cause of the anomaly.
[0114] Optionally, the program generation module 340 is specifically used for:
[0115] Based on multi-dimensional features and reinforcement learning strategies, the target program template is determined from the program template library according to the anomaly type of the anomaly root cause.
[0116] Extract target parameters from the abnormal dataset and the evidence chain;
[0117] Input the target parameters into the target program template to generate initial operation and maintenance code;
[0118] The initial operation and maintenance code is subjected to a syntax check, and the initial operation and maintenance code that passes the syntax check is determined to be executable operation and maintenance code.
[0119] Optionally, the resource configuration module 330 includes:
[0120] Resource requirement information is determined based on the anomaly type of the root cause of the anomaly;
[0121] Based on the resource requirement information, predict the load status of multiple execution nodes executing the executable operation and maintenance program;
[0122] The target execution node is determined based on the load status of the multiple execution nodes.
[0123] Optionally, the program execution module 350 is specifically used for:
[0124] A sandbox isolation environment is constructed based on the resource allocation strategy, wherein the sandbox isolation environment includes resource isolation and process isolation;
[0125] Execute the executable operation and maintenance program in the sandbox isolation environment;
[0126] Record the status data during program execution. The status data includes status data and performance status data. The status data includes variable data, execution results, and execution trajectory of key nodes. The performance status data includes resource utilization and response time.
[0127] The operation and maintenance strategy evaluation device provided in the embodiments of the present invention can execute the operation and maintenance strategy evaluation method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0128] Example 4
[0129] Figure 4 A schematic diagram of an electronic device 10, which can be used to implement embodiments of the present invention, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0130] like Figure 4As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0131] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0132] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as operational strategy evaluation methods.
[0133] In some embodiments, the operation and maintenance strategy evaluation method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the operation and maintenance strategy evaluation method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the operation and maintenance strategy evaluation method by any other suitable means (e.g., by means of firmware).
[0134] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0135] In some embodiments, the operation and maintenance strategy evaluation method can be implemented as a computer program, which is implicitly included in a computer program product. When executed by a processor, the computer program implements the operation and maintenance strategy evaluation method of the present invention. The computer program product can be understood as a software product that primarily implements its solution through a computer program. The computer program used to implement the method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer program causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer program can be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a standalone software package, or entirely on a remote machine or server.
[0136] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0137] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0138] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0139] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0140] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0141] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for evaluating operation and maintenance strategies, characterized in that, Obtain the exception dataset from the business system; The abnormal dataset includes: abnormal alarm data and monitoring data; Root cause analysis was performed on the abnormal dataset to obtain the root cause analysis results; An executable operation and maintenance program is generated based on the abnormal dataset, the root cause analysis results, and the target program template. Based on the root cause analysis results, determine the resource configuration strategy corresponding to the executable operation and maintenance program; The executable operation and maintenance program is executed in a sandbox isolation environment according to the resource configuration policy, and the status data of the executable operation and maintenance program during the execution process is recorded. An operation and maintenance strategy evaluation report is generated based on the root cause analysis results, the executable operation and maintenance program, and the status data.
2. The method according to claim 1, characterized in that, Root cause analysis was performed on the abnormal dataset to obtain the root cause analysis results, including: The abnormal dataset is semantically parsed and target mapped using a domain knowledge graph to obtain the target to be analyzed. Based on the target to be analyzed and the abnormal dataset, anomaly reasoning is performed using a reasoning mechanism that combines forward and backward chains to obtain root cause analysis results; the root cause analysis results include abnormal root causes and associated evidence chains.
3. The method according to claim 2, characterized in that, The step of performing anomaly inference based on the target to be analyzed and the abnormal dataset, using a reasoning mechanism combining forward and backward chains, to obtain root cause analysis results includes: Starting from the known facts in the abnormal dataset and the target to be analyzed, the root cause of the anomaly is derived by applying rules in the rule base. Starting from the root cause of the anomaly, we search backwards from the fact base for a chain of evidence supporting the root cause of the anomaly.
4. The method according to claim 3, characterized in that, The step of generating an executable operation and maintenance program based on the abnormal dataset, the root cause analysis results, and the target program template includes: Based on multi-dimensional features and reinforcement learning strategies, the target program template is determined from the program template library according to the anomaly type of the anomaly root cause. Extract target parameters from the abnormal dataset and the evidence chain; Input the target parameters into the target program template to generate initial operation and maintenance code; The initial operation and maintenance code is subjected to a syntax check, and the initial operation and maintenance code that passes the syntax check is determined to be executable operation and maintenance code.
5. The method according to claim 3, characterized in that, The step of determining the resource configuration strategy corresponding to the executable operation and maintenance program based on the root cause analysis results includes: Resource requirement information is determined based on the anomaly type of the root cause of the anomaly; Based on the resource requirement information, predict the load status of multiple execution nodes executing the executable operation and maintenance program; The target execution node is determined based on the load status of the multiple execution nodes.
6. The method according to claim 1, characterized in that, The executable operation and maintenance program is executed in a sandbox isolation environment according to the resource configuration policy, and the status data of the executable operation and maintenance program during execution is recorded, including: A sandbox isolation environment is constructed based on the resource allocation strategy, wherein the sandbox isolation environment includes resource isolation and process isolation; Execute the executable operation and maintenance program in the sandbox isolation environment; Record the status data during program execution. The status data includes status data and performance status data. The status data includes variable data, execution results, and execution trajectory of key nodes. The performance status data includes resource utilization and response time.
7. An operation and maintenance strategy evaluation device, characterized in that, include: The dataset acquisition module is used to acquire the abnormal dataset of the business system; The abnormal dataset includes: abnormal alarm data and monitoring data; The root cause analysis module is used to perform root cause analysis on the abnormal dataset and obtain the root cause analysis results. The program generation module is used to generate an executable operation and maintenance program based on the abnormal dataset, the root cause analysis results, and the program template. The resource configuration module is used to determine the resource configuration strategy corresponding to the executable operation and maintenance program based on the root cause analysis results. The program execution module is used to execute the executable operation and maintenance program in a sandbox isolation environment according to the resource configuration policy, and record the status data of the executable operation and maintenance program during the execution process; The operation and maintenance assessment module is used to generate an operation and maintenance strategy assessment report based on the root cause analysis results, the executable operation and maintenance program, and the status data.
8. An electronic device, characterized in that, The electronic device includes: At least one processor; and a memory communicatively connected to the at least one processor; The memory stores a computer program that can be executed by the at least one processor, which is then executed by the at least one processor to enable the at least one processor to perform the operation and maintenance strategy evaluation method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that are used to cause a processor to execute the operation and maintenance strategy evaluation method according to any one of claims 1-6.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the operation and maintenance strategy evaluation method as described in any one of claims 1-6.