Operation instruction execution method and device, storage medium and electronic equipment
Through the automated operation instruction execution method, abnormal behavior recognition, problem diagnosis and instruction generation models are used to solve the problems of low efficiency and low accuracy in Kubernetes cluster operation and maintenance, and efficient and accurate automated operation and maintenance are achieved.
Patent Information
- Application Number
- CN202510137568.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-27
AI Technical Summary
During the operation and maintenance of Kubernetes clusters, manual operation and maintenance efficiency are low and the accuracy of execution results is low, resulting in low execution efficiency and poor accuracy of operation instructions.
By obtaining the original container operation and maintenance data generated by Kubernetes during the container operation and maintenance process, preprocessing it to obtain standard data; then, based on the preset abnormal behavior recognition model, problem diagnosis model and instruction generation model, the data is abnormal behavior recognition, problem diagnosis and operation instructions generated, and these instructions are automatically executed to solve abnormal problems.
It realizes automatic execution of operation instructions, improves the execution efficiency of operation instructions, reduces the dependence of manual operation and maintenance, and enhances the accuracy of execution results.
Smart Images

Figure CN120045287A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of computer technologies, and in particular, to a method for executing operation instructions, an apparatus for executing operation instructions, a computer-readable storage medium, and an electronic device. Background Art
[0002] In some related operation and maintenance solutions for Kubernetes clusters, it can be achieved manually. However, during the manual operation and maintenance process, since manual monitoring, analysis, and problem-solving are required, there are problems of low execution efficiency of operation instructions and low accuracy rate of execution results of operation instructions.
[0003] It should be noted that the information disclosed in the above background art is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0004] The purpose of the present disclosure is to provide a method for executing operation instructions, an apparatus for executing operation instructions, a computer-readable storage medium, and an electronic device, so as to at least to some extent overcome the problems of low execution efficiency of operation instructions and low accuracy rate of execution results of operation instructions caused by the limitations and defects of related technologies.
[0005] According to one aspect of the present disclosure, there is provided a method for executing operation instructions, including:
[0006] Obtain the original container operation and maintenance data generated during the container operation and maintenance process of Kubernetes, and preprocess the original container operation and maintenance data to obtain standard container operation and maintenance data;
[0007] Based on a preset abnormal behavior recognition model, perform abnormal behavior recognition on the standard container operation and maintenance data to obtain an abnormal behavior recognition result during the container operation and maintenance process of Kubernetes;
[0008] Based on a preset problem diagnosis model, perform problem diagnosis and prediction on the abnormal behavior recognition result to obtain a predicted result of abnormal problems that will occur during the container operation and maintenance process of Kubernetes;
[0009] Input the predicted result of abnormal problems into a preset instruction generation model to obtain operation instructions for solving the abnormal problems, and execute the operation instructions to solve the abnormal problems.
[0010] In an exemplary embodiment of the present disclosure, obtaining the original container operation and maintenance data generated by Kubernetes during the container operation and maintenance process includes: calling the application programming interface service of Kubernetes and obtaining the original cluster status data generated by Kubernetes during the container operation and maintenance process from the application programming interface service; wherein, the original cluster status data includes at least one of container status data, node status data, and node resource data; calling the status monitoring tool of the Kubernetes and obtaining the original monitoring metric data generated by the Kubernetes during the container operation and maintenance process from the status monitoring tool; wherein, the original monitoring metric data includes CPU usage rate and / or memory usage rate; calling the log collection tool of the Kubernetes and obtaining the original log data generated by the Kubernetes during the container operation and maintenance process from the log collection tool.
[0011] In an exemplary embodiment of the present disclosure, the preset abnormal behavior recognition model includes an embedding mapping layer, an encoding layer, and a mixture of experts model layer;
[0012] Among them, performing abnormal behavior recognition on the standard container operation and maintenance data based on the preset abnormal behavior recognition model to obtain the abnormal behavior recognition result of Kubernetes during the container operation and maintenance process includes: generating the first basic information to be predicted according to the standard cluster status data and standard monitoring metric data in the standard container operation and maintenance data, and generating the first context information to be predicted according to the standard log data in the standard container operation and maintenance data and the preset first model prompt parameters; performing embedding mapping processing on the first basic information to be predicted based on the embedding mapping layer to obtain the first operation and maintenance feature, and performing embedding mapping processing on the first context information to be predicted based on the embedding mapping layer to obtain the first context flag sequence; performing encoding processing on the first operation and maintenance feature and the first context flag sequence based on the encoding layer to obtain the first context overall representation, and performing prediction on the first context flag sequence and the first context overall representation based on the mixture of experts model layer to obtain the abnormal behavior recognition result of Kubernetes during the container operation and maintenance process.
[0013] In an exemplary embodiment of the present disclosure, the mixture of experts model layer includes a gated network model and an expert neural network model;
[0014] Among them, based on the hybrid expert model layer, predictions are made on the first context flag sequence and the first overall context representation to obtain the abnormal behavior recognition result during the container operation and maintenance of Kubernetes, including: determining the model weights of the expert neural network model based on the first context flag sequence according to the gating network model; determining the target neural network model required for predicting the abnormal behavior recognition result from the expert neural network according to the model weights; inputting the first context flag sequence and the first overall context representation into the target neural network model to obtain the abnormal behavior recognition result during the container operation and maintenance of Kubernetes; wherein, the abnormal behavior recognition result includes at least one of the following: the restart frequency of the containers included in Kubernetes is greater than a first preset threshold, the CPU usage rate of the nodes included in Kubernetes is greater than a second preset threshold, and the network latency of the nodes is greater than or equal to a third preset threshold.
[0015] In an exemplary embodiment of the present disclosure, the preset abnormal behavior recognition model is obtained in the following manner: acquiring historical container operation and maintenance data generated during the container operation and maintenance of Kubernetes and the actual behavior labels of Kubernetes under the historical container operation and maintenance data; training a low-rank adaptation model based on the historical container operation and maintenance data and the actual behavior labels to obtain the low-rank matrix parameters of Kubernetes in the container operation and maintenance scenario; fine-tuning a large language model based on the low-rank matrix parameters to obtain the preset abnormal behavior recognition model.
[0016] In an exemplary embodiment of the present disclosure, based on a preset problem diagnosis model, problem diagnosis and prediction are performed on the abnormal behavior recognition result to obtain the abnormal problem prediction result that will occur during the container operation and maintenance of Kubernetes, including: generating second basic information to be predicted according to the abnormal behavior recognition result and generating second context information to be predicted according to preset second model prompt parameters; inputting the second basic information to be predicted and the second context information to be predicted into the preset problem diagnosis model to obtain the abnormal problem prediction result that will occur during the container operation and maintenance of Kubernetes; wherein, the abnormal problem prediction result includes at least one of the following: the restart frequency of the container is greater than a first preset threshold, and the container cannot work properly, resulting in the inability to execute tasks normally; the CPU usage rate of the nodes included in Kubernetes is greater than a second preset threshold, and the resources are about to be exhausted, resulting in the inability to execute tasks normally; the network latency of the nodes is greater than a third preset threshold, resulting in a time delay in task execution.
[0017] In an exemplary embodiment of the present disclosure, inputting the abnormal problem prediction result into a preset instruction generation model to obtain an operation instruction for solving the abnormal problem includes: generating third basic information to be predicted according to the abnormal problem prediction result, and generating third context information to be predicted according to preset third model prompt parameters; inputting the third basic information to be predicted and the third context information to be predicted into the preset instruction generation model to obtain an operation instruction for solving the abnormal problem; wherein, the operation instruction includes at least one of the following: restarting the container, expanding the capacity of the node, and allocating bandwidth resources to the node.
[0018] In an exemplary embodiment of the present disclosure, executing the operation instruction to solve the abnormal problem includes: determining the instruction category of the operation instruction, and matching an instruction execution tool for executing the operation instruction according to the instruction category; wherein, the instruction tool includes at least one of a container restart tool, a computing resource allocation tool, and a bandwidth resource allocation tool; calling the instruction execution tool, and controlling the instruction execution tool to execute the operation instruction to solve the abnormal behavior.
[0019] In an exemplary embodiment of the present disclosure, the execution method of the operation instruction further includes:
[0020] In response to a touch operation on an interactive control on an interactive interface corresponding to a Kubernetes container orchestration cluster, determining a target interactive control and display interface information corresponding to the target interactive control; wherein, the display interface information includes at least one of the running state information of the container orchestration cluster, the operation instruction to be executed, the instruction execution result corresponding to the operation instruction, and other custom configuration information;
[0021] Obtaining the display interface information and displaying the display interface information to facilitate viewing of the running state information of the container orchestration cluster and / or the operation instruction to be executed and / or the instruction execution result corresponding to the operation instruction and / or other custom configuration information.
[0022] According to one aspect of the present disclosure, there is provided an execution device for an operation instruction, including:
[0023] A data preprocessing module, configured to obtain original container operation and maintenance data generated during the container operation and maintenance process of Kubernetes, and preprocess the original container operation and maintenance data to obtain standard container operation and maintenance data;
[0024] An abnormal behavior recognition module, configured to perform abnormal behavior recognition on the standard container operation and maintenance data based on a preset abnormal behavior recognition model to obtain an abnormal behavior recognition result during the container operation and maintenance process of Kubernetes;
[0025] A problem diagnosis and prediction module, configured to perform problem diagnosis and prediction on the abnormal behavior recognition result based on a preset problem diagnosis model, so as to obtain a prediction result of abnormal problems that will occur during the container operation and maintenance of Kubernetes;
[0026] An operation instruction execution module, configured to input the prediction result of the abnormal problem into a preset instruction generation model to obtain an operation instruction for solving the abnormal problem, and execute the operation instruction to solve the abnormal problem.
[0027] According to one aspect of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the execution method of the operation instruction described in any one of the above is implemented.
[0028] According to one aspect of the present disclosure, there is provided an electronic device, including:
[0029] A processor; and
[0030] A memory, configured to store executable instructions of the processor;
[0031] Wherein, the processor is configured to execute the execution method of the operation instruction described in any one of the above by executing the executable instructions.
[0032] An execution method of an operation instruction provided by an embodiment of the present disclosure. On the one hand, by acquiring original container operation and maintenance data generated during the container operation and maintenance of Kubernetes, and preprocessing the original container operation and maintenance data to obtain standard container operation and maintenance data; then, based on a preset abnormal behavior recognition model, performing abnormal behavior recognition on the standard container operation and maintenance data to obtain an abnormal behavior recognition result during the container operation and maintenance of Kubernetes; furthermore, based on a preset problem diagnosis model, performing problem diagnosis and prediction on the abnormal behavior recognition result to obtain a prediction result of abnormal problems that will occur during the container operation and maintenance of Kubernetes; finally, inputting the prediction result of the abnormal problem into a preset instruction generation model to obtain an operation instruction for solving the abnormal problem, and executing the operation instruction to solve the abnormal problem, which realizes the automatic execution of the operation instruction, solves the problem of low execution efficiency caused by the need to execute the operation instruction manually in the related art, and improves the execution efficiency of the operation instruction; on the other hand, it also solves the problem of low accuracy of the execution result caused by the need to execute the operation instruction manually in the related art.
[0033] It should be understood that the above general description and subsequent detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Description of the Drawings
[0034] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments in line with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other accompanying drawings based on these drawings without creative efforts.
[0035] Figure 1 A flowchart schematically showing a method for executing an operation instruction according to an exemplary embodiment of the present disclosure.
[0036] Figure 2 A schematic structural diagram showing a Kubernetes container orchestration cluster according to an exemplary embodiment of the present disclosure.
[0037] Figure 3 A schematic diagram showing nodes included in a Kubernetes container orchestration cluster and container Pods included in the nodes according to an exemplary embodiment of the present disclosure.
[0038] Figure 4 A schematic block diagram showing a structure of an operation instruction execution system according to an exemplary embodiment of the present disclosure.
[0039] Figure 5 A schematic structural diagram showing a preset abnormal behavior recognition model according to an exemplary embodiment of the present disclosure.
[0040] Figure 6 A flowchart of a method for fine-tuning a preset abnormal behavior recognition model according to an exemplary embodiment of the present disclosure.
[0041] Figure 7 A schematic structural diagram showing an operation instruction execution device according to an exemplary embodiment of the present disclosure.
[0042] Figure 8 A schematic diagram showing an electronic device for implementing a method for executing an operation instruction according to an exemplary embodiment of the present disclosure. Detailed implementation manners
[0043] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the concept of example embodiments to those skilled in the art. The features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will recognize that the technical solutions of the present disclosure may be practiced without one or more of the specific details, or may be implemented using other methods, components, devices, steps, etc. In other cases, well-known technical solutions are not shown or described in detail so as not to obscure the various aspects of the present disclosure.
[0044] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus their repeated description will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0045] Currently, the operation and maintenance of Kubernetes clusters face challenges such as high complexity, reliance on manual experience, and difficulty in coping with emergencies. Specifically, the Kubernetes operation and maintenance described herein involve multiple aspects such as monitoring the cluster status, deploying and managing applications, adjusting resource configurations, and handling faults. Usually, it requires operation and maintenance personnel to use the kubectl command to view the monitoring panel (such as Prometheus or Grafana, etc.) and analyze logs to complete; at the same time, Prometheus and Grafana involved herein are common Kubernetes monitoring tools that can collect and visualize various metric data of the cluster to help operation and maintenance personnel understand the system status.
[0046] Specifically, as the scale of the cluster expands and the complexity of applications increases, operation and maintenance personnel need to process a large amount of monitoring data, log information, and events in order to discover and solve problems in a timely manner. However, traditional monitoring and alerting mechanisms based on rules or thresholds are prone to false alarms and missed alarms, making it difficult to conduct in-depth root cause analysis and predictive maintenance. In addition, the writing and maintenance costs of automated operation and maintenance scripts are relatively high, and it is difficult to adapt to constantly changing operation and maintenance scenarios. To solve this problem, some related technical solutions have proposed operation and maintenance solutions based on AI (Artificial Intelligence). In this operation and maintenance solution, machine learning models can be used for anomaly detection, fault prediction, capacity planning, etc. to implement specific operation and maintenance processes. However, these solutions usually require training models for specific scenarios, have poor generality, and require a high threshold for AI technology.
[0047] In addition, there are currently some rule-based automated operation and maintenance tools. For example, by configuring corresponding alert rules, an alert is issued when the system metrics exceed the thresholds configured in the alert rules. However, these rules need to be manually defined, cannot automatically adapt to system changes, and are difficult to handle complex problems. Further, although some solutions have tried to use machine learning models for anomaly detection, the machine learning models they use are not large models. At the same time, during the model training process, features need to be manually selected for model training, and the generalization ability of the obtained model is also weak. Furthermore, due to the weak generalization ability of the obtained model, the obtained model usually can only analyze specific metrics, cannot consider the overall situation, and is even more difficult to conduct root cause analysis of problems. On this basis, although some cloud providers also provide anomaly detection services based on machine learning, their models are usually black boxes and are difficult to optimize specifically. It can be seen from this that the operation and maintenance solutions for Kubernetes clusters described above have the following defects:
[0048] On the one hand, manual operation and maintenance is inefficient, that is, because it requires manual monitoring, analysis, and problem-solving, there are problems of time-consuming, laborious, and easy to make mistakes. On the other hand, the operation and maintenance threshold is high, that is, because it requires professional Kubernetes knowledge and AI technology, the operation and maintenance cost is increased. On the other hand, the rules are fixed and it is difficult to handle complex problems, that is, rule-based automated solutions are difficult to adapt to dynamic environments and cannot handle unknown problems. Further, the generality of AI models is poor, that is, machine learning-based solutions usually require training models for specific scenarios, cannot be reused, and require a large amount of labeled data. Finally, the interpretability of the model is poor, that is, the decision-making process of existing machine learning models is not transparent enough to be understood, and it is difficult for operation and maintenance personnel to trust.
[0049] Based on this, the exemplary embodiments of the present disclosure provide a method for executing operation instructions. This method for executing operation instructions can utilize artificial intelligence technology, specifically an AI (Artificial Intelligence) Agent composed of large models (i.e., an AI intelligent agent), to improve the intelligent level of Kubernetes cluster operation and maintenance, and improve operation and maintenance efficiency and the reliability of the cluster on the basis of reducing operation and maintenance costs. Further, in the actual application process, the exemplary embodiments of the present disclosure can fine-tune a general large language model (LLM) in the Kubernetes scenario, and then build a Kubernetes Agent according to the fine-tuned large language model and prompt engineering; on this premise, operation instructions can be generated and executed through the Kurbernetes Agent.
[0050] Further, in the process of generating operation instructions, first, analyze and understand the logs and monitoring of Kuberentes, and then automatically detect whether an anomaly occurs according to the analysis results; then, automatically match the operation and maintenance tools (Agent Tool) required for the problem and execute the corresponding operation instructions based on the operation and maintenance tools; finally, perform a backtest on the returned results to determine whether the execution meets the expectations; in the actual application process, since the whole process is automated, it can reduce manual intervention on the basis of improving the operation and maintenance efficiency of Kubernetes; moreover, the large model Agent can also lower the operation and maintenance threshold of Kubernetes, enabling personnel without professional Kubernetes knowledge to perform simple operation and maintenance management, realizing the automation and intelligence of Kubernetes operation and maintenance, and can automatically detect, predict, and solve corresponding problems, providing a general and flexible AI operation and maintenance solution, and can also adapt to different Kubernetes environments; on this basis, due to the inherent generality of the LLM, it can improve the interpretability of problem diagnosis and resource optimization, thereby improving operation and maintenance efficiency and reliability.
[0051] In the present exemplary embodiment, a method for executing operation instructions is first provided. This method can run on a terminal device, a server, a server cluster, a cloud server, etc. where Kurbernetes is located; of course, those skilled in the art can also run the method of the present disclosure on other platforms according to requirements, and no special limitation is made in this exemplary embodiment. Specifically, referring to Figure 1 as shown, the method for executing operation instructions may include the following steps:
[0052] Step S110. Obtain the original container operation and maintenance data generated during the container operation and maintenance of Kubernetes, and preprocess the original container operation and maintenance data to obtain standard container operation and maintenance data;
[0053] Step S120. Based on a preset abnormal behavior recognition model, perform abnormal behavior recognition on the standard container operation and maintenance data to obtain the abnormal behavior recognition result during the container operation and maintenance of Kubernetes;
[0054] Step S130. Based on a preset problem diagnosis model, perform problem diagnosis and prediction on the abnormal behavior recognition result to obtain the abnormal problem prediction result that will occur during the container operation and maintenance of Kubernetes;
[0055] Step S140. Input the abnormal problem prediction result into a preset instruction generation model to obtain an operation instruction for solving the abnormal problem, and execute the operation instruction to solve the abnormal problem.
[0056] In the above-described method for executing the operation instruction, on the one hand, by obtaining the original container operation and maintenance data generated during the container operation and maintenance of Kubernetes, and preprocessing the original container operation and maintenance data to obtain standard container operation and maintenance data; then, based on a preset abnormal behavior recognition model, perform abnormal behavior recognition on the standard container operation and maintenance data to obtain the abnormal behavior recognition result during the container operation and maintenance of Kubernetes; furthermore, based on a preset problem diagnosis model, perform problem diagnosis and prediction on the abnormal behavior recognition result to obtain the abnormal problem prediction result that will occur during the container operation and maintenance of Kubernetes; finally, input the abnormal problem prediction result into a preset instruction generation model to obtain an operation instruction for solving the abnormal problem, and execute the operation instruction to solve the abnormal problem, the automatic execution of the operation instruction is realized, and the problem of low execution efficiency caused by the need to execute the operation instruction manually in the related art is solved, and the execution efficiency of the operation instruction is improved; on the other hand, the problem of low accuracy of the execution result caused by the need to execute the operation instruction manually in the related art is also solved.
[0057] Next, the present disclosure will be described in conjunction with the accompanying drawings for the example embodiments described herein.
[0058] Kubernetes Container Orchestration Cluster: Kubernetes provides a powerful mechanism based on which automated deployment, scaling, and management of containerized applications can be achieved. In a Kubernetes container orchestration cluster, core concepts can include, but are not limited to, Pod (container application), Service (service), and Deployment (deployment), etc. During actual application, cluster management can be carried out through components in the Control Plane such as API Server (application programming interface service), Scheduler (scheduler), Controller Manager (control manager), etc. Further, monitoring data collection can be achieved through tools such as Metrics-Server and Prometheus; log data collection can be achieved through tools such as Fluentd and Elasticsearch, etc. Event recording is the basic data source for Kubernetes operation and maintenance. Even further, in a Kubernetes container orchestration cluster, there can also be multiple computing nodes, and each computing node can include one or more containers. For specific example diagrams, reference can be made to Figure 2 and Figure 3 as shown
[0059] Artificial Intelligence and Deep Learning: In recent years, significant progress has been made in artificial intelligence and deep learning technologies, especially in fields such as natural language processing, time series analysis, and anomaly detection. Further, the emergence of large models such as the Transformer model enables machines to better understand and process complex data patterns, providing new possibilities for intelligent operation and maintenance.
[0060] AI Agent Technology: An AI Agent is an intelligent entity that can perceive the environment, perform reasoning, and take actions. In the software field, it refers to a deep learning model trained based on a large-scale dataset, which has powerful natural language understanding, generation, reasoning, and decision-making capabilities and can autonomously execute tasks like an intelligent entity (such as generating operation instructions and executing operation instructions). An AI Agent can be designed as an intelligent program for automatically executing specific tasks or assisting users in decision-making. In the exemplary embodiments of the present disclosure, it specifically refers to an intelligent software entity for analyzing Kubernetes operation and maintenance data, understanding operation and maintenance intentions, generating operation and maintenance strategies, and executing operation and maintenance operations.
[0061] LLM is the abbreviation of Large Language Model, which is a natural language processing model based on deep learning. By training on a large amount of text data, LLM can learn the structure and semantics of language, and is capable of generating, understanding, and processing complex natural language content. In actual applications, the large language model can adapt to multiple languages and tasks without having to redesign the model for each task. It can handle complex problems and generate natural and fluent text.
[0062] Secondly, the technical implementation principle of the exemplary embodiments of the present disclosure will be explained and described.
[0063] A method for executing operation instructions recorded in the exemplary embodiments of the present disclosure can be used to achieve intelligent operation and maintenance of the Kubernetes container orchestration cluster based on the large model AI Agent. In actual applications, the pre-trained large model AI Agent can be used to perform deep learning and analysis on various operation and maintenance data of the Kubernetes cluster, so as to understand the operation and maintenance intent of the container orchestration cluster and predict potential risks, and further achieve the purpose of assisting in fault diagnosis and finally realizing automated operation and maintenance operations. Specifically, data such as metrics (monitoring data), logs (log data), and events (status information) of the container orchestration cluster can be obtained, and then these data are input into the large model AI Agent for processing to achieve automatic execution of operation instructions. On this basis, the AI Agent can identify abnormal patterns, locate the root cause of faults, and predict resource bottlenecks based on its powerful understanding and reasoning capabilities, and generate corresponding operation and maintenance strategies and operation instructions. Finally, the execution module is responsible for executing these instructions to achieve automated operation and maintenance.
[0064] Furthermore, the operation instruction execution system involved in the exemplary embodiments of the present disclosure will be explained and described. Specifically, refer to Figure 4As shown in the figure, the execution system of the operation instruction may include a data acquisition component 410, an intelligent agent 420 for generating operation instructions, an operation instruction execution component 430, and a user interaction component 440. Among them, the data acquisition component is responsible for collecting monitoring data, log data, configuration information, etc. from the API Server (Application Programming Interface Server) in Kubernetes, the status monitoring tool Prometheus, and the log system; the intelligent agent for generating operation instructions (i.e., AI Agent) can be used to analyze data, detect anomalies, predict problems, generate solutions, and execute operation instructions; the operation instruction execution component is the Agent tool set, which contains the developed and encapsulated Kubernetes machine operation and maintenance tools; in the actual application process, the operation instruction execution component can be used to execute the operation instructions generated by the AI Agent, such as querying resources, adjusting resources through the Kubernetes API, deploying downstream troubleshooting tools, and executing commands in nodes or containers; the user interaction component can be used to provide a user-interactive interface, so that the user can view relevant information such as system status, AI Agent analysis results, execution operations, and custom configurations based on the user-interactive interface.
[0065] In the actual application process, the data collection component continuously collects data from the Kubernetes system and sends it to the AI Agent; then, the AI Agent analyzes the data for anomaly detection and problem prediction, generates corresponding operation instructions, and sends the operation instructions to the operation instruction execution component; after receiving the operation instructions, the operation instruction execution component calls the Kubernetes API to execute the corresponding operations; finally, the user interaction component displays the system status and the analysis results of the AI Agent and allows the user to interact. It should be noted here that the focus of this disclosure is on how to generate operation instructions and how to execute operation instructions; on this premise, in order to generate operation instructions, it is first necessary to build an AI Agent; specifically, in the process of building the AI Agent, it is first necessary to select a suitable pre-trained large model; then, fine-tune and build a Prompt project for the Kubernetes operation and maintenance scenario to enable it to understand Kubernetes data, analyze problems, and generate operation instructions; further, in the process of generating operation instructions, it is also necessary to develop a data interface to obtain monitoring data, logs, and configuration information from the Kubernetes API Server, Prometheus, log system, etc. based on this data interface; furthermore, in the process of executing instructions, it is also necessary to develop tools for executing Kubernetes operations, such as deploying applications, adjusting resources, restarting Pods, etc., and provide an API for the AI Agent to call; finally, in the process of user interaction, it is also necessary to develop a user interaction interface to allow the user to view the analysis results of the AI Agent, execute operations, and customize configurations; and, in the process of developing tools for executing Kubernetes operations, it is also necessary to provide Prompt customization and function interface descriptions for each tool so that the LLM Agent can understand the execution scenario, function, and execution expected results of each Agent tool.
[0066] Hereinafter, the agents involved in the exemplary embodiments of the present disclosure will be explained and described. Specifically, the agents involved in the exemplary embodiments of the present disclosure may include a preset abnormal behavior recognition model, a preset problem diagnosis model, and a preset instruction generation model; in the actual application process, the model structures of the preset abnormal behavior recognition model, the preset problem diagnosis model, and the preset instruction generation model are the same and can all be implemented based on a large language model; at the same time, in the process of specific model fine-tuning, the training data set used can be set according to actual needs, and this example does not make special restrictions on this.
[0067] Hereinafter, taking the preset abnormal behavior recognition model as an example, the specific model structure will be explained and described. Specifically, refer to Figure 5As shown in the figure, the preset abnormal behavior recognition model may include a first input layer 501, a first embedding mapping layer 502, a first encoding layer 503, a first mixture-of-experts model layer 504, and a first output layer 505; among them, the mixture-of-experts model layer described here may include a gating network model and an expert neural network model; at the same time, the implementation functions of each model layer in the specific data processing process will be described in detail later, and no further elaboration will be made here. It should also be added that the model structures of the preset problem diagnosis model and the preset instruction generation model are similar to the model result of the preset abnormal behavior recognition model, and no further elaboration will be made here.
[0068] Hereinafter, taking the preset abnormal behavior recognition model as an example, the specific model fine-tuning process will be explained and described. Specifically, refer to Figure 6 As shown in the figure, the fine-tuning process of the preset abnormal behavior recognition model may include the following steps:
[0069] Step S610, obtaining historical container operation and maintenance data generated by Kubernetes during the container operation and maintenance process and the actual behavior labels of Kubernetes under the historical container operation and maintenance data;
[0070] Step S620, training the low-rank adaptation model based on the historical container operation and maintenance data and the actual behavior labels to obtain the low-rank matrix parameters of Kubernetes in the container operation and maintenance scenario;
[0071] Step S630, fine-tuning the large language model based on the low-rank matrix parameters to obtain the preset abnormal behavior recognition model.
[0072] Hereinafter, steps S610 - S630 will be explained and described. Specifically, the large language model described here is also the LLM; in the actual application process, in order to obtain the abnormal behavior recognition model required in this embodiment, the LLM can be fine-tuned and trained by models such as LoRA (i.e., Low-Rank Adaptation) to obtain the abnormal behavior recognition model; at the same time, the historical container operation and maintenance data described here refers to the operation and maintenance data of the Kubernetes container orchestration cluster in a certain period of the past; the actual behavior labels described here refer to the actual behaviors of the containers in the Kubernetes container orchestration cluster when executing tasks under the premise of the existence of historical container operation and maintenance data; for example, the container runs normally or is restarted or computing resources or bandwidth resources are added to the container, etc.
[0073] Hereinafter, in combination with Figures 2 - 6 For Figure 1Further explain and illustrate the execution method of the operation instructions shown in the figure. Specifically:
[0074] In step S110, obtain the original container operation and maintenance data generated by Kubernetes during container operation and maintenance, and preprocess the original container operation and maintenance data to obtain standard container operation and maintenance data.
[0075] In this exemplary embodiment, first, obtain the original container operation and maintenance data generated by Kubernetes during container operation and maintenance; specifically, it can be implemented in the following way: call the application programming interface service of Kubernetes, and obtain the original cluster status data generated by Kubernetes during container operation and maintenance from the application programming interface service; among them, the original cluster status data may include, but is not limited to, container status data (i.e., Pod status), node status data, and node resource data (i.e., the resource usage of nodes, etc.); call the status monitoring tool of Kubernetes, and obtain the original monitoring metric data generated by Kubernetes during container operation and maintenance from the status monitoring tool; among them, the original monitoring metric data includes CPU usage rate and / or memory usage rate; call the log collection tool of Kubernetes, and obtain the original log data generated by Kubernetes during container operation and maintenance from the log collection tool; among them, the log collection tool recorded here may include, but is not limited to, Prometheus or other monitoring tools or log systems, etc.
[0076] Secondly, preprocess the original container operation and maintenance data to obtain standard container operation and maintenance data; specifically, the preprocessing process recorded here may include, but is not limited to, security detection, legality detection, cleaning, filtering, filling, and normalization, etc.; in the actual application process, corresponding preprocessing can be performed according to the actual situation of the data, and this example does not make special restrictions on this.
[0077] In step S120, based on a preset abnormal behavior recognition model, perform abnormal behavior recognition on the standard container operation and maintenance data to obtain the abnormal behavior recognition result of Kubernetes during container operation and maintenance.
[0078] Specifically, the specific determination process of the abnormal behavior recognition result can be achieved through the following steps: First, generate the first basic information to be predicted based on the standard cluster status data and standard monitoring index data in the standard container operation and maintenance data, and generate the first context information to be predicted based on the standard log data in the standard container operation and maintenance data and the preset first model prompt parameters; Second, perform embedding mapping processing on the first basic information to be predicted based on the embedding mapping layer to obtain the first operation and maintenance feature, and perform embedding mapping processing on the first context information to be predicted based on the embedding mapping layer to obtain the first context flag sequence; Then, perform encoding processing on the first operation and maintenance feature and the first context flag sequence based on the encoding layer to obtain the first overall context representation, and perform prediction on the first context flag sequence and the first overall context representation based on the mixture of experts model layer to obtain the abnormal behavior recognition result of Kubernetes in the container operation and maintenance process. Specifically, the embedding mapping layer described here can include an Embedding embedding mapping layer and a Bert embedding mapping layer. During the embedding mapping process, the first basic information to be predicted can be processed by the Embedding embedding mapping layer to obtain the first operation and maintenance feature, and the first context information to be predicted can be processed by the Bert embedding mapping layer to obtain the first context flag sequence; At the same time, the encoding layer described here can be a bidirectional multi-layer Transformer, and of course it can also be other structures, and this example does not make special restrictions on this.
[0079] In an exemplary embodiment, based on the mixture-of-experts model layer, predicting the first context flag sequence and the first overall context representation to obtain the recognition result of abnormal behaviors during the container operation and maintenance of Kubernetes can be achieved in the following manner: Based on the gated network model, determining the model weights of the expert neural network model according to the first context flag sequence; determining the target neural network model required to perform the prediction task of the abnormal behavior recognition result from the expert neural network according to the model weights; inputting the first context flag sequence and the first overall context representation into the target neural network model to obtain the recognition result of abnormal behaviors during the container operation and maintenance of Kubernetes; wherein, the recognition result of abnormal behaviors may include, but is not limited to, that the restart frequency of containers included in Kubernetes is greater than a first preset threshold, the CPU usage rate of nodes included in Kubernetes is greater than a second preset threshold, and the network latency of nodes is greater than or equal to a third preset threshold, etc. That is to say, in the actual application process, since the large model has the ability to understand Kubernetes data through pre-training; therefore, the abnormal behaviors can be directly recognized according to the received data; wherein, the obtained abnormal behaviors may be, for example, frequent restart of Pods, abnormal increase in the CPU usage rate of nodes, and increase in the network latency of nodes, etc.; it should be further supplemented and explained here that the abnormal behaviors described here are illustrated by taking the container management scenario as an example. Correspondingly, it can also be applied to the automated deployment scenario and the container or node expansion scenario, and this example does not make special restrictions on this.
[0080] In step S130, based on a preset problem diagnosis model, diagnosing and predicting the abnormal behavior recognition result to obtain the predicted result of abnormal problems that will occur during the container operation and maintenance of Kubernetes.
[0081] Specifically, the specific determination process of the abnormal problem prediction result can be implemented in the following manner: generating the second basic information to be predicted according to the abnormal behavior recognition result, and generating the second context information to be predicted according to the preset second model prompt parameter; inputting the second basic information to be predicted and the second context information to be predicted into the preset problem diagnosis model to obtain the abnormal problem prediction result generated during the container operation and maintenance process of Kubernetes; wherein, the abnormal problem prediction result may include, but is not limited to: the restart frequency of the container is greater than the first preset threshold, and the container cannot work properly, resulting in the inability to execute tasks normally; the CPU usage rate of the nodes included in Kubernetes is greater than the second preset threshold, and the resources are about to be exhausted, resulting in the inability to execute tasks normally; the network latency of the nodes is greater than the third preset threshold, resulting in a time delay in task execution. Specifically, since the specific model structure of the preset problem diagnosis model recorded here is the same as the model structure of the preset abnormal behavior recognition model recorded above, the specific determination process of the abnormal problem prediction result is also generally similar to the specific determination process of the abnormal behavior recognition result recorded above, so it will not be further elaborated here. At the same time, the preset second model prompt parameter recorded here can be set according to actual needs, and this example does not make special restrictions on this. That is, in the actual application process, the preset problem diagnosis model first needs to conduct problem diagnosis. For example, it is necessary to determine whether the abnormality is caused by an application error, insufficient resources, or a network problem; further, after obtaining the diagnosis result, it is necessary to predict possible future problems according to the diagnosis result. For example, it is necessary to give an early warning that the CPU resources will be exhausted, resulting in the inability to execute tasks, or give an early warning that the container cannot work properly, resulting in the inability to execute tasks normally; or give an early warning that the network latency is too large, resulting in a time delay in task execution, etc.
[0082] In step S140, input the abnormal problem prediction result into the preset instruction generation model to obtain an operation instruction for solving the abnormal problem, and execute the operation instruction to solve the abnormal problem.
[0083] In the present exemplary embodiment, first, the abnormal problem prediction result is input into a preset instruction generation model to obtain an operation instruction for solving the abnormal problem; specifically, it can be implemented in the following manner: generate third basic information to be predicted according to the abnormal problem prediction result, and generate third context information to be predicted according to preset third model prompt parameters; input the third basic information to be predicted and the third context information to be predicted into the preset instruction generation model to obtain an operation instruction for solving the abnormal problem; wherein, the operation instruction includes at least one of the following: restart the container, expand the resources of the node, and allocate bandwidth resources for the node. Specifically, since the specific model structure of the preset instruction generation model recorded here is the same as the model structure of the preset abnormal behavior recognition model recorded above, the specific determination process of the operation instruction is also generally similar to the specific determination process of the abnormal behavior recognition result recorded above, so it will not be further elaborated here. At the same time, the preset third model prompt parameters recorded here can be set according to actual needs, and no special restrictions are imposed in this example. Further, in the actual application process, if it is found that a certain Pod is frequently restarted, the AI Agent may suggest restarting the Pod; if it is found that the CPU utilization rate of the node is too high, resources can be expanded for the node; if it is found that the latency of the node is too long, more bandwidth resources can be allocated for the node, and so on.
[0084] In one exemplary embodiment, the preset first model prompt parameter recorded above can be: prompt: You are a Kubernetes intelligent operation and maintenance assistant, responsible for analyzing the running status of the Kubernetes cluster and taking measures when necessary. Currently, the CPU usage of pod {{pod_name}} has been continuously higher than 80%. Please use the following tools to analyze and give suggestions. The tools available to you include: 1. get_pod_metrics: Obtain the metric data of the Pod. 2. get_pod_logs: Obtain the log data of the Pod. 3. describe_pod: Obtain the description information of the Pod. 4. restart_pod: Restart the specified Pod. Please use the appropriate tools to analyze and give the final operation steps. And please output strictly in the following format: 1. [Analysis]: <Analysis of the problem>; 2. [Operation]: <The operation to be executed, including the tool called and the input parameters>; 3. [Final suggestion]: <The final suggestion, including whether manual intervention is required>.
[0085] Second, after obtaining the operation instruction, the operation instruction can be executed to solve the abnormal problem; specifically, the specific execution process of the operation instruction can be implemented in the following way: determine the instruction category of the operation instruction, and match and execute the instruction execution tool of the operation instruction according to the instruction category; among them, the instruction tools recorded here include container restart tools, computing resource allocation tools, bandwidth resource allocation tools, etc.; call the instruction execution tool, and control the instruction execution tool to execute the operation instruction to solve the abnormal behavior. Among them, matching and executing the instruction execution tool of the operation instruction according to the instruction category can be implemented in the following way: obtain the tool attribute information of the original tool, and determine the structured document of the original tool according to the tool attribute information; determine the tool instruction information, tool operation information and tool response information of the original tool according to the structured document, and generate a tool usage instance corresponding to the original tool according to the tool instruction information, tool operation information and tool response information; based on the instruction category and the tool usage instance, match and execute the instruction execution tool of the operation instruction from the original tools. That is, in the actual application process, when the instruction execution component receives the operation instruction of the AI Agent, it can call the corresponding instruction execution tool based on the Kubernetes API to execute the corresponding operation to solve the possible abnormal problems, so as to realize the automatic operation and maintenance of Kubernetes. It should also be supplemented here that in the process of executing the operation instruction, it can also be realized by sending a custom operation instruction to the Kubernetes cluster for operation, or by remotely executing Linux commands to execute the operation instruction. This example does not make special restrictions on this.
[0086] Hereinafter, the instruction execution tools involved in the exemplary embodiments of the present disclosure will be explained and described. Specifically: Tool 1: restart_pod (restart Pod); the specific tool description is: restart the specified Pod; the required input is: pod_name (string): the name of the Pod, such as "my-app-pod"; namespace (string): the namespace where the Pod is located, such as "default"; the specific output is: status (string): the operation status, such as "success" or "failed"; the specific implementation is: call the Kubernetes API to delete the Pod, triggering the Pod to restart.
[0087] Further, after the operation instruction is executed, the execution method of the operation instruction may further include: in response to a touch operation on an interactive control on an interactive interface corresponding to the Kubernetes container orchestration cluster, determining a target interactive control and the information of the interface to be displayed corresponding to the target interactive control; wherein, the information of the interface to be displayed includes the running state information of the container orchestration cluster, the operation instructions to be executed, the instruction execution results corresponding to the operation instructions, and other custom configuration information, etc.; obtaining the information of the interface to be displayed and displaying the information of the interface to be displayed, so as to view the running state information of the container orchestration cluster and / or the operation instructions to be executed and / or the instruction execution results corresponding to the operation instructions and / or other custom configuration information. That is to say, in the actual application process, if a user needs to view a certain piece of information or multiple pieces of information, it can be achieved based on this interactive interface; when viewing information, the user can click on the interactive control corresponding to the information.
[0088] Hereinafter, specific embodiments will be combined to further explain and illustrate the execution method of the operation instruction recorded in the exemplary embodiments of the present disclosure.
[0089] Specifically, in the actual application process, first, when the system starts, configure the parameters of the data collection component (such as the API Server address, Prometheus address, etc.), and configure the parameters of the intelligent agent AI Agent (such as the model address, Prompt template, etc.); secondly, fine-tune the large model to improve its analysis effect in specific scenarios; then, some rules can also be customized, such as alarm rules, resource adjustment rules, etc., and these rules can be used as a supplement to the AI Agent to further improve the intelligence level of the system. Based on this, the system accuracy and user experience can be improved on the basis of improving the system flexibility. Among them, improving the system flexibility is mainly manifested as: the initial configuration and user-defined rules enable the system to flexibly adapt to different Kubernetes environments and user needs; improving the system accuracy is mainly manifested as: model fine-tuning can improve the analysis effect of the large model in specific scenarios, thereby improving the accuracy of system diagnosis and prediction; improving the user experience is mainly manifested as: user-defined rules enable users to adjust according to their own needs, increasing the user's sense of participation and control.
[0090] On the premise of the above-described content, in order to generate operation instructions, first, source data is collected; specifically, the source data described herein is generated during the operation of Kubernetes, including the cluster status information provided by the Kubernetes APIServer, the metric data collected by Prometheus, and the application logs collected by the logging system; this source data is characterized by real-time, diversity, and high-dimensionality; among them, real-time means that the data will be continuously updated as the system runs; diversity means that it includes structured metric data and unstructured log data; high-dimensional means that the data contains information in multiple dimensions, such as Pod names, node names, metric names, etc. Secondly, target data is determined based on the source data; among them, what can be described herein may include the adjustment plan of Kubernetes configuration or the possible failures of a certain service, etc.; at the same time, the adjustment plan of Kubernetes configuration can be used to automatically adjust the resource configuration of Pods to avoid resource waste or resource shortage; the possible failures of a certain service can be used to perform capacity planning in advance to avoid service interruption.
[0091] So far, the execution method of the operation instructions recorded in the exemplary embodiments of the present disclosure has been fully implemented. Based on the foregoing content, it can be known that the execution method of the operation instructions recorded in the exemplary embodiments of the present disclosure has at least the following advantages: on the one hand, the large model can be used to drive the AI Agent to realize the intelligent operation and maintenance of Kubernetes; at the same time, through the real-time analysis of the Kubernetes operation data, problem diagnosis, prediction, and resource optimization can be carried out; on the other hand, a fully automated operation and maintenance process is realized; that is, from data collection to operation execution, the entire process does not require manual intervention, greatly improving the operation and maintenance efficiency; further, this solution can also adapt to different Kubernetes environments and can be customized through configuration; on the other hand, in the exemplary embodiments of the present disclosure, by integrating the large model AI Agent into the Kubernetes operation and maintenance system, the system can automatically collect and analyze the operation data of the Kubernetes system, and on this basis, perform problem diagnosis, fault prediction, and resource optimization; compared with the prior art, it does not require manual intervention and greatly reduces the complexity of operation and maintenance, improving the operation and maintenance efficiency; at the same time, the present disclosure can also predict potential risks, perform resource planning in advance, and ensure the stable operation of the system; in addition, the present disclosure uses a large model for analysis, so it is easier to understand and analyze complex problems, improving the accuracy and efficiency of problem diagnosis.
[0092] The following is an embodiment of the present disclosure device, which can be used to execute the embodiment of the method of the present disclosure. For the details not disclosed in the embodiment of the present disclosure device, please refer to the embodiment of the method of the present disclosure.
[0093] The exemplary embodiments of the present disclosure also provide an execution device for operation instructions. Specifically, referring to Figure 7 As shown, the execution device for operation instructions may include a data preprocessing module 710, an abnormal behavior recognition module 720, a problem diagnosis and prediction module 730, and an operation instruction execution module 740. Among them:
[0094] The data preprocessing module 710 can be used to obtain the original container operation and maintenance data generated by Kubernetes during the container operation and maintenance process, and preprocess the original container operation and maintenance data to obtain standard container operation and maintenance data;
[0095] The abnormal behavior recognition module 720 can be used to perform abnormal behavior recognition on the standard container operation and maintenance data based on a preset abnormal behavior recognition model to obtain the abnormal behavior recognition result of Kubernetes during the container operation and maintenance process;
[0096] The problem diagnosis and prediction module 730 can be used to perform problem diagnosis and prediction on the abnormal behavior recognition result based on a preset problem diagnosis model to obtain the abnormal problem prediction result that will be generated by Kubernetes during the container operation and maintenance process;
[0097] The operation instruction execution module 740 can be used to input the abnormal problem prediction result into a preset instruction generation model to obtain an operation instruction for solving the abnormal problem, and execute the operation instruction to solve the abnormal problem.
[0098] In an exemplary embodiment of the present disclosure, obtaining the original container operation and maintenance data generated by Kubernetes during the container operation and maintenance process includes: calling the application programming interface service of Kubernetes, and obtaining the original cluster status data generated by Kubernetes during the container operation and maintenance process from the application programming interface service; wherein, the original cluster status data includes at least one of container status data, node status data, and node resource data; calling the status monitoring tool of Kubernetes, and obtaining the original monitoring metric data generated by Kubernetes during the container operation and maintenance process from the status monitoring tool; wherein, the original monitoring metric data includes CPU usage rate and / or memory usage rate; calling the log collection tool of Kubernetes, and obtaining the original log data generated by Kubernetes during the container operation and maintenance process from the log collection tool.
[0099] In an exemplary embodiment of the present disclosure, the preset abnormal behavior recognition model includes an embedding mapping layer, an encoding layer, and a mixture-of-experts model layer; wherein, based on the preset abnormal behavior recognition model, abnormal behavior recognition is performed on the standard container operation and maintenance data to obtain the abnormal behavior recognition result of Kubernetes during the container operation and maintenance process, including: generating first basic information to be predicted according to the standard cluster status data and standard monitoring index data in the standard container operation and maintenance data, and generating first context information to be predicted according to the standard log data in the standard container operation and maintenance data and the preset first model prompt parameters; performing embedding mapping processing on the first basic information to be predicted based on the embedding mapping layer to obtain first operation and maintenance features, and performing embedding mapping processing on the first context information to be predicted based on the embedding mapping layer to obtain a first context flag sequence; performing encoding processing on the first operation and maintenance features and the first context flag sequence based on the encoding layer to obtain a first context overall representation, and performing prediction on the first context flag sequence and the first context overall representation based on the mixture-of-experts model layer to obtain the abnormal behavior recognition result of Kubernetes during the container operation and maintenance process.
[0100] In an exemplary embodiment of the present disclosure, the mixture-of-experts model layer includes a gated network model and an expert neural network model; wherein, based on the mixture-of-experts model layer, performing prediction on the first context flag sequence and the first context overall representation to obtain the abnormal behavior recognition result of Kubernetes during the container operation and maintenance process, including: determining the model weights of the expert neural network model based on the first context flag sequence by the gated network model; determining the target neural network model required for performing the prediction task of the abnormal behavior recognition result from the expert neural network according to the model weights; inputting the first context flag sequence and the first context overall representation into the target neural network model to obtain the abnormal behavior recognition result of Kubernetes during the container operation and maintenance process; wherein, the abnormal behavior recognition result includes at least one of the restart frequency of the containers included in Kubernetes being greater than a first preset threshold, the CPU usage rate of the nodes included in Kubernetes being greater than a second preset threshold, and the network latency of the nodes being greater than or equal to a third preset threshold.
[0101] In an exemplary embodiment of the present disclosure, the preset abnormal behavior recognition model is obtained in the following manner: Obtain historical container operation and maintenance data generated by Kubernetes during container operation and maintenance, and the actual behavior labels of Kubernetes under the historical container operation and maintenance data; Train a low-rank adaptation model based on the historical container operation and maintenance data and the actual behavior labels to obtain low-rank matrix parameters of Kubernetes in the container operation and maintenance scenario; Fine-tune a large language model based on the low-rank matrix parameters to obtain the preset abnormal behavior recognition model.
[0102] In an exemplary embodiment of the present disclosure, based on a preset problem diagnosis model, problem diagnosis and prediction are performed on the abnormal behavior recognition result to obtain the abnormal problem prediction result that Kubernetes will generate during container operation and maintenance, including: Generating second basic information to be predicted according to the abnormal behavior recognition result, and generating second context information to be predicted according to preset second model prompt parameters; Inputting the second basic information to be predicted and the second context information to be predicted into the preset problem diagnosis model to obtain the abnormal problem prediction result that Kubernetes will generate during container operation and maintenance; Wherein, the abnormal problem prediction result includes at least one of the following: The restart frequency of the container is greater than a first preset threshold, and the container cannot work properly, resulting in the inability to execute tasks normally; The CPU usage rate of the nodes included in Kubernetes is greater than a second preset threshold, and the resources are about to be exhausted, resulting in the inability to execute tasks normally; The network latency of the nodes is greater than a third preset threshold, resulting in a time delay in task execution.
[0103] In an exemplary embodiment of the present disclosure, inputting the abnormal problem prediction result into a preset instruction generation model to obtain an operation instruction for solving the abnormal problem, including: Generating third basic information to be predicted according to the abnormal problem prediction result, and generating third context information to be predicted according to preset third model prompt parameters; Inputting the third basic information to be predicted and the third context information to be predicted into the preset instruction generation model to obtain an operation instruction for solving the abnormal problem; Wherein, the operation instruction includes at least one of the following: Restarting the container, expanding the capacity of the node, and allocating bandwidth resources to the node.
[0104] In an exemplary embodiment of the present disclosure, executing the operation instruction to solve the abnormal problem includes: determining the instruction category of the operation instruction, and matching an instruction execution tool for executing the operation instruction according to the instruction category; wherein, the instruction tool includes at least one of a container restart tool, a computing resource allocation tool, and a bandwidth resource allocation tool; invoking the instruction execution tool, and controlling the instruction execution tool to execute the operation instruction to solve the abnormal behavior.
[0105] In an exemplary embodiment of the present disclosure, matching an instruction execution tool for executing the operation instruction according to the instruction category includes: obtaining the tool attribute information of the original tool, and determining the structured document of the original tool according to the tool attribute information; determining the tool instruction information, tool operation information, and tool response information of the original tool according to the structured document, and generating a tool usage instance corresponding to the original tool according to the tool instruction information, tool operation information, and tool response information; based on the instruction category and the tool usage instance, matching an instruction execution tool for executing the operation instruction from the original tools.
[0106] In an exemplary embodiment of the present disclosure, the execution device of the operation instruction may further include:
[0107] A to-be-displayed interface information determination module, which can be used to determine a target interactive control and to-be-displayed interface information corresponding to the target interactive control in response to a touch operation on an interactive control on an interactive interface corresponding to a Kubernetes container orchestration cluster; wherein, the to-be-displayed interface information includes at least one of the running state information of the container orchestration cluster, operation instructions to be executed, instruction execution results corresponding to the operation instructions, and other custom configuration information.
[0108] A to-be-displayed interface information display module, which can be used to obtain the to-be-displayed interface information and display the to-be-displayed interface information, so as to view the running state information of the container orchestration cluster and / or operation instructions to be executed and / or instruction execution results corresponding to the operation instructions and / or other custom configuration information.
[0109] The specific details of each module in the above execution device of the operation instruction have been described in detail in the corresponding execution method of the operation instruction, so they will not be repeated here.
[0110] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0111] In addition, although the steps of the methods in the present disclosure are described in a specific order in the drawings, this does not require or imply that these steps must be executed in that specific order, or that all the steps shown must be executed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.
[0112] In an exemplary embodiment of the present disclosure, an electronic device capable of implementing the above method is also provided. Those skilled in the art can understand that various aspects of the present disclosure can be implemented as a system, method, or program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as circuits, modules, or systems here.
[0113] Next, refer to Figure 8 to describe the electronic device 800 according to this embodiment of the present disclosure. Figure 8 The electronic device 800 shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0114] As Figure 8 shown, the electronic device 800 is presented in the form of a general-purpose computing device. The components of the electronic device 800 may include but are not limited to: at least one of the above-mentioned processing units 810, at least one of the above-mentioned storage units 820, a bus 830 connecting different system components (including the storage unit 820 and the processing unit 810), and a display unit 840.
[0115] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 810, so that the processing unit 810 executes the steps according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification. For example, the processing unit 810 can execute as Figure 1Step S110 shown in the figure: Obtain the original container operation and maintenance data generated during the container operation and maintenance process of Kubernetes, and preprocess the original container operation and maintenance data to obtain standard container operation and maintenance data; Step S120: Based on a preset abnormal behavior recognition model, perform abnormal behavior recognition on the standard container operation and maintenance data to obtain the abnormal behavior recognition result during the container operation and maintenance process of Kubernetes; Step S130: Based on a preset problem diagnosis model, perform problem diagnosis and prediction on the abnormal behavior recognition result to obtain the abnormal problem prediction result that will occur during the container operation and maintenance process of Kubernetes; Step S140: Input the abnormal problem prediction result into a preset instruction generation model to obtain an operation instruction for solving the abnormal problem, and execute the operation instruction to solve the abnormal problem.
[0116] The storage unit 820 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 8201 and / or a cache storage unit 8202, and may further include a read-only storage unit (ROM) 8203.
[0117] The storage unit 820 may further include a program / utilities 8204 having a set (at least one) of program modules 8205. Such program modules 8205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.
[0118] The bus 830 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of the multiple bus structures.
[0119] The electronic device 800 can also communicate with one or more external devices 900 (such as a keyboard, a pointing device, a Bluetooth device, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 800, and / or communicate with any device that enables the electronic device 800 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 850. Moreover, the electronic device 800 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 860. As shown in the figure, the network adapter 860 communicates with other modules of the electronic device 800 through the bus 830. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 800, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0120] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software, or can be implemented by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, and the software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0121] In an exemplary embodiment of the present disclosure, there is also provided a computer-readable storage medium, on which a program product capable of implementing the above method of this specification is stored. In some possible implementation manners, various aspects of the present disclosure can also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to enable the terminal device to execute the steps according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification.
[0122] The program product for implementing the above method according to the embodiments of the present disclosure can adopt a portable compact disc read-only memory (CD-ROM) and include program code, and can run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, the readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device.
[0123] The program product may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the foregoing. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0124] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal may take many forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable signal medium may also be any readable medium other than a readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the readable medium may be transmitted with any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0125] The program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on the remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0126] In addition, the above drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present disclosure, and are not for limiting purposes. It is easily understood that the processes shown in the above drawings do not indicate or limit the time sequence of these processes. Additionally, it is also easily understood that these processes may be executed, for example, synchronously or asynchronously in multiple modules.
[0127] Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not invented by the present disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.
Claims
1. A method for executing an operation instruction, characterized in that: include: Obtaining the original container operation and maintenance data generated by Kubernetes during the container operation and maintenance process, and preprocessing the original container operation and maintenance data to obtain standard container operation and maintenance data; Based on a preset abnormal behavior recognition model, abnormal behavior recognition is performed on the standard container operation and maintenance data to obtain abnormal behavior recognition results of the Kubernetes during the container operation and maintenance process; Based on a preset problem diagnosis model, problem diagnosis and prediction are performed on the abnormal behavior identification result to obtain the prediction result of abnormal problems that may occur in the container operation and maintenance process of Kubernetes; The abnormal problem prediction result is input into a preset instruction generation model to obtain an operation instruction for solving the abnormal problem, and the operation instruction is executed to solve the abnormal problem.
2. The method for executing an operation instruction according to claim 1, characterized in that: Obtain the original container operation and maintenance data generated by Kubernetes during the container operation and maintenance process, including: Calling the Kubernetes application program interface service, and obtaining the original cluster state data generated by Kubernetes during the container operation and maintenance process from the application program interface service; wherein the original cluster state data includes at least one of container state data, node state data, and node resource data; Calling the Kubernetes status monitoring tool, and obtaining the original monitoring indicator data generated by the Kubernetes during the container operation and maintenance process from the status monitoring tool; wherein the original monitoring indicator data includes CPU usage and / or memory usage; The log collection tool of the Kubernetes is called, and the original log data generated by the Kubernetes during the container operation and maintenance process is obtained from the log collection tool.
3. The method for executing an operation instruction according to claim 1, characterized in that: The preset abnormal behavior recognition model includes an embedding mapping layer, a coding layer and a hybrid expert model layer; Among them, based on the preset abnormal behavior recognition model, the standard container operation and maintenance data is subjected to abnormal behavior recognition, and the abnormal behavior recognition result of the Kubernetes in the container operation and maintenance process is obtained, including: Generate first basic information to be predicted according to standard cluster status data and standard monitoring indicator data in the standard container operation and maintenance data, and generate first context information to be predicted according to standard log data in the standard container operation and maintenance data and preset first model prompt parameters; Performing embedding mapping processing on the first basic information to be predicted based on the embedding mapping layer to obtain a first operation and maintenance feature, and performing embedding mapping processing on the first context information to be predicted based on the embedding mapping layer to obtain a first context flag sequence; Based on the encoding layer, the first operation and maintenance feature and the first context flag sequence are encoded to obtain a first context overall representation, and based on the hybrid expert model layer, the first context flag sequence and the first context overall representation are predicted to obtain an abnormal behavior recognition result of the Kubernetes during the container operation and maintenance process.
4. The method for executing an operation instruction according to claim 3, characterized in that: The hybrid expert model layer includes a gated network model and an expert neural network model; The first context flag sequence and the first context overall representation are predicted based on the hybrid expert model layer to obtain the abnormal behavior recognition result of the Kubernetes in the container operation and maintenance process, including: Determining a model weight of the expert neural network model based on the gating network model according to the first contextual flag sequence; Determining a target neural network model required for performing a prediction task of abnormal behavior recognition results from the expert neural network according to the model weights; Inputting the first context flag sequence and the overall representation of the first and last context into the target neural network model to obtain the abnormal behavior recognition result of Kubernetes during the container operation and maintenance process; Among them, the abnormal behavior identification result includes at least one of the following: the restart frequency of the container included in the Kubernetes is greater than a first preset threshold, the CPU usage of the node included in the Kubernetes is greater than a second preset threshold, and the network delay of the node is greater than or equal to a third preset threshold.
5. The method for executing an operation instruction according to claim 3 or 4, characterized in that: The preset abnormal behavior recognition model is obtained in the following manner: Obtain historical container operation and maintenance data generated by Kubernetes during the container operation and maintenance process and actual behavior labels of the Kubernetes under the historical container operation and maintenance data; Based on the historical container operation and maintenance data and the actual behavior label pairs, the low-rank adaptation model is trained to obtain the low-rank matrix parameters of Kubernetes in the container operation and maintenance scenario; The large language model is fine-tuned based on the low-rank matrix parameters to obtain the preset abnormal behavior recognition model.
6. The method for executing an operation instruction according to claim 1, characterized in that: Based on the preset problem diagnosis model, the abnormal behavior identification result is diagnosed and predicted to obtain the abnormal problem prediction results that may be generated by the Kubernetes during the container operation and maintenance process, including: Generate second basic information to be predicted according to the abnormal behavior recognition result, and generate second context information to be predicted according to a preset second model prompt parameter; The second basic information to be predicted and the second context information to be predicted are input into a preset problem diagnosis model to obtain a prediction result of an abnormal problem that may be generated by Kubernetes during the container operation and maintenance process; wherein the prediction result of the abnormal problem includes at least one of the following: The restart frequency of the container is greater than a first preset threshold, and the container cannot work normally, resulting in failure to perform tasks normally; The CPU usage of the nodes included in Kubernetes is greater than the second preset threshold, and the resources are about to be exhausted, resulting in the inability to perform tasks normally; The network delay of the node is greater than a third preset threshold, resulting in a delay in task execution.
7. The method for executing an operation instruction according to claim 1, characterized in that: The abnormal problem prediction result is input into a preset instruction generation model to obtain an operation instruction for solving the abnormal problem, including: Generating third basic information to be predicted according to the abnormal problem prediction result, and generating third context information to be predicted according to a preset third model prompt parameter; Inputting the third basic information to be predicted and the third context information to be predicted into a preset instruction generation model to obtain an operation instruction to solve the abnormal problem; The operation instruction includes at least one of the following: restarting the container, expanding the capacity of the node, and allocating bandwidth resources to the node.
8. The method for executing an operation instruction according to claim 1, characterized in that: Executing the operation instruction to solve the abnormal problem includes: Determine the instruction category of the operation instruction, and match an instruction execution tool for executing the operation instruction according to the instruction category; wherein the instruction tool includes at least one of a container restart tool, a computing resource allocation tool, and a bandwidth resource allocation tool; The instruction execution tool is called, and the instruction execution tool is controlled to execute the operation instruction to resolve the abnormal behavior.
9. The method for executing an operation instruction according to claim 1, characterized in that: The method for executing the operation instruction also includes: In response to a touch operation on an interactive control on an interactive interface corresponding to a Kubernetes container orchestration cluster, determine a target interactive control and interface information to be displayed corresponding to the target interactive control; wherein the interface information to be displayed includes at least one of the running status information of the container orchestration cluster, the operation instruction to be executed, the instruction execution result corresponding to the operation instruction, and other custom configuration information; The interface information to be displayed is obtained and displayed, so as to facilitate viewing of the operating status information of the container orchestration cluster and / or the operation instructions to be executed and / or the instruction execution results corresponding to the operation instructions and / or other custom configuration information.
10. An execution device for an operation instruction, characterized in that: include: The data preprocessing module is used to obtain the original container operation and maintenance data generated by Kubernetes during the container operation and maintenance process, and preprocess the original container operation and maintenance data to obtain standard container operation and maintenance data; An abnormal behavior identification module is used to identify abnormal behaviors of the standard container operation and maintenance data based on a preset abnormal behavior identification model, and obtain abnormal behavior identification results of the Kubernetes during the container operation and maintenance process; A problem diagnosis and prediction module is used to diagnose and predict the abnormal behavior identification results based on a preset problem diagnosis model, and obtain the prediction results of abnormal problems that may occur in the container operation and maintenance process of Kubernetes; The operation instruction execution module is used to input the abnormal problem prediction result into a preset instruction generation model to obtain the operation instruction for solving the abnormal problem, and execute the operation instruction to solve the abnormal problem.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for executing the operation instruction according to any one of claims 1 to 9 is implemented.
12. An electronic device, characterized in that: include: processor; as well as A memory, configured to store executable instructions of the processor; The processor is configured to execute the method for executing the operating instruction according to any one of claims 1 to 9 by executing the executable instruction.