Container fault processing method and device, electronic equipment and computer program product

By using large language models in containers for troubleshooting, automatically collecting and analyzing log information and hardware information, quickly locate the cause of failure and provide repair solutions, the problem of low accuracy of container fault diagnosis is solved, and the safety and reliability of containers are improved.

CN120508429APending Publication Date: 2025-08-19INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510636239.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

The accuracy of container fault diagnosis is low, and the prior art is difficult to quickly and accurately identify faults and provide effective repair solutions in complex and variable container environments.

Method used

By obtaining the log information and hardware information of the target container, using a large language model for troubleshooting, extracting the cause of the failure and traceability path, and formulating and implementing a fault repair plan.

Benefits of technology

It improves the accuracy and efficiency of container fault diagnosis, enhances the safety and reliability of containerized applications, and reduces dependence on professional and technical personnel.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120508429A_ABST
    Figure CN120508429A_ABST
Patent Text Reader

Abstract

The invention discloses a container fault processing method and device, electronic equipment and a computer program product. The method comprises the steps that log information of a target container is acquired, and hardware information of cluster nodes deployed by the target container is acquired; inputting the log information and the hardware information into the target model to obtain a fault diagnosis result; under the condition that the fault diagnosis result represents that the target container has the fault, extracting a fault reason and a fault tracing path from the fault diagnosis result; and determining a fault repair scheme based on the fault reason and the fault tracing path, and executing the fault repair scheme on the target container. Through the method and the device, the problem of low fault diagnosis accuracy of the container in related technologies is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and more specifically, to a method, device, electronic device, and computer program product for troubleshooting a container. Background Art

[0002] Amid the rapid development of cloud computing and the Internet of Things (IoT), container technology has become the preferred solution for microservice deployment due to its lightweight, fast startup, and high resource utilization. Container technology allows applications and their dependencies to be packaged and run in a lightweight, self-contained environment, greatly simplifying application deployment and management processes and improving development and operations efficiency. However, the widespread adoption of container technology also presents new security and operations challenges. Container environments can share the host machine's kernel and resources. While this sharing mechanism significantly improves resource utilization, it also blurs the boundaries between containers, increasing security risks. Misconfiguration errors, excessive resource sharing, and insecure network communication policies can all pose security risks in container environments, providing malicious attackers with avenues for intrusion. For example, incorrect container configuration can allow attackers to gain control of the host machine through container escape techniques; improper resource sharing can lead to resource contention for applications running within the container; and network communication vulnerabilities can be exploited by attackers to launch network-level attacks.

[0003] Related technologies for container fault diagnosis rely on extensive log analysis and expert experience, making fault diagnosis particularly challenging in container environments. Log analysis is a time-consuming and tedious task, requiring sifting, correlating, and analyzing massive amounts of log data to identify the root cause of a problem. While expert knowledge can quickly resolve issues in some cases, it relies on individual experience and is difficult to quantify and scale. This limitation is particularly pronounced in complex and ever-changing container environments. Furthermore, the high dynamics and complexity of container environments further exacerbate the difficulty of fault diagnosis. Containers have short lifecycles and may start and stop frequently; communication between containers is complex, potentially involving different network policies and resource allocation; and resource competition and interaction between containers, the host, and other containers can lead to unforeseen failures. These characteristics require fault diagnosis methods to not only accurately identify faults but also possess real-time and intelligent capabilities to adapt to the rapid changes in the container environment.

[0004] Currently, no effective solution has been proposed to the problem of low accuracy in fault diagnosis of containers in related technologies. Summary of the Invention

[0005] The main purpose of this application is to provide a container fault handling method, device, electronic equipment and computer program product to solve the problem of low accuracy of container fault diagnosis in related technologies.

[0006] To achieve the above objectives, according to one aspect of the present application, a container fault handling method is provided. The method comprises: obtaining log information of a target container and hardware information of a cluster node on which the target container is deployed; inputting the log information and hardware information into a target model to obtain a fault diagnosis result; if the fault diagnosis result indicates that the target container is faulty, extracting the fault cause and fault tracing path from the fault diagnosis result; determining a fault repair plan based on the fault cause and fault tracing path, and executing the fault repair plan on the target container.

[0007] Optionally, after inputting the log information and hardware information into the target model to obtain the fault diagnosis result, the method further includes: obtaining the fault diagnosis result of the target container within a preset period, and determining the fault type distribution, fault frequency and fault repair result of the target container from the fault diagnosis result within the preset period; obtaining the log information and hardware information of the target container within the preset period, and determining the computing resource usage status of the target container based on the log information and hardware information within the preset period; determining the operating status evaluation value of the target container based on the computing resource usage status, and issuing a warning message when the operating status evaluation value is greater than or equal to the evaluation value threshold, wherein the warning message is used to prompt that there is an operating risk for the target container.

[0008] Optionally, the target model is obtained by: obtaining historical log information and historical hardware information of the target container, and determining the historical fault diagnosis results of each set of historical log information and corresponding historical hardware information; determining each set of historical log information, historical hardware information and historical fault diagnosis results as a set of training samples to obtain multiple sets of training samples; training the large language model through multiple sets of training samples to obtain the target model.

[0009] Optionally, the method also includes: collecting new log information, new hardware information and new fault diagnosis results of the target container every optimization cycle; determining each group of new log information, new hardware information and new fault diagnosis results as a group of first new training samples to obtain multiple groups of first new training samples; combining the multiple groups of first new training samples with the multiple groups of training samples to obtain optimized multiple groups of training samples, and training the target model with the optimized multiple groups of training samples to obtain an optimized target model.

[0010] Optionally, obtaining the log information of the target container and obtaining the hardware information of the cluster node where the target container is deployed includes: deploying the target container to the cluster node in the cluster environment; collecting the log information of the target container through a preset collection tool, wherein the log information includes at least one of the following: standard output, error log and event information, and the standard output includes execution information and abnormal information of the application in the target container; collecting the hardware information of the cluster node through a preset collection tool, wherein the hardware information includes at least one of the following: central processing unit, memory, disk storage and network status.

[0011] Optionally, after inputting the log information and hardware information into the target model and obtaining the fault diagnosis result, the method further includes: when the fault diagnosis result indicates that there is no fault in the target container, obtaining the operating status of the target container within a preset time period; when there is an abnormality in the operating status of the target container within the preset time period, determining the cause of the abnormality; determining the cause of the abnormality as a new fault cause, determining a new fault tracing path for the new fault cause, and determining the new fault cause and the new fault tracing path as the target fault diagnosis result; determining the log information, hardware information, and target fault diagnosis result of the target container within the preset time period as a second new training sample, combining the second new training sample with multiple groups of training samples to obtain updated multiple groups of training samples, and updating the target model based on the updated multiple groups of training samples.

[0012] Optionally, after obtaining the log information of the target container and the hardware information of the cluster node where the target container is deployed, the method further includes: converting the unstructured data in the log information into structured data to obtain the converted log information; filtering the noise data from the converted log information and hardware information, and standardizing the hardware information to obtain updated log information and hardware information; and executing the step of inputting the log information and hardware information into the target model based on the updated log information and hardware information.

[0013] To achieve the above-mentioned objectives, according to another aspect of the present application, a container fault handling device is provided. The device comprises: an acquisition unit for acquiring log information of a target container and hardware information of a cluster node on which the target container is deployed; an input unit for inputting the log information and hardware information into a target model to obtain a fault diagnosis result; an extraction unit for extracting the fault cause and fault tracing path from the fault diagnosis result when the fault diagnosis result indicates that the target container has a fault; and a first execution unit for determining a fault repair plan based on the fault cause and fault tracing path, and executing the fault repair plan on the target container.

[0014] In an embodiment of the present application, the log information of the target container is obtained, and the hardware information of the cluster node where the target container is deployed is obtained; the log information and hardware information are input into the target model to obtain a fault diagnosis result; when the fault diagnosis result indicates that there is a fault in the target container, the cause of the fault and the fault tracing path are extracted from the fault diagnosis result; a fault repair plan is determined based on the fault cause and the fault tracing path, and the fault repair plan is executed on the target container. By automatically collecting and analyzing the log information and hardware information of the container, the cause of the fault is quickly located using the target model, and an accurate fault tracing path and a feasible fault repair plan are provided, thereby achieving the purpose of enhancing the security and reliability of containerized applications, thereby achieving the technical effect of improving the fault diagnosis accuracy of the container, and further solving the technical problem of low fault diagnosis accuracy of the container in the related art. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:

[0016] Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing a container fault handling method is shown;

[0017] Figure 2 This is a flowchart of a method for troubleshooting a container according to an embodiment of the present application;

[0018] Figure 3 is a schematic diagram of an optional container fault handling method provided in an embodiment of the present application;

[0019] Figure 4 is a schematic diagram of a fault handling device for a container provided in accordance with an embodiment of the present application;

[0020] Figure 5 This is a structural block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0021] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0022] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0023] It should be noted that the collected information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in this application are information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with relevant laws, regulations and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation portals for users to choose to authorize or refuse. For example, an interface is set up between this system and relevant users or institutions to provide users with corresponding operation portals for users to choose to agree or refuse the automated decision-making results; if the user chooses to refuse, the expert decision-making process will be entered.

[0024] Example 1

[0025] According to an embodiment of the present application, an embodiment of a method for handling container faults is also provided. It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system, such as a set of computer-executable instructions, and although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown.

[0026] The method embodiment provided in the first embodiment of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 The hardware structure block diagram of a computer terminal (or mobile device) for implementing a container fault handling method is shown. Figure 1As shown, the computer terminal 10 (or mobile device) may include one or more (illustrated as 102a, 102b, ..., 102n in the figure) processors 102 (the processor 102 may include but is not limited to a processing device such as an MCU (Microcontroller Unit) or an FPGA (Field-Programmable Gate Array), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a USB (Universal Serial Bus) port (which may be included as one of the ports of a BUS (Business, bus)), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.

[0027] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry". The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuitry may be a single independent processing module, or may be incorporated in whole or in part into any of the other components of the computer terminal 10 (or mobile device). As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0028] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the container fault handling method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implementing the above-mentioned container fault handling method. The memory 104 may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0029] The transmission device 106 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the computer terminal 10. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.

[0030] The display may be, for example, a touch screen liquid crystal display that enables a user to interact with a user interface of the computer terminal 10 (or mobile device).

[0031] In the above operating environment, this application provides a container fault handling method. Figure 2 Flowchart of a method for troubleshooting a container according to an embodiment of the present application. Figure 2 As shown, the method includes:

[0032] Step S201: Obtain log information of a target container and obtain hardware information of a cluster node where the target container is deployed.

[0033] In step S201, the log information of the target container is collected through the log collection device. In a production environment, logs are not directly obtained in real time through the command line. Instead, the log information of all containers is continuously collected into a centralized log management system. The collected log information may come from different applications and have different formats. In order to facilitate subsequent analysis, the logs need to be formatted into a unified structure, and a structured text format can be used. After the logs are standardized, they are filtered and cleaned to remove irrelevant or redundant information and retain key log entries and information to improve the accuracy and efficiency of subsequent fault diagnosis.

[0034] Hardware information can include CPU (Central Processing Unit) usage, memory usage, disk I / O data volume, network I / O data volume, etc. The collected hardware information needs to be standardized, such as by performing differential processing to reflect actual resource consumption and normalizing to eliminate dimensional differences between different indicators to facilitate unified comparison and analysis.

[0035] Step S202: input the log information and hardware information into the target model to obtain the fault diagnosis result.

[0036] In step S202, the pre-processed log information may include key fields such as timestamp, log level, error type, container ID, and log text. An offline target model that has been trained using a large amount of historical log information, historical hardware information, and historical fault diagnosis results is called. The target model can be a big data or language model with the ability to recognize complex log patterns and understand context. The prepared log information and hardware information are input into the target model in a structured or vectorized form. After receiving the input data, the target model executes its internal prediction and analysis algorithms to generate fault diagnosis results.

[0037] Step S203 : When the fault diagnosis result indicates that a fault exists in the target container, the fault cause and the fault tracing path are extracted from the fault diagnosis result.

[0038] In step S203, the fault diagnosis results may include a series of numerical values or labels related to the container status and log anomalies. Labels may include the severity of the anomaly, the anomaly type, and the specific log entries or hardware metrics that may have caused the anomaly. Within the fault diagnosis results output by the target model, the anomalies most relevant to the target container failure are identified. Examples include abnormally high frequencies of certain log entries, specific system metrics outside of normal ranges, or discrepancies between the model's prediction of the container status and the actual state.

[0039] Extract the cause of the fault from the fault diagnosis results. For example, a sudden increase in memory usage, combined with log information, may be due to a memory leak. Locate the cause of the fault by analyzing abnormal hardware indicators, such as a surge in CPU usage, disk I / O bottlenecks, network latency, etc. For example, a persistently high CPU usage of 99% may be due to poor performance code or infinite loops in the application running in the container. Combined with the container's operating environment, configuration parameters, and the overall status of the cluster, conduct contextual analysis to further refine the cause of the fault. For example, if a container suddenly becomes abnormal after an upgrade, it may be due to version compatibility issues.

[0040] Use time series data to track changes in system status before and after a failure occurs. Combined with the timestamps of container logs, the order of events can be reconstructed, thereby building a timeline from the onset of the failure to the expansion of its impact. Analyze the correlation between log entries and hardware indicators to identify which events or operations triggered the abnormal state. For example, if a large number of database query failures are recorded in the container log, and at the same time the network I / O indicators are abnormal, it can be inferred that network problems cause database access delays, thereby affecting container performance. Based on the fault diagnosis results, trace the fault source and reversely infer which log entries or hardware events may be the root cause of the fault. For example, the model may indicate that the container network connection is abnormal. Further tracing through network logs and network monitoring data, it is found that the cause is network congestion on the host machine.

[0041] Step S204: determine a fault repair solution based on the fault cause and the fault tracing path, and execute the fault repair solution on the target container.

[0042] In step S204, the cause of the failure is analyzed in depth, whether it is a resource bottleneck (such as CPU, memory), configuration error, software defect or network problem. At the same time, the fault tracing path is reviewed to understand which operations or events ultimately led to the failure. Based on the cause of the failure, targeted repair measures are formulated. For example, for memory leaks, you can update the code or restart the container; for network problems, you can check the network configuration or adjust the network policy. If the cause of the failure is related to resource allocation, such as insufficient CPU or memory, the fault repair solution can be to adjust the resource requests and limits of the container, or reschedule the container in the cluster.

[0043] Perform operations on the target container according to the developed fault remediation plan. This can include updating the container image, modifying configuration files, adjusting resource allocation, or performing system-level operations. After executing the remediation plan, continuously monitor the container status and cluster health to verify that the fault has been resolved, while also noting whether the remediation operations have introduced new issues. For common fault types, you can design automated scripts to execute the remediation process, reducing the need for manual intervention and improving response speed and efficiency.

[0044] The container fault handling method provided in the embodiment of the present application obtains the log information of the target container and the hardware information of the cluster node deployed by the target container; inputs the log information and hardware information into the target model to obtain a fault diagnosis result; when the fault diagnosis result indicates that there is a fault in the target container, extracts the fault cause and the fault tracing path from the fault diagnosis result; determines a fault repair plan based on the fault cause and the fault tracing path, and executes the fault repair plan on the target container. By automatically collecting and analyzing the log information and hardware information of the container, the target model is used to quickly locate the fault cause, and an accurate fault tracing path and a feasible fault repair plan are provided, thereby achieving the purpose of enhancing the security and reliability of containerized applications, thereby realizing the technical effect of improving the fault diagnosis accuracy of the container, and further solving the technical problem of low fault diagnosis accuracy of the container in the related art.

[0045] In order to ensure the safe operation of the container, after the container fault is determined, a warning message is promptly used to prompt the container operation risk. Optionally, in the container fault handling method provided in the embodiment of the present application, after the log information and hardware information are input into the target model to obtain the fault diagnosis result, the method also includes: obtaining the fault diagnosis result of the target container within a preset period, and determining the fault type distribution, fault frequency and fault repair result of the target container from the fault diagnosis result within the preset period; obtaining the log information and hardware information of the target container within the preset period, and determining the computing resource usage status of the target container based on the log information and hardware information within the preset period; determining the operating status evaluation value of the target container based on the computing resource usage status, and issuing a warning message when the operating status evaluation value is greater than or equal to the evaluation value threshold, wherein the warning message is used to prompt that there is an operation risk of the target container.

[0046] In some embodiments, a preset period can be 24 hours or a week to collect fault diagnosis results for the target container. The frequency of occurrence of different fault types is calculated from the fault diagnosis results within the preset period. This helps identify common failure modes and guides the development of future preventive measures. The distribution of fault types, fault frequency, and fault repair results for the target container within the period are calculated. The fault repair results record the repair status of each fault, including whether the repair was successful, the repair time, and the repair measures taken. This helps evaluate the container's self-healing capabilities and the effectiveness of the fault handling process.

[0047] Continuously collect computing resource usage data such as CPU, memory, disk I / O, and network I / O of the target container to understand its computing resource usage trends and peak values. Based on the computing resource usage status within a preset period, calculate the average, standard deviation, and peak value of resource utilization to reflect the resource demand and fluctuation of the container. Combine the distribution and frequency of failure types to score the failure stability of the container. For example, the greater the number of failures, the lower the score. Weight the resource usage index and the failure index to obtain the operating status evaluation value of the target container. The evaluation value can be calculated using a variety of statistical analysis methods, such as exponentially weighted moving average, nearest neighbor algorithm, or machine learning model prediction.

[0048] Set a threshold for the health assessment value to distinguish between normal operation and risk. The threshold can be based on historical data and the normal operating range of the container. When the health assessment value is greater than or equal to the threshold, a warning message is automatically triggered. The warning message can be sent via email, SMS, system notification, or displayed directly on the management interface. The warning message can include a detailed risk description, such as specific indicators of resource overuse, fault type and frequency information, and preliminary fault remediation suggestions or response strategies.

[0049] This embodiment periodically evaluates the operating status of the target container, promptly identifies and warns of potential operating risks, and takes preventive and optimization measures to ensure the continuity and reliability of container services.

[0050] In order to perform fault diagnosis on the target container, a target model needs to be trained. Optionally, in the fault handling method for the container provided in the embodiment of the present application, the target model is obtained by: obtaining historical log information and historical hardware information of the target container, and determining the historical fault diagnosis results of each set of historical log information and corresponding historical hardware information; determining each set of historical log information, historical hardware information and historical fault diagnosis results as a set of training samples to obtain multiple sets of training samples; training a large language model using the multiple sets of training samples to obtain a target model.

[0051] In some embodiments, all historical log information and historical hardware information during the operation of the container are collected from the monitoring system of the target container and its cluster. The historical log information may include standard output, error logs, event information, etc. The historical hardware information may include indicators such as CPU usage, memory usage, disk I / O data volume, and network I / O data volume. The historical fault diagnosis results are used as labels and associated with the corresponding historical log information and historical hardware information. Each set of historical log information, historical hardware information, and historical fault diagnosis results together constitutes a training sample. Ensure that the training sample contains sufficient information so that the model can learn the association between faults and log and hardware indicators.

[0052] A large language model is selected as a foundation. Pre-built training samples, including historical log text, hardware metric data, and fault diagnosis result labels, are fed into the model for supervised learning training. During training, the model automatically extracts features from historical log and hardware information and identifies the correlation between these features and fault diagnosis results. The model's accuracy and generalization capabilities are continuously improved by adjusting model parameters, increasing training rounds, or using larger datasets. The optimization objective is to optimize the model's performance on the test dataset, ensuring that the target model not only performs well on the training data but also accurately predicts unseen container failures.

[0053] It should be noted that the target model can be an offline model, using big data and a large language model for fault analysis and model training. First, preprocessed data is input into the large language model, which is labeled and classified using historical data to identify common fault types. The labeled data is stored in the database and serves as the basis for subsequent model training. The offline model then undergoes continuous incremental training based on real-time data to improve its adaptability to new faults and changes. This incremental training optimizes the model by updating weights or adding new samples. Finally, by evaluating the model's prediction accuracy, the model parameters are regularly adjusted to improve its performance in fault diagnosis. After sufficient training, the offline model can quickly and accurately locate new faults and provide corresponding solution recommendations.

[0054] This embodiment trains a target model, which can predict container failures based on log information and hardware information, and analyze the causes of failures, providing strong support for container operation and maintenance and improving the accuracy of container fault detection.

[0055] In order to ensure the timeliness of the target model, the target model needs to be optimized regularly. Optionally, in the fault handling method for the container provided in the embodiment of the present application, the method further includes: collecting new log information, new hardware information and new fault diagnosis results of the target container every optimization cycle; determining each group of new log information, new hardware information and new fault diagnosis results as a group of first new training samples to obtain multiple groups of first new training samples; combining the multiple groups of first new training samples with the multiple groups of training samples to obtain multiple optimized groups of training samples, and training the target model with the optimized multiple groups of training samples to obtain an optimized target model.

[0056] In some embodiments, in each optimization cycle (which can be set to daily, weekly, or on-demand), new log information, new hardware information, and new fault diagnosis results generated during this period of the target container are collected. These data reflect the recent operating status and fault events of the container. Ensure that the new log information and hardware information are accurately paired with the corresponding fault diagnosis results to form the first new training sample. Combine multiple groups of first new training samples with multiple groups of existing training samples to form an optimized training sample set. Input the optimized multiple groups of training samples into the target model for incremental training. The target model will use the new data to further adjust its parameters to improve the target model's ability to recognize new fault scenarios.

[0057] By continuously optimizing the target model, this embodiment can make the diagnosis of container failures more accurate and better adapt to changes in the container environment, thereby improving the stability and operation and maintenance efficiency of the container service.

[0058] Whether a container fails is determined by obtaining the log information and hardware information of the container. Optionally, in the container fault handling method provided in an embodiment of the present application, obtaining the log information of the target container and obtaining the hardware information of the cluster node where the target container is deployed include: deploying the target container to the cluster node in the cluster environment; collecting the log information of the target container through a preset collection tool, wherein the log information includes at least one of the following: standard output, error log and event information, and the standard output includes execution information and abnormal information of the application in the target container; collecting the hardware information of the cluster node through a preset collection tool, wherein the hardware information includes at least one of the following: central processing unit, memory, disk storage and network status.

[0059] In some embodiments, a suitable cluster environment is selected to deploy the target container, ensuring that the container can run and provide services in the cluster. A pre-configured log collection tool is configured or used to monitor and collect container log information in real time. The collected target container log information may include standard output, error logs, and event information. Standard output contains application execution information, including normal operations and exception information; error logs record errors and warnings during container operation; and event information provides detailed information about container lifecycle events, such as container start, stop, and restart.

[0060] Use pre-set collection tools to continuously monitor the hardware information of the cluster nodes where the target container is deployed, including central processing unit (CPU) usage, memory usage, disk storage usage, and network status. Hardware information includes CPU usage, which reflects the container's demand for computing resources; memory usage, which indicates the container's memory consumption; disk storage usage, which monitors the container's data storage needs; and network status, including network I / O data volume and latency, which reflects the container's network communication status.

[0061] This example collects log information from target containers and hardware information from cluster nodes, providing a solid data foundation for subsequent fault diagnosis and analysis. This data will be used to train a large language model to improve the automation and accuracy of container fault diagnosis, thereby reducing container service downtime and improving system reliability and operational efficiency.

[0062] In order to avoid inaccurate fault diagnosis results of the target model, the container is continuously monitored for abnormalities within a preset period. Optionally, in the fault handling method for the container provided in the embodiment of the present application, after the log information and hardware information are input into the target model to obtain the fault diagnosis result, the method further includes: when the fault diagnosis result indicates that there is no fault in the target container, obtaining the operating status of the target container within the preset period; when there is an abnormality in the operating status of the target container within the preset period, determining the cause of the abnormality; determining the cause of the abnormality as a new fault cause, determining a new fault tracing path for the new fault cause, and determining the new fault cause and the new fault tracing path as the target fault diagnosis result; determining the log information, hardware information, and target fault diagnosis result of the target container within the preset period as a second new training sample, combining the second new training sample with multiple groups of training samples to obtain updated multiple groups of training samples, and updating the target model based on the updated multiple groups of training samples.

[0063] In some embodiments, a preset monitoring period is set, such as the last few minutes, hours or day, to observe the operating status of the target container. During the preset period, the operating status data of the target container is continuously collected, including but not limited to business indicators such as the application's response time, throughput, and failed request rate, as well as operation and maintenance indicators such as the number of times the container is started and stopped. The collected operating status data is analyzed using statistical analysis, anomaly detection algorithms or expert rules to determine whether there are any anomalies. After confirming that there are anomalies in the operating status of the target container, the specific cause of the anomaly is located through log analysis, resource monitoring, and application performance monitoring. The cause of the anomaly is refined into specific fault types, such as resource bottlenecks, improper configuration, software defects, or external service interruptions. The fault tracing path of the cause of the anomaly is determined, including the time point when the fault occurred, the operating environment of the container, related system events, and service interaction processes.

[0064] The identified anomaly cause is used as a new fault cause, and the fault tracing path is used as a new fault tracing path. Together, these two constitute the target fault diagnosis result. The log information and hardware information of the target container within a preset time period are combined with the target fault diagnosis result to form a second set of newly added training samples. Multiple sets of the second set of newly added training samples are merged with the original sets of training samples to form updated sets of training samples. Based on the updated sets of training samples, the target model is retrained or incrementally updated. By learning the newly added fault scenarios and causes, the model's fault diagnosis capabilities will be further enhanced, especially for those hidden faults that are difficult to identify in the early stages.

[0065] This embodiment continuously feeds deeper anomaly detection results back into model training, thereby continuously optimizing the target model's fault diagnosis capabilities. This allows it to more accurately and comprehensively identify various fault scenarios in containers, thereby improving the overall stability and O&M efficiency of the container service.

[0066] In order to improve the training efficiency of the target model, the log information and hardware information can be structured and denoised. Optionally, in the fault handling method for the container provided in the embodiment of the present application, after obtaining the log information of the target container and the hardware information of the cluster node where the target container is deployed, the method further includes: converting the unstructured data in the log information into structured data to obtain the converted log information; filtering the noise data from the converted log information and hardware information, and standardizing the hardware information to obtain updated log information and hardware information; and executing the step of inputting the log information and hardware information into the target model based on the updated log information and hardware information.

[0067] In some embodiments, the unstructured data in the log information is converted into structured data, and the converted log information is stored in a tabular form, with each row representing a log entry and each column corresponding to a structured field. In this way, the target model can more easily understand and process the log data, thereby improving the efficiency and accuracy of fault diagnosis. The converted log information and hardware information are subjected to noise data filtering to exclude irrelevant or duplicate information. For example, log entries that are not related to the current fault diagnosis, or occasional abnormal values in hardware indicators are filtered out. Hardware information (such as CPU usage, memory usage) is converted into a unified dimension and range, such as converting all CPU usage into a proportional value between 0 and 1, or converting all memory usage into GB as a unit. Standardization helps to eliminate the magnitude differences between different hardware indicators, enabling the model to more fairly evaluate the usage of various resources.

[0068] This embodiment improves the diagnostic capability and prediction accuracy of the model by ensuring that the log information and hardware information input into the target model are structured, clean, and standardized.

[0069] According to another embodiment of the present application, an optional container fault handling method is also provided. Figure 3 is a schematic diagram of an optional container fault handling method provided in an embodiment of the present application, such as Figure 3 As shown, the method includes:

[0070] Step S301: deploy the application container into the cluster environment.

[0071] Step S302: The collection tool on the cluster node will continuously collect the container's standard output, error logs, event information, and the usage and health status of the node's hardware resources such as CPU, memory, storage, and network.

[0072] Step S303: The information collected in step S302 is sent to the information preprocessing module through the data collection model. The data collection module is responsible for collecting container logs and operating environment data, and the information preprocessing module is responsible for structured processing of the collected data.

[0073] Step S304: The information preprocessing module converts the unstructured information into structured data. It then randomly extracts a portion of the information as training data and sends it to the offline model training module. The remaining data is then sent to the online fault diagnosis module. The model training module uses the processed data to train a large model, while the fault diagnosis module uses the model to perform fault diagnosis, traceability, and real-time monitoring of the container.

[0074] Step S305: The offline model training module receives the training data sent in step S304.

[0075] Step S306: The data received in step S305 is automatically labeled and stored. The labeled data enters the database and serves as the basis for subsequent model training.

[0076] Step S307: The offline model continuously performs incremental training on the labeled data to improve the model's adaptability to new faults and changes. This incremental training optimizes the model by updating weights or adding new samples.

[0077] Step S308: After the incremental training is completed, the model parameters are adjusted by evaluating the prediction accuracy of the model to improve its performance in fault diagnosis.

[0078] Step S309: The online diagnosis module receives the data sent in step S304 and performs fault diagnosis using the offline model.

[0079] Step S310: The real-time monitoring module analyzes the uploaded data and issues real-time early warning and health status warning.

[0080] Step S311: The fault tracing module calls the offline trained model to locate the fault and generate possible solutions based on the real-time fault information.

[0081] Step S312: The visualization display module graphically displays the fault analysis results, traceability process and historical data.

[0082] This embodiment uses an optional container fault handling method to automatically collect and analyze container log data, utilizes machine learning technology to quickly locate the cause of the fault, and provides accurate fault tracing information and feasible solutions, thereby significantly improving the efficiency and accuracy of fault handling, reducing dependence on professional technicians, and enhancing the security and reliability of containerized applications. Through efficient data processing and intelligent fault diagnosis, the fault detection, analysis, and tracing capabilities in the container environment are effectively improved, significantly improving the reliability and maintainability of the container system. It is suitable for the operation and maintenance management of large-scale container clusters, helping to reduce system downtime and improve system stability and reliability.

[0083] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0084] Example 2

[0085] The present application also provides a container fault handling device. It should be noted that the container fault handling device of the present application can be used to execute the container fault handling method provided in the present application. The following describes the container fault handling device provided in the present application.

[0086] According to an embodiment of the present application, a device for implementing the above-mentioned container fault handling method is also provided. Figure 4 Schematic diagram of a fault handling device for a container according to an embodiment of the present application. Figure 4 As shown, the device includes:

[0087] An acquisition unit 401 is configured to acquire log information of a target container and hardware information of a cluster node where the target container is deployed.

[0088] Input unit 402, used to input log information and hardware information into the target model to obtain fault diagnosis results;

[0089] An extraction unit 403 is configured to extract the fault cause and the fault tracing path from the fault diagnosis result when the fault diagnosis result indicates that the target container has a fault;

[0090] The first execution unit 404 is configured to determine a fault repair solution based on the fault cause and the fault tracing path, and execute the fault repair solution on the target container.

[0091] The container fault handling device provided in the embodiment of the present application obtains the log information of the target container and the hardware information of the cluster node deployed by the target container through the acquisition unit 401; the input unit 402 inputs the log information and hardware information into the target model to obtain a fault diagnosis result; the extraction unit 403 extracts the fault cause and the fault tracing path from the fault diagnosis result when the fault diagnosis result indicates that the target container has a fault; the first execution unit 404 determines the fault repair plan based on the fault cause and the fault tracing path, and executes the fault repair plan on the target container. By automatically collecting and analyzing the log information and hardware information of the container, the target model is used to quickly locate the fault cause, and an accurate fault tracing path and a feasible fault repair plan are provided, thereby achieving the purpose of enhancing the security and reliability of containerized applications, thereby realizing the technical effect of improving the fault diagnosis accuracy of the container, and further solving the technical problem of low fault diagnosis accuracy of the container in the related art.

[0092] Optionally, in the fault handling device for a container provided in an embodiment of the present application, the device further includes: a first determination unit, used to obtain the fault diagnosis results of the target container within a preset period, and determine the fault type distribution, fault frequency and fault repair results of the target container from the fault diagnosis results within the preset period; a second determination unit, used to obtain the log information and hardware information of the target container within the preset period, and determine the computing resource usage status of the target container based on the log information and hardware information within the preset period; a warning unit, used to determine the operating status evaluation value of the target container based on the computing resource usage status, and issue a warning message when the operating status evaluation value is greater than or equal to the evaluation value threshold, wherein the warning message is used to prompt that there is an operating risk for the target container.

[0093] Optionally, in the fault handling device for a container provided in an embodiment of the present application, the device further includes: a third determination unit, used to obtain historical log information and historical hardware information of the target container, and determine the historical fault diagnosis results of each set of historical log information and corresponding historical hardware information; a fourth determination unit, used to determine each set of historical log information, historical hardware information and historical fault diagnosis results as a set of training samples to obtain multiple sets of training samples; and a first training unit, used to train a large language model using multiple sets of training samples to obtain a target model.

[0094] Optionally, in the fault handling device for a container provided in an embodiment of the present application, the device further includes: a collection unit, used to collect new log information, new hardware information and new fault diagnosis results of the target container every optimization cycle; a fifth determination unit, used to determine each group of new log information, new hardware information and new fault diagnosis results as a group of first new training samples, to obtain multiple groups of first new training samples; a second training unit, used to combine the multiple groups of first new training samples with the multiple groups of training samples to obtain multiple optimized groups of training samples, and train the target model using the optimized multiple groups of training samples to obtain an optimized target model.

[0095] Optionally, in the fault handling device for a container provided in an embodiment of the present application, the acquisition unit 401 includes: a deployment module, used to deploy the target container to a cluster node in a cluster environment; a first acquisition module, used to collect log information of the target container through a preset acquisition tool, wherein the log information includes at least one of the following: standard output, error log and event information, and the standard output includes execution information and abnormal information of the application in the target container; a second acquisition module, used to collect hardware information of the cluster node through a preset acquisition tool, wherein the hardware information includes at least one of the following: central processing unit, memory, disk storage and network status.

[0096] Optionally, in the fault handling device for the container provided in the embodiment of the present application, the device further includes: an operating status acquisition unit, for acquiring the operating status of the target container within a preset time period when the fault diagnosis result indicates that there is no fault in the target container; a sixth determination unit, for determining the cause of the abnormality when there is an abnormality in the operating status of the target container within the preset time period; a seventh determination unit, for determining the cause of the abnormality as a new fault cause, determining a new fault tracing path for the new fault cause, and determining the new fault cause and the new fault tracing path as the target fault diagnosis result; a third training unit, for determining the log information, hardware information, and target fault diagnosis result of the target container within the preset time period as a second new training sample, combining the second new training sample with multiple groups of training samples to obtain updated multiple groups of training samples, and updating the target model based on the updated multiple groups of training samples.

[0097] Optionally, in the fault handling device for a container provided in an embodiment of the present application, the device further includes: a conversion unit for converting unstructured data in the log information into structured data to obtain converted log information; a filtering unit for filtering noise data from the converted log information and hardware information, and performing standardization processing on the hardware information to obtain updated log information and hardware information; and a second execution unit for executing the step of inputting the log information and hardware information into the target model based on the updated log information and hardware information.

[0098] It should be noted that the acquisition unit 401, input unit 402, extraction unit 403, and first execution unit 404 correspond to steps S201 to S204 in Example 1. The four units and the corresponding steps implement the same examples and application scenarios, but are not limited to the contents disclosed in the above-mentioned Example 1. It should be noted that the above-mentioned modules or units can be hardware components or software components stored in a memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). The above-mentioned modules or units can also be part of a device and can be run in the computer terminal 10 provided in Example 1.

[0099] Example 3

[0100] An embodiment of the present application may provide an electronic device, Figure 5 This is a structural block diagram of an electronic device according to an embodiment of the present application. Figure 5 As shown, the electronic device may include: one or more ( Figure 5 Only one is shown) processor 502, memory 504, storage controller, and peripheral interface, wherein the peripheral interface is connected to the radio frequency module, audio module and display.

[0101] Among them, the memory can be used to store software programs and modules, such as program instructions / modules corresponding to the methods and devices in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implementing the above-mentioned method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely arranged relative to the processor, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network and a combination thereof.

[0102] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: obtain the log information of the target container and the hardware information of the cluster node where the target container is deployed; input the log information and hardware information into the target model to obtain the fault diagnosis result; when the fault diagnosis result indicates that there is a fault in the target container, extract the fault cause and the fault tracing path from the fault diagnosis result; determine the fault repair plan based on the fault cause and the fault tracing path, and execute the fault repair plan on the target container.

[0103] The processor can also call the information and application stored in the memory through the transmission device to perform the following steps: obtain the fault diagnosis results of the target container within a preset period, and determine the fault type distribution, fault frequency and fault repair results of the target container from the fault diagnosis results within the preset period; obtain the log information and hardware information of the target container within the preset period, and determine the computing resource usage status of the target container based on the log information and hardware information within the preset period; determine the operating status evaluation value of the target container based on the computing resource usage status, and issue a warning message when the operating status evaluation value is greater than or equal to the evaluation value threshold, wherein the warning message is used to prompt that there is an operating risk for the target container.

[0104] The processor can also call the information and application stored in the memory through the transmission device to perform the following steps: obtain the historical log information and historical hardware information of the target container, and determine the historical fault diagnosis results of each set of historical log information and corresponding historical hardware information; determine each set of historical log information, historical hardware information and historical fault diagnosis results as a set of training samples to obtain multiple sets of training samples; train the large language model through multiple sets of training samples to obtain the target model.

[0105] The processor can also call the information and application programs stored in the memory through the transmission device to perform the following steps: collect the new log information, new hardware information and new fault diagnosis results of the target container every optimization cycle; determine each group of new log information, new hardware information and new fault diagnosis results as a group of first new training samples to obtain multiple groups of first new training samples; combine the multiple groups of first new training samples with the multiple groups of training samples to obtain the optimized multiple groups of training samples, and train the target model with the optimized multiple groups of training samples to obtain the optimized target model.

[0106] The processor can also call the information and applications stored in the memory through the transmission device to perform the following steps: deploy the target container to the cluster node in the cluster environment; collect the log information of the target container through a preset collection tool, wherein the log information includes at least one of the following: standard output, error log and event information, and the standard output includes the execution information and abnormal information of the application in the target container; collect the hardware information of the cluster node through the preset collection tool, wherein the hardware information includes at least one of the following: central processing unit, memory, disk storage and network status.

[0107] The processor can also call the information and application stored in the memory through the transmission device to perform the following steps: when the fault diagnosis result indicates that there is no fault in the target container, obtain the operating status of the target container within a preset time period; when there is an abnormality in the operating status of the target container within the preset time period, determine the cause of the abnormality; determine the cause of the abnormality as a new fault cause, determine a new fault tracing path for the new fault cause, and determine the new fault cause and the new fault tracing path as the target fault diagnosis result; determine the log information, hardware information, and target fault diagnosis result of the target container within the preset time period as a second new training sample, combine the second new training sample with multiple groups of training samples to obtain updated multiple groups of training samples, and update the target model based on the updated multiple groups of training samples.

[0108] The processor can also call the information and applications stored in the memory through the transmission device to perform the following steps: converting the unstructured data in the log information into structured data to obtain the converted log information; filtering the noise data of the converted log information and hardware information, and standardizing the hardware information to obtain updated log information and hardware information; and executing the step of inputting the log information and hardware information into the target model based on the updated log information and hardware information.

[0109] By adopting the embodiment of the present application, a method is provided for obtaining the log information of the target container and the hardware information of the cluster node where the target container is deployed; the log information and hardware information are input into the target model to obtain the fault diagnosis result; when the fault diagnosis result indicates that the target container has a fault, the fault cause and the fault tracing path are extracted from the fault diagnosis result; based on the fault cause and the fault tracing path, a fault repair plan is determined, and the fault repair plan is executed on the target container. By automatically collecting and analyzing the log information and hardware information of the container, using the target model to quickly locate the cause of the fault, and providing an accurate fault tracing path and a feasible fault repair plan, the purpose of enhancing the security and reliability of containerized applications is achieved, thereby achieving the technical effect of improving the fault diagnosis accuracy of the container, and further solving the technical problem of low fault diagnosis accuracy of the container in the related art.

[0110] It can be understood by those skilled in the art that Figure 5 The structure shown is for illustration only, and the electronic device may also be a smart phone, a tablet computer, a PDA, a mobile internet device (MID), a PAD or other terminal device. Figure 5 It does not limit the structure of the above electronic device. For example, the electronic device may also include Figure 5 More or fewer components (such as network interfaces, display devices, etc.) shown in, or with Figure 5Different configurations shown.

[0111] A person skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0112] Example 4

[0113] The embodiment of the present application further provides a storage medium. Optionally, in this embodiment, the storage medium can be used to store the program code executed by the container fault handling method provided in the first embodiment.

[0114] Optionally, in this embodiment, the storage medium may be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group.

[0115] The present application also provides a computer program product, which, when executed on a data processing device, is suitable for executing the steps of the container fault handling method.

[0116] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0117] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0118] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0119] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0120] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0121] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0122] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A method for handling a container failure, characterized in that: include: Obtain log information of the target container and hardware information of the cluster node where the target container is deployed; Inputting the log information and the hardware information into a target model to obtain a fault diagnosis result; If the fault diagnosis result indicates that the target container has a fault, extracting the fault cause and the fault tracing path from the fault diagnosis result; A fault repair solution is determined based on the fault cause and the fault tracing path, and the fault repair solution is executed on the target container.

2. The method according to claim 1, characterized in that After inputting the log information and the hardware information into a target model to obtain a fault diagnosis result, the method further includes: Obtaining fault diagnosis results of the target container within a preset period, and determining the fault type distribution, fault frequency, and fault repair result of the target container from the fault diagnosis results within the preset period; Obtaining log information and hardware information of the target container within the preset period, and determining a computing resource usage status of the target container based on the log information and hardware information within the preset period; An operating status evaluation value of the target container is determined based on the computing resource usage status, and a warning message is issued when the operating status evaluation value is greater than or equal to an evaluation value threshold, wherein the warning message is used to prompt that there is an operating risk for the target container.

3. The method according to claim 1, characterized in that The target model is obtained in the following way: Obtaining historical log information and historical hardware information of the target container, and determining historical fault diagnosis results for each set of historical log information and corresponding historical hardware information; Determine each set of historical log information, the historical hardware information, and the historical fault diagnosis results as a set of training samples, to obtain multiple sets of training samples; The target model is obtained by training a large language model using the multiple sets of training samples.

4. The method according to claim 1, wherein The method further comprises: Collecting new log information, new hardware information, and new fault diagnosis results of the target container every optimization cycle; Determine each group of the newly added log information, the newly added hardware information, and the newly added fault diagnosis result as a group of first newly added training samples, to obtain multiple groups of first newly added training samples; The multiple groups of first newly added training samples are combined with multiple groups of training samples to obtain multiple optimized groups of training samples, and the target model is trained using the multiple optimized groups of training samples to obtain an optimized target model.

5. The method according to claim 1, wherein Obtaining the log information of the target container and the hardware information of the cluster node where the target container is deployed includes: Deploy the target container to a cluster node in a cluster environment; Collecting log information of the target container using a preset collection tool, wherein the log information includes at least one of the following: standard output, error log, and event information, and the standard output includes execution information and abnormal information of the application in the target container; The hardware information of the cluster nodes is collected by the preset collection tool, wherein the hardware information includes at least one of the following: central processing unit, memory, disk storage and network status.

6. The method according to claim 1, characterized in that After inputting the log information and the hardware information into a target model to obtain a fault diagnosis result, the method further includes: If the fault diagnosis result indicates that the target container has no fault, obtaining the operating status of the target container within a preset time period; If the operating state of the target container is abnormal within the preset time period, determining the cause of the abnormality; Determine the abnormal cause as a new fault cause, determine a new fault tracing path for the new fault cause, and determine the new fault cause and the new fault tracing path as a target fault diagnosis result; The log information, hardware information, and target fault diagnosis results of the target container within the preset time period are determined as second newly added training samples, the second newly added training samples are combined with multiple groups of training samples to obtain updated multiple groups of training samples, and the target model is updated based on the updated multiple groups of training samples.

7. The method according to claim 1, characterized in that After obtaining the log information of the target container and obtaining the hardware information of the cluster node where the target container is deployed, the method further includes: Converting unstructured data in the log information into structured data to obtain converted log information; filtering noise data from the converted log information and the hardware information, and performing standardization processing on the hardware information to obtain updated log information and hardware information; The step of inputting the log information and the hardware information into a target model is performed based on the updated log information and hardware information.

8. A container fault handling device, characterized in that: include: An acquisition unit, configured to acquire log information of a target container and hardware information of a cluster node where the target container is deployed; An input unit, configured to input the log information and the hardware information into a target model to obtain a fault diagnosis result; An extraction unit, configured to extract a fault cause and a fault tracing path from the fault diagnosis result if the fault diagnosis result indicates that the target container has a fault; The first execution unit is configured to determine a fault repair solution based on the fault cause and the fault tracing path, and execute the fault repair solution on the target container.

9. An electronic device, characterized in that: include: a memory storing an executable program; A processor is configured to run the program, wherein the program, when running, executes the container fault handling method according to any one of claims 1 to 7.

10. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by a processor, the steps of the container fault handling method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Fault diagnosis method and device, storage medium and electronic equipment

    CN120687329A