Abnormality processing method, service platform and computing device
By performing service anomaly detection and recovery before container startup, the problem of low efficiency in handling service anomalies in existing technologies is solved, resulting in a more efficient and stable service platform operation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-10
- Publication Date
- 2026-03-20
AI Technical Summary
In existing technologies, the exception handling methods for containerized services cannot effectively identify the correctness of the service itself, which leads to increased downtime and manual intervention costs due to container restarts, reducing the efficiency of exception handling and the stability of the service platform.
Before starting the container, perform service anomaly detection by checking environment variables, dependent services, and file integrity from multiple dimensions to ensure that the service is normal before starting the container; if the start fails, perform targeted recovery processing based on the anomaly information.
It improves the comprehensiveness and accuracy of anomaly identification, reduces the number of restarts and downtime, enhances the stability and availability of the service platform, and reduces the cost of manual intervention.
Smart Images

Figure CN121705083A_ABST
Abstract
Description
Technical Field
[0001] The embodiments in this specification relate to the fields of containerized service technology and computer technology, and in particular to an exception handling method, a service platform, and a computing device. Background Technology
[0002] With the rapid development of computer technology, containerized deployment has become the mainstream deployment method for providing various containerized application services.
[0003] Currently, service platforms typically configure service parameters through environment variables and rely on external services to complete specific functions. For containerized services, anomaly detection usually involves using probes to verify the service status after the container starts, and resolving anomalies by killing and restarting the container or through manual intervention.
[0004] However, in practical applications, due to the complex dependencies between various types of containerized services, and the fact that the correctness of the service itself is not directly related to the running state of the container, simply restarting the container for exception handling cannot guarantee the correctness of the service itself, nor can it provide targeted exception recovery for the container. Furthermore, container restarting not only increases service interruption time but also raises the cost of manual intervention, reducing the efficiency of exception handling and the stability of the service platform. Therefore, there is an urgent need for an exception handling method that can overcome the above shortcomings. Summary of the Invention
[0005] In view of this, embodiments of this specification provide an exception handling method. One or more embodiments of this specification also relate to a service platform, a computing device, a computer-readable storage medium, and a computer program product, to address technical deficiencies in the prior art.
[0006] According to a first aspect of the embodiments of this specification, an exception handling method is provided, applied to a service platform, the service platform including a container, the container being used to run containerized services, the method comprising: Perform service anomaly detection on the containerized service and obtain service detection results; If the containerized service is determined to be normal based on the service detection results, the container is started. If startup fails, startup exception information is obtained, and based on the startup exception information, exception recovery processing is performed on the container.
[0007] According to a second aspect of the embodiments of this specification, a service platform is provided, including a container, a service anomaly detection module, a container startup module, and a container anomaly recovery module; The container is used to run containerized services; The service anomaly detection module is used to respond to the service anomaly detection command, perform service anomaly detection on the containerized service, obtain the service detection result, and, if it is determined that the containerized service is normal based on the service detection result, generate a container startup command and send the container startup command to the container startup module. The container startup module is used to start the container in response to the container startup command, and generate a container exception recovery command in the event that the container startup fails, and send the container exception recovery command to the container exception recovery module. The container anomaly recovery module is used to respond to the container anomaly recovery command, obtain startup anomaly information, and perform anomaly recovery processing on the container based on the startup anomaly information.
[0008] According to a third aspect of the embodiments of this specification, a computing device is provided, comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer program / instructions, which, when executed by the processor, implement the steps of the above-described exception handling method.
[0009] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores a computer program / instructions that, when executed by a processor, implement the steps of the above-described exception handling method.
[0010] According to a fifth aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described exception handling method.
[0011] One embodiment of this specification implements an exception handling method applied to a service platform. The service platform includes a container for running containerized services. The method includes: performing service exception detection on the containerized services and obtaining service detection results; starting the container if the containerized services are determined to be normal based on the service detection results; and if the start fails, obtaining start exception information and performing exception recovery processing on the container based on the start exception information.
[0012] By performing service anomaly detection on the containerized services running in the container before starting the container on the service platform, obtaining service detection results, and only starting the container if the containerized service is normal, the correct configuration and availability of the containerized services provided by the container are ensured. At the same time, by directly obtaining startup anomaly information in the event of container startup failure, and performing targeted anomaly recovery processing on the container based on the startup anomaly information, the comprehensiveness and accuracy of anomaly identification for the service platform are improved, the targeting and efficiency of anomaly recovery processing are enhanced, the cost of manual intervention is reduced, and the stability and availability of the containerized services provided by the service platform are guaranteed. Attached Figure Description
[0013] Figure 1 This is a flowchart of an exception handling method provided in one embodiment of this specification; Figure 2 This is a timing diagram illustrating the execution process of an exception handling method provided in one embodiment of this specification. Figure 3 This is a schematic diagram of the structure of a service platform provided in one embodiment of this specification; Figure 4 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation
[0014] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0015] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0016] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0017] Furthermore, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0018] In one or more embodiments of this specification, a large model refers to a deep learning model with a large number of model parameters, typically containing hundreds of millions, tens of billions, hundreds of billions, trillions, or even tens of trillions of model parameters. A large model can also be called a foundation model. It is pre-trained using large-scale unlabeled corpora to produce a pre-trained model with hundreds of millions of parameters. Such models can adapt to a wide range of downstream tasks and have good generalization ability. Examples include Large Language Models (LLMs) and multi-modal pre-training models.
[0019] In practical applications, large models only require a small number of samples to fine-tune the pre-trained model before they can be applied to different tasks. Large models can be widely used in fields such as Natural Language Processing (NLP) and Computer Vision. Specifically, they can be applied to computer vision tasks such as Visual Question Answering (VQA), Image Captioning (IC), and Image Generation, as well as NLP tasks such as text-based sentiment classification, text summarization, and machine translation. The main application scenarios for large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.
[0020] First, the terms and concepts used in one or more embodiments of this specification will be explained.
[0021] Kubernetes (K8S) is an open-source container orchestration system used to automate the deployment, scaling, and management of containerized applications, supporting resource scheduling and high availability across host clusters.
[0022] Containers (Docker) are lightweight, portable software runtime units that ensure that applications or services can run consistently in different scenarios by encapsulating applications and their dependencies in an isolated environment.
[0023] Large Language Model (LLM) is a deep learning model trained on large amounts of text data, enabling it to generate natural language text or understand the meaning of language text. It can perform complex tasks such as spell checking, grammar correction, text summarization, machine translation, sentiment analysis, dialogue generation, and content recommendation.
[0024] Computer vision (CV) is a branch of artificial intelligence that enables machines to analyze and understand image / video content.
[0025] Natural Language Processing (NLP) is a branch of artificial intelligence that focuses on enabling interaction between computers and humans through natural language. It allows computers to understand, interpret, and generate human language, supporting applications such as text analysis, question answering, conversational language, speech recognition, and speech synthesis.
[0026] An Application Programming Interface Service (API) is a programming service that provides predefined rules and protocols that allow different software applications to communicate and interact with each other.
[0027] HTTP status codes are three-digit codes returned by a server in response to an HTTP request. They indicate the server's response status to the client's request and can be divided into five categories: informational status codes, success status codes, redirection status codes, client error status codes, and server error status codes.
[0028] With the rapid development of computer technology, containerized deployment has become the mainstream deployment method for providing various containerized application services.
[0029] Currently, service platforms typically configure service parameters through environment variables and rely on external services to complete specific functions. For containerized services, anomaly detection usually involves using probes to verify the service status after the container starts. If anomalies occur, the anomaly is resolved by killing and restarting the container, or through manual intervention. For example, service platforms centered around container orchestration platforms (such as Kubernetes) commonly employ probe checks such as post-startup health checks (Liveness Probe) and readiness checks to detect and recover from anomalies in containerized services. Readiness checks determine if the container is ready to receive traffic, while post-startup health checks determine if the container is running normally. If the checks fail, the platform typically triggers a container restart or recreate operation.
[0030] However, in practical applications, due to the complex dependencies between various types of containerized services, and the fact that the correctness of the service itself is not directly related to the running state of the container, standard probes used to inspect containers can typically only verify whether the web server process is alive or responsive. They cannot further verify, for example, whether model files are loaded correctly, or whether service inference dependencies (such as vector databases and feature engineering services) are truly available. They cannot verify the correctness of the service itself.
[0031] Meanwhile, the aforementioned detection method only performs periodic checks after the container starts, and "restarts the container" if any anomalies are found. This simplistic approach can lead to ineffective startups or inaccurate recovery. Specifically, if anomalies in the container service cannot be recovered by restarting, the container will repeatedly fail to start due to configuration errors; similarly, it cannot recover from specific internal anomalies, such as dependency service anomalies or model file errors, reducing the efficiency of anomaly handling for containerized services and the stability of the service platform.
[0032] In view of this, this specification provides an anomaly handling method to solve various technical problems existing in the practical application of the aforementioned container anomaly detection method based on a general health probe. This specification also relates to a service platform, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.
[0033] See Figure 1 , Figure 1 A flowchart of an exception handling method according to an embodiment of this specification is shown. The exception handling method is applied to a service platform, which includes containers for running containerized services. The exception handling method specifically includes the following steps.
[0034] Step 102: Perform service anomaly detection on the containerized service and obtain the service detection results.
[0035] The exception handling methods provided in one or more embodiments of this specification can be applied to various containerized service deployment scenarios with high availability requirements, such as the deployment of artificial intelligence model services, database services, and various services in microservice architectures. The corresponding application objects can include various service platforms, such as enterprise-level artificial intelligence service platforms, cloud computing platforms, and big data processing platforms.
[0036] A service platform is a system platform used to deploy, manage, and run various containerized services, such as a Kubernetes-based container orchestration platform or a containerized service platform optimized based on artificial intelligence models. Specifically, a service platform can include the runtime environment of containers and corresponding containerized service management modules. It provides the operational foundation and support for containerized services and can ensure the correctness and availability of containerized services by performing service anomaly detection, thereby avoiding service anomalies caused by service configuration errors or dependency issues, and improving the stability and service availability of the service platform.
[0037] A container is a lightweight, portable software runtime unit that encapsulates an application and its dependencies in an isolated environment, ensuring consistent operation across different scenarios. Containers can include Docker containers, operating system containers, security-enhanced containers, and more. Specifically, as the core runtime unit of a service platform, containers ensure the independence and security of the services they run through isolation. The runtime state of a container is closely related to the availability of the containerized service. In practice, the runtime state of a container and the service state of the containerized service it runs may not be consistent. A container may be running normally, but the containerized service it provides may not necessarily be running normally. Therefore, it is necessary to perform service anomaly detection on the containerized service before starting the container.
[0038] Containerized services are applications or application services encapsulated within containers. In other words, containerized services are based on container technology, encapsulating applications and their dependent environments into standardized, lightweight, and portable independent units. After the container starts, containerized services can provide specific services through interfaces, such as data generation, data processing, and signal response. Containerized services can include different types such as model services, database services, and API services. Model services can include large language model (LLM) services, computer vision (CV) image processing services, and natural language processing (NLP) services. Specifically, containerized services run through containers, and their correctness and availability directly affect the stability of the service platform's external services. The normality of containerized services can be determined by environment variables, external dependent services, model files, etc., and can be checked before the container starts. If the checks are normal, the container is started to ensure the service can function correctly.
[0039] Service anomaly detection is a multi-dimensional anomaly check performed on containerized services before container startup. Specifically, service anomaly detection for containerized services may include verifying environment variable configurations, the availability of dependent services, and file integrity, determining whether the containerized service meets the conditions for normal operation from multiple dimensions. Service anomaly detection is crucial for ensuring the normal operation of containerized services, and the results directly determine whether to execute the subsequent container startup process.
[0040] Service detection results are assessment information about the status of containerized services output during service anomaly detection. This includes whether the containerized service is functioning correctly, and if anomalies are present, it also includes service anomaly information, such as the type and cause of the anomaly, and allows for the identification of corresponding anomaly handling tasks for recovery. Service detection results can serve as a basis for deciding whether to start the container. The accuracy of the detection affects the correctness of the services provided, as well as the efficiency and effectiveness of anomaly recovery handling in abnormal situations. For example, a service detection result might be "Environment variables configured correctly, dependent services available, model files complete…", indicating that the containerized service can start normally; or it might be "Environment variables missing, dependent services unavailable", indicating that the containerized service cannot start normally.
[0041] In practical applications, containers in a service platform can start and execute containerized services through their internal runtime environment and configuration. When a container starts, it loads environment variables, dependent services, and other resources according to the preset configuration and starts the containerized service process.
[0042] Before the containerized services on the service platform start, service anomaly detection can be performed on the running containerized services. Specifically, service anomaly detection can include environment variable configuration checks, dependency service availability checks, and file integrity checks. If any detection fails, the corresponding service anomaly information is recorded, and service detection results are generated.
[0043] In one optional embodiment, for file integrity detection in service anomaly detection, the integrity of model files (such as weight files of large language models) can be verified.
[0044] Specifically, the hash value of the model file (e.g., SHA256 hash) can be calculated and compared with a preset baseline hash value. The preset baseline hash value can be generated and stored by the model generation system when the model file is released. During the service anomaly detection phase before container startup, the initialization script can perform hash calculation and comparison operations. If the hash values match, the model file is determined to be complete and untampered; if they do not match, the file is determined to be corrupted or abnormal. At the same time, the correctness of the file format can also be verified by parsing the header information of the model file (e.g., reading the file header magic number and determining whether it is in PyTorch.pt or ONNX format).
[0045] In this step, by performing service anomaly detection on the containerized services to be run by the container before starting the container in the service platform, abnormal containerized services caused by configuration errors, missing dependencies, or corrupted files can be detected and prevented from going online in advance. This avoids service anomalies caused by service configuration errors or dependency issues, provides a decision-making basis for subsequent container startup processes, improves the comprehensiveness and accuracy of anomaly detection, reduces the probability of service anomalies after container startup, reduces the number of container restarts caused by service anomalies, thereby reducing service interruption time and manual intervention costs, and improving the availability of the service platform.
[0046] Step 104: If the containerized service is determined to be normal based on the service detection results, start the container.
[0047] A containerized service is considered normal when it has passed service anomaly checks before startup. All dimensions of the service evaluation, such as environment configuration, dependent services, and file integrity, meet the conditions for normal operation. This indicates that the containerized service has passed multi-dimensional service anomaly checks and can provide stable and correct services after the container starts.
[0048] In practical applications, determining that a containerized service is functioning correctly based on service testing results can be done in several ways.
[0049] One possible approach is to determine whether the service detection result includes a service normal label, such as "containerized service normal". If it does, then the service anomaly detection is directly passed and the containerized service is normal.
[0050] Another option is to determine whether the service detection result includes service exception information. If the service detection result does not carry service exception information or the service exception information is empty, it can be determined that the containerized service is in a normal state and the container startup process can continue. If the service exception information is included, the container startup process needs to be paused, and based on the service exception information, the corresponding service exception handling task should be identified and called to perform exception recovery processing.
[0051] Another option is to set corresponding thresholds and make a comprehensive judgment based on the specific indicators provided in the service test results, such as the accuracy of environment variable configuration, the availability score of dependent services, and the integrity verification results of model files. If all key indicators reach the preset thresholds, the containerized service can be judged to be normal.
[0052] Once the containerized service is confirmed to be functioning correctly based on service detection results, the container can be started. Specifically, the container startup module in the service platform, upon receiving the container startup command generated by the service anomaly detection module, can call the container's runtime interface to start the container according to preset container configuration parameters (such as container image address, resource limits, network configuration, etc.). During startup, preset environment variables are loaded, dependent services are linked, and model files are initialized, thereby ensuring that the corresponding containerized service can run correctly within the container. Optionally, upon successful container startup, information such as the container's corresponding identifier (ID) and startup status can also be returned.
[0053] In this step, by starting the container when the containerized service is confirmed to be normal based on the service detection results, service anomalies caused by configuration errors or dependency issues of the containerized service can be avoided. This improves the stability and availability of the containerized service, reduces the number of container restarts caused by service anomalies, thereby reducing service interruption time. This allows the service platform to provide high-quality containerized services, improves the stability and availability of the service platform, avoids the problem of discovering service anomalies after the container has started and needing to restart the container or perform other processing, reduces manual intervention costs, and improves the automation and efficiency of anomaly handling.
[0054] Step 106: If startup fails, obtain startup exception information and perform exception recovery processing for the container based on the startup exception information.
[0055] Startup failure occurs when a container fails to complete initialization and enter the running state during the startup process. Specifically, startup failure can be identified by error messages or exception codes returned during container startup, indicating that the container startup process failed to execute successfully for some abnormal reason. This failure can be caused by an abnormality in the containerized service being run, or it can be caused by an abnormality in the container's configuration or runtime environment, such as a corrupted container image, insufficient computing resources, or incorrect network configuration. In cases of startup failure, it is necessary to perform abnormal recovery processing on the container itself.
[0056] Startup exception information is a descriptive message returned when a container fails to start, directly explaining the reason for the failure. Specifically, startup exception information can be a text description, a tag-based exception identifier, or a specific exit code, such as "Container startup failed: Image does not exist" or "Exit code 3". Startup exception information directly reflects the specific reason for the container startup failure and can serve as a basis for handling container exception recovery.
[0057] An exit code is a form of startup exception information returned when a container fails to start. Specifically, it can be a specific numerical code used to identify different reasons for startup failure. Specifically, exit codes can adopt common exit code standards of the service platform, such as HTTP status codes for containers; exit codes can also be custom exit codes, corresponding to specific exception reasons encountered in actual applications. For example, exit code 1 can indicate an environment variable configuration error, exit code 2 can indicate that a dependent service is unavailable, exit code 3 can indicate a container image problem, exit code 4 can indicate insufficient computing resources, and exit code 5 can indicate a network configuration error, etc. Optionally, for each exit code, a corresponding exception handling task can be further determined. This determination process can be based on a preset mapping relationship.
[0058] Anomaly recovery handling refers to the recovery operations performed in the event of an anomaly. Specifically, anomaly recovery handling for containers involves repair and recovery operations performed on the container itself when the container fails to start. These operations may include, but are not limited to, restarting the container, re-pulling the container image, reloading container resource configurations, and switching to a backup network configuration.
[0059] In practical applications, startup exception information can be directly obtained from the error messages or exception codes returned during container startup. If container startup fails, specific error information will be automatically returned, which may include an exception description or exit code. Retrieving startup exception information requires no additional detection or processing steps for the container; it provides direct feedback in the event of a container startup failure.
[0060] Once startup exception information is obtained, container recovery processing can be performed based on the startup exception. Specifically, there are several ways to perform container recovery processing based on startup exception information.
[0061] One optional approach is to determine and invoke the corresponding startup exception handling task for exception recovery based on the startup exception information. Specifically, the corresponding exception recovery task can be determined based on the exit code or exception description included in the startup exception information, according to a preset mapping relationship. For example, exit code "3" can correspond to "container image error," and the corresponding startup exception handling task could be "re-pull the container image," etc.
[0062] Another option is to perform semantic analysis and classification based on the anomaly description in the startup anomaly information, categorize the anomaly into a preset anomaly type, and then use the anomaly handling task corresponding to that anomaly type for anomaly recovery. For example, if the anomaly description includes the keyword "network configuration," then network configuration repair strategies can be automatically applied.
[0063] Optionally, during the process of abnormal recovery of containers based on startup exception information, if the service platform cannot complete the abnormal recovery process based on the startup exception information, for example, if the startup exception recovery task corresponding to the startup exception information is not configured, or if the container startup process still has exceptions after multiple abnormal recovery processes, an alarm message can be generated and sent to external users, such as the service platform administrator or users, for manual intervention in the abnormal recovery process, such as modifying the configuration file or resetting the network connection.
[0064] In this step, by directly obtaining the startup exception information when the container fails to start, and performing targeted exception recovery processing on the container based on this information, the accuracy and efficiency of the exception recovery for the container can be improved. This avoids the inefficient method of simply restarting the container, reduces the service interruption time caused by the failure of container startup, and provides reliable basic support for the normal operation of containerized services.
[0065] In the embodiments described in this specification, service anomaly detection is performed on the containerized services running in the container before the container of the service platform is started. Service detection results are obtained, and the container is started only when the containerized service is normal. This ensures the correct configuration and availability of the containerized services provided by the container. At the same time, in the case of container startup failure, startup anomaly information is directly obtained, and targeted anomaly recovery processing is performed on the container based on the startup anomaly information. This improves the comprehensiveness and accuracy of anomaly identification for the service platform, enhances the targeting and efficiency of anomaly recovery processing, reduces the cost of manual intervention, and ensures the stability and availability of the containerized services that the service platform can provide.
[0066] In one optional embodiment of this specification, service anomaly detection includes service dependency anomaly detection; Perform service anomaly detection on containerized services and obtain service detection results, including: Perform service dependency anomaly detection on containerized services and determine the service detection result based on the response status of other services that the containerized service depends on.
[0067] Dependency anomaly detection is a targeted inspection of external services that containerized services depend on during service anomaly detection. By verifying the availability or normal response of external services, dependency anomaly detection determines whether the containerized service can function correctly and is a crucial component of service anomaly detection. The results of dependency anomaly detection directly impact the accuracy of service detection results. For containerized services, dependency anomaly detection can identify service anomalies caused by dependency service issues before the container starts. For example, for the large language model service within a model service, dependency anomaly detection can check the availability of inference framework services or caching services that the large language model depends on.
[0068] Other services that a containerized service depends on are external services that the containerized service needs to call or connect to during normal operation. These can include database services, API services, and other external model services. The dependencies between these other services and the containerized service are fundamental conditions for the containerized service to function properly. For example, a large language model service within a model service might depend on an inference framework service for text inference capabilities, a caching service for context or historical dialogue records, and a feature service for data feature extraction. The correctness of these dependencies between other services and the containerized service needs to be verified before container startup to ensure the containerized service can function correctly.
[0069] The response status is the feedback result from other external services to the containerized service's requests. This can include, for example, service connection status, service response time, and service response content. The response status can be used to determine whether the external service is functioning correctly and generate corresponding service detection results for further assessment of the containerized service's performance.
[0070] In practical applications, service anomaly detection for containerized services can include dependency anomaly detection. Specifically, dependency anomaly detection can be performed by sending test requests to other services that the containerized service depends on through a predefined dependency service check interface, and obtaining the corresponding response status for evaluation. The execution of dependency anomaly detection can employ parallel detection, sending test requests to multiple dependent services simultaneously; sequential detection, checking each dependent service one by one; or a timeout mechanism to control the detection process of each other service, preventing the overall detection efficiency from being affected by the timeout of a single service.
[0071] In one alternative embodiment, other services on which the containerized service depends may include Large Language Model (LLM) service, vector database service, feature extraction service, etc. Verification of these other services can be performed by sending standardized lightweight test requests to the services and determining their availability based on the response status.
[0072] Specifically, taking the containerized service as a large language model service as an example, the dependency detection process for the large language model service includes the following steps: sending a preset test request (e.g., an HTTP POST request, the request body of which can be {"query":"ping"} or an empty string), and setting a preset timeout threshold (e.g., 3 seconds); expecting the LLM service to return a preset success response under normal conditions (e.g., HTTP status code 200, and the response body containing {"status":"ok"} or a similar health status indicator); if no response is received within the preset timeout period, or the received response status code indicates a server error, or the response body does not contain the expected health status indicator, then the large language model service is determined to be unavailable.
[0073] Once the response status of other services that the containerized service depends on is obtained, the service detection result can be determined based on the response status.
[0074] Specifically, the response status of other services can be compared with a preset normal response standard. If the response status meets the normal standard, the dependent service is considered normal; otherwise, it is determined to be abnormal. Alternatively, the response status of multiple other services can be comprehensively judged. For example, if all dependent services respond normally, the service detection result is normal; if at least one dependent service is abnormal, the service detection result is abnormal. Furthermore, weights can be set according to the importance of other dependent services, with higher weights assigned to critical dependent services. For example, in the large language model service, the inference framework service and the caching service can be considered critical dependencies. An abnormal response status of these services can directly determine that the service detection result is abnormal, while an abnormality of secondary dependent services can only affect the service detection result score, and further judgment on whether it is abnormal can be made based on the score.
[0075] In the embodiments of this specification, by including dependency anomaly detection in the service anomaly detection, it is possible to discover whether other services that the containerized service depends on are available during the anomaly detection process for the containerized service. This avoids the containerized service from failing to run properly due to dependency service anomalies, improves the comprehensiveness and accuracy of service anomaly detection, reduces service anomalies caused by dependency service issues after container startup, reduces service interruption time and manual intervention costs, and provides a reference for anomaly recovery handling of containerized services, making anomaly recovery handling more targeted and efficient. For example, when dependency anomaly detection finds that the LLM service is unavailable, a backup LLM service switching strategy can be triggered immediately.
[0076] In one optional embodiment of this specification, after performing service anomaly detection on the containerized service and obtaining the service detection result, the method further includes: If a containerized service is determined to be abnormal based on the service detection results, service abnormality information is obtained, and abnormal recovery processing is performed on the containerized service based on the service abnormality information.
[0077] Containerized service anomalies are the states determined by service anomaly detection before container startup. These states may include at least one of the key elements such as environment configuration, dependent services, and file integrity failing to meet normal operating conditions. This state indicates that the containerized service cannot provide stable or correct services. For example, abnormal environment variable configuration may cause the containerized service to become unresponsive, or the unavailability of dependent services may cause the containerized service to return an empty result.
[0078] Service exception information is descriptive information about the cause of an anomaly included in the service detection results when a containerized service experiences anomalies. Specifically, service exception information can be a text description of the anomaly, an exception identifier in the form of a tag, or a specific exit code, such as "environment variable missing" or "exit code 2." Service exception information reflects the specific cause of the containerized service anomaly and can serve as the basis for handling anomaly recovery procedures for containerized services.
[0079] In practical applications, after performing service anomaly detection on containerized services and obtaining the service detection results, it can be used to determine whether the containerized service is abnormal. Specifically, if the service detection results include service anomaly information, it indicates that the containerized service is abnormal. For example, the service detection result could be "environment variables are missing, dependent services are unavailable." Alternatively, it can be judged through specific indicators provided in the service detection results. For example, if the environment variable configuration accuracy rate is lower than a preset threshold, or the availability score of dependent services is lower than a threshold, it indicates that the containerized service is abnormal.
[0080] If a containerized service is determined to be abnormal based on service detection results, the container startup process can be paused, and the containerized service can be restored based on the service anomaly information. Specifically, there are several ways to restore the containerized service based on the service anomaly information.
[0081] One optional approach is to determine and invoke the corresponding service exception handling task based on the service exception information to perform exception recovery processing on the containerized service. Specifically, the corresponding exception handling task can be determined based on the exit code or exception description included in the service exception information, according to a preset mapping relationship. For example, exit code "1" can correspond to "environment variable configuration error," and the corresponding service exception handling task could be "reload environment variable configuration"; exit code "2" can correspond to "dependent service unavailable," and the corresponding service exception handling task could be "re-fetch dependent service or switch to an alternative dependency," etc.
[0082] Specifically, the implementation mechanism for "switching to standby dependent services" can be based on the discovery method and traffic scheduling of containerized services. The service platform can maintain a list of available standby services and their access endpoints, which can be dynamically managed by the configuration center or pre-configured in the platform. When "switching to standby dependencies" is required, the exception recovery module can obtain the list of currently available standby service addresses by querying the configuration center or local service registry (such as the DNS discovery mechanism based on Kubernetes Service or the Consul service registry). Subsequently, the recovery engine or its cooperating Sidecar proxy can update the dependency connection configuration or local routing policy of the containerized service, switching traffic from the failed primary service to the selected standby service node. The switching trigger condition for the standby dependency can be based on the statistics of the number of consecutive failures of the primary service (e.g., reaching a preset threshold) and the type of startup exception information or service exception information.
[0083] Another option is to perform semantic analysis and classification based on the exception description in the service exception information, categorize the exception into a preset exception type, and use the exception handling task corresponding to that exception type for exception recovery. For example, if the exception description includes the keyword "environment variables," the configuration can be automatically reloaded.
[0084] Optionally, during the process of handling the abnormal recovery of containerized services based on service abnormal information, if the service platform cannot complete the abnormal recovery process based on the service abnormal information, for example, if the service abnormal recovery task corresponding to the service abnormal information is not configured, or if the containerized service still has abnormalities after multiple abnormal recovery processes, an alarm message can be generated and sent to external users, such as the service platform administrator or users, to manually intervene in the abnormal recovery process, such as reconfiguring the environment variable file or specifying new dependent services.
[0085] In the embodiments of this specification, by further obtaining service anomaly information when the containerized service is determined to be abnormal based on the service detection results, and performing anomaly recovery processing on the containerized service based on the service anomaly information, anomalies in the containerized service can be detected and resolved before the container starts. This improves the efficiency and accuracy of anomaly handling, reduces the number of container restarts caused by containerized service anomalies, reduces service interruption time and manual intervention costs, and ensures the correctness and availability of the containerized service through anomaly recovery, thereby improving the stability and service availability of the service quality provided by the service platform.
[0086] In one optional embodiment of this specification, based on service anomaly information, anomaly recovery processing is performed on containerized services, including: Based on the service exception information, the corresponding service exception handling task is invoked to perform exception recovery processing on the containerized service.
[0087] Service exception handling tasks, corresponding to service exception information, are a set of recovery strategies or operations used to handle exceptions in containerized services. The types of service exception handling tasks can include files, code, threads, instructions, etc., and the specific type can be determined based on the different containerized service exceptions in the actual application. Service exception handling tasks correspond to service exception information; for example, they can correspond to the exit code included in the service exception information, or the exception category determined by the description information of the service exception information. For example, when the service exception information includes the exit code "1", the corresponding service exception handling task could be "reload environment variable configuration"; when the description information of the service exception information includes the keyword "dependent service unavailable", the corresponding service exception handling task could be "re-fetch dependent services or switch to alternative dependencies", etc.
[0088] In practical applications, once service exception information is identified, the corresponding service exception handling task can be determined and invoked based on that information. Specifically, a pre-defined mapping relationship between service exception information and service exception handling tasks can be used to match the exit code or exception type in the service exception information with the pre-defined exception handling task. In other words, upon obtaining service exception information, the pre-defined mapping relationship can be automatically searched to determine the service exception handling task corresponding to that exception information.
[0089] Optionally, before invoking the service exception handling task corresponding to the service exception information to perform exception recovery processing on the containerized service based on the service exception information, the following may also be included: Based on the historical service detection results of containerized services, configure a mapping table between service anomaly information and service anomaly handling tasks; Based on service exception information, the corresponding service exception handling task is invoked to perform exception recovery processing on the containerized service, including: Based on the service exception information, query the mapping relationship table to determine the service exception handling task; Invoke the service exception handling task to perform exception recovery processing on the containerized service.
[0090] Once a service exception handling task has been identified, that task can be invoked to perform exception recovery operations on the containerized service.
[0091] Specifically, the service exception handling task can be executed to perform exception recovery for containerized services, and the execution can be tailored to the specific type of the service exception handling task. For example, if the service exception handling task is "reload environment variable configuration", the service platform can automatically execute the loading instruction, load the environment variable configuration file, and update the environment variable settings of the containerized service, enabling the containerized service to run normally. If the service exception handling task is "switch to a backup dependency service", the service platform can automatically execute a dependency switching thread, connect to the current dependency service, and enable the backup dependency service, ensuring that the containerized service can continue to provide normal service.
[0092] In the embodiments described in this specification, by invoking the corresponding service exception handling task based on service exception information to perform exception recovery processing on containerized services, targeted recovery of containerized service exceptions can be achieved, avoiding the simple recovery method of restarting the container, improving the accuracy and efficiency of service exception recovery, and reducing service interruption time caused by containerized service exceptions. At the same time, since there is a correspondence between the service exception handling task and the service exception information, the targetedness and effectiveness of exception recovery processing are ensured, reducing the cost of manual intervention and improving the stability and service availability of the service platform.
[0093] In one optional embodiment of this specification, based on startup exception information, exception recovery processing is performed on the container, including: Based on the startup exception information, the startup exception handling task corresponding to the startup exception information is invoked to perform exception recovery processing on the container.
[0094] The startup exception handling task corresponding to the startup exception information is a set of recovery strategies or operations used to handle container startup failure exceptions. The type of startup exception handling task can also include files, code, threads, instructions, etc., and the specific type can be determined based on the different container startup exceptions in the actual application. The startup exception handling task corresponds to the startup exception information; for example, it can correspond to the exit code included in the startup exception information, or to the exception category determined by the description information of the startup exception information. For example, when the startup exception information includes the exit code "5", the corresponding startup exception handling task could be "re-pull the script file of the container image"; when the description information of the startup exception information includes the keyword "network configuration error", the corresponding startup exception handling task could be "reconfigure network parameters".
[0095] In practical applications, once startup exception information is identified, the corresponding startup exception handling task can be determined and invoked based on this information. Specifically, a pre-defined mapping relationship between startup exception information and startup exception handling tasks can be used to match the exit code or exception type in the startup exception information with the pre-defined exception handling task. In other words, upon obtaining startup exception information, the pre-defined mapping relationship can be automatically searched to determine the startup exception handling task corresponding to that exception information.
[0096] Optionally, before invoking the startup exception handling task corresponding to the startup exception information to perform exception recovery processing on the container based on the startup exception information, the following may also be included: Based on the historical startup detection results of containers, configure a mapping table between startup exception information and startup exception handling tasks; Based on the startup exception information, the corresponding startup exception handling task is invoked to perform exception recovery processing on the container, including: Based on the startup exception information, query the mapping table to determine the startup exception handling task; Invoke the startup exception handling task to perform exception recovery processing on the container.
[0097] Once a startup exception handling task has been identified, that task can be invoked to perform exception recovery operations on the container.
[0098] Specifically, invoking the startup exception handling task to perform exception recovery for the container can be tailored to the actual type of the startup exception handling task. For example, when the startup exception handling task is "reload environment variable configuration," the startup platform can automatically execute the loading instruction, load the environment variable configuration file, and update the environment variable settings for containerized startup, enabling containerized startup to run normally. When the startup exception handling task is "re-pull container image," the container platform can automatically pull the correct container image from the specified image repository and reload it into the container environment, enabling the container to start normally. When the startup exception handling task is "reconfigure network parameters," the service platform can automatically update the container's network configuration parameters, such as IP address and port settings, ensuring that the container can connect to the correct network environment.
[0099] In the embodiments of this specification, by invoking the corresponding startup exception handling task based on startup exception information to perform exception recovery processing on the container, targeted recovery of container startup failure problems can be achieved, avoiding the simple recovery method of restarting the container, improving the accuracy and efficiency of container startup exception recovery, reducing service interruption time caused by container startup failure. At the same time, since there is a correspondence between the startup exception handling task and the startup exception information, the targeting and effectiveness of exception recovery processing are ensured, reducing the cost of manual intervention and improving the stability and service availability of the service platform.
[0100] In an optional embodiment of this specification, after starting the container and determining that the containerized service is functioning normally based on service testing results, the method further includes: If startup is successful, container operation checks will be performed based on the preset detection cycle to obtain the operation check results; If a container is found to be malfunctioning based on the results of operational monitoring, abnormal operation information is obtained, and abnormal recovery processing is performed on the container based on the abnormal operation information.
[0101] The preset detection period is a time interval pre-set in the service platform to perform container runtime checks after the container has successfully started and is running. The preset detection period can be a fixed length, such as every 5 minutes or every 10 minutes; or it can be a dynamically adjusted period based on the type of service provided. For example, a shorter detection period can be set for services with high availability requirements, while a longer detection period can be set for services with higher stability requirements. Setting the preset detection period determines the frequency and timeliness of container runtime checks, thereby enabling timely detection of operational anomalies during container operation.
[0102] Container runtime monitoring refers to the periodic runtime status checks performed on containerized services running within a container after successful container startup. Specifically, container runtime monitoring can verify the availability and correctness of containerized services by calling the runtime status (health) check interfaces provided by the containerized services. For example, the " / health" interface of the containerized service can be called every 5 minutes to check the container runtime status code or response content returned by the interface. Container runtime monitoring can include checks on various aspects such as service response time, response content, and interface availability.
[0103] The runtime detection results are output during the container runtime detection process. They are assessment information about the runtime status of containerized services and can be used to determine whether the container is running normally. If there are abnormalities in the container runtime, the results can also include corresponding runtime abnormality information. The runtime abnormality information can reflect information such as the type of runtime abnormality and the cause of the runtime abnormality. For example, the runtime detection result can be "Service response is normal, response time is within the threshold range", indicating that the container is running normally; or it can be "Service response timed out, interface returned error code 404", indicating that the container is running abnormally.
[0104] Container runtime anomalies refer to the state where a container, after successful startup, fails to provide services normally during operation. These anomalies can include service response timeouts, API errors, and unresponsive services. Container runtime anomalies indicate that a problem has occurred while the containerized service is running within the container, requiring anomaly recovery handling. The determination of container runtime anomalies can be based on runtime monitoring results. When runtime monitoring results show that the state of the containerized service running within the container does not meet preset normal standards, it is determined to be a container runtime anomaly.
[0105] Exception information is descriptive information output by the container runtime detection process when a container encounters an exception, describing the cause of the exception. Exception information can be a text description, a tag-based exception identifier, or a specific exit code, such as "service response timeout" or "exit code 7". Exception information directly reflects the specific reason for the container's runtime exception and can serve as a basis for handling container exception recovery.
[0106] In practical applications, once a container has successfully started, container runtime checks can be performed on the container based on a preset check period in the service platform. Specifically, system-level timers or service-level scheduled tasks can be used to ensure that container runtime checks are automatically executed at the end of each preset check period. For example, when the preset check period is 5 minutes, the service platform will automatically perform container runtime checks on the containerized service every 5 minutes through the runtime status check interface. For multiple containers, single-threaded sequential checks or multi-threaded parallel checks can be used to improve check efficiency.
[0107] During container runtime monitoring, the monitoring can be performed by calling the monitoring interface provided by the containerization service. The interface returns container runtime status information or service response content, which is then analyzed according to preset judgment criteria to generate an evaluation result regarding the container's runtime status. Specifically, this may include steps such as sending a container runtime monitoring request, receiving service response data, parsing the response content, and comparing it with preset normal standards. For example, an "HTTP GET" request can be sent to the " / health" interface of the containerized service running in the container to obtain the service response status code or response body. Further analysis can then determine whether the status code is "200" or whether the response body contains information such as "healthy," thereby determining the container runtime monitoring result.
[0108] The " / health" health check interface provided by the containerized service can be configured internally to aggregate or reflect the status of key pre-check items. That is, when the health check interface is called, a lightweight runtime self-check can be performed. The exit code or status information returned is consistent with the exit code system defined in the service anomaly detection phase, or has a clear mapping relationship (for example, exit code 1 can indicate an environment configuration error, exit code 2 can indicate a dependent service anomaly, etc.). It should be noted that the health check of the " / health" interface mainly focuses on verifying the basic operational capabilities of the containerized service (such as process liveness, memory status, basic dependency connectivity, etc.). Under normal circumstances, it does not typically trigger a complete end-to-end service inference chain to ensure higher efficiency and lower load in the check process.
[0109] After performing container operation checks on the container based on a preset detection cycle and obtaining the operation check results, it can be determined whether the container is operating abnormally based on the operation check results.
[0110] Specifically, based on the operation detection results, a comprehensive evaluation can be conducted through the set anomaly judgment rules. If the operation detection results contain operation anomaly information, it indicates that there is an anomaly in the container operation. For example, the operation detection results may include "exit code 7". Alternatively, it can be judged by whether the key indicators (such as response time, response status code, etc.) included in the operation detection results are lower or higher than a preset threshold. For example, if the response time included in the operation detection results exceeds 3 seconds, it can be determined that the container operation is abnormal.
[0111] If the runtime detection results include runtime exception information, it indicates that the current container is experiencing an anomaly. Further recovery processing can then be performed based on this exception information. The runtime exception information can be determined by parsing the exception description information contained in the runtime detection results, or by directly obtaining the exception cause description from the container runtime detection process. For example, if the " / health" interface of a containerized service returns "500 Internal Server Error," the runtime exception information can be determined as "Internal service error, error code 500."
[0112] Once the runtime exception information is identified, recovery processing can be performed on the container based on that information. Specifically, there are several ways to perform container recovery processing based on runtime exception information.
[0113] One optional approach is to determine and invoke a runtime exception handling task based on the runtime exception information to perform exception recovery processing on the container. Specifically, the corresponding runtime exception handling task can be determined based on the exit code or exception description included in the runtime exception information, according to a preset mapping relationship. For example, when the runtime exception information is "Internal service error, error code 500", the corresponding runtime exception handling task could be "Restart container"; when the runtime exception information is "Dependency service unavailable", the corresponding runtime exception handling task could be "Switch to backup dependency service", and so on.
[0114] Another option is to perform semantic analysis and classification based on the exception description in the runtime exception information, categorize the exception into a preset exception type, and use the exception handling task corresponding to that exception type for exception recovery. For example, if the exception description includes the keyword "container image error," the container image can be automatically re-pulled.
[0115] Optionally, during the process of abnormal recovery of containers based on runtime abnormal information, if the service platform cannot complete the abnormal recovery process based on the service abnormal information, for example, if the runtime abnormal recovery task corresponding to the runtime abnormal information is not configured, or if the container's running status is still abnormal after multiple abnormal recovery processes, an alarm message can be generated and sent to external users, such as the service platform's administrator or users, for manual intervention in the abnormal recovery process, such as modifying configuration files or resetting network connections.
[0116] In the embodiments described in this specification, by performing container operation checks based on a preset detection cycle after the container starts successfully, abnormal container operation can be detected in a timely manner during container operation, avoiding service interruptions caused by persistent abnormalities and improving the availability and stability of the service platform. By obtaining the operation check results and determining the container operation abnormality, service anomalies can be accurately identified, reducing false positives and false negatives. By obtaining the operation anomaly information and performing targeted anomaly recovery processing based on this information, the simple method of restarting the container can be avoided, improving the accuracy and efficiency of anomaly recovery and reducing service interruption time. At the same time, the periodic container operation check can form a closed loop with the service anomaly check before container startup, covering the entire lifecycle of containerized services, ensuring the high availability of containerized services, reducing the cost of manual intervention, and improving the stability and service availability of the service platform.
[0117] In one optional embodiment of this specification, based on runtime exception information, exception recovery processing is performed on the container, including: Based on the runtime exception information, the corresponding runtime exception handling task is invoked to perform exception recovery processing on the container.
[0118] The runtime exception handling task corresponding to the runtime exception information is a set of recovery strategies or operations used to handle exceptions that occur during container operation. The type of runtime exception handling task can also include files, code, threads, instructions, etc., and the specific type can be determined based on the specific circumstances of the container runtime exception in the actual application. The runtime exception handling task corresponds to the runtime exception information; for example, it can correspond to the exit code included in the runtime exception information, or to the exception category determined by the description information of the runtime exception information. For example, if the runtime exception information includes the exit code "7", the corresponding runtime exception handling task could be "restart the container"; if the description information of the runtime exception information includes the keyword "container image error", the corresponding runtime exception handling task could be "automatically re-pull the container image", etc.
[0119] In practical applications, once runtime exception information is identified, the corresponding runtime exception handling task can be determined and invoked based on this information. Specifically, a pre-defined mapping relationship between runtime exception information and runtime exception handling tasks can be used to match the exit code or exception type in the runtime exception information with the pre-defined exception handling task. In other words, upon obtaining runtime exception information, the pre-defined mapping relationship can be automatically searched to determine the runtime exception handling task corresponding to that exception information.
[0120] Optionally, before invoking the runtime exception handling task corresponding to the runtime exception information to perform exception recovery processing on the container based on the runtime exception information, the following may also be included: Based on the historical operation detection results of containers, configure a mapping table between operation anomaly information and operation anomaly handling tasks; Based on the runtime exception information, the corresponding runtime exception handling task is invoked to perform exception recovery processing on the container, including: Based on the runtime exception information, query the mapping relationship table to determine the runtime exception handling task; Invoke the runtime exception handling task to perform exception recovery processing on the container.
[0121] Once a runtime exception handling task has been identified, that task can be invoked to perform exception recovery operations on the container.
[0122] Specifically, the execution of exception handling tasks for container recovery can be tailored to the specific type of the task. For example, if the task is "restart container," the service platform can automatically execute the restart command to restart the container and restore its normal operation. If the task is "re-pull container image," the container platform can automatically pull the correct container image from the specified image repository and reload it into the container environment to ensure normal operation. If the task is "reconfigure network parameters," the service platform can automatically update the container's network configuration parameters, such as IP address and port settings, to ensure the container can connect to the correct network environment.
[0123] In the embodiments described in this specification, by invoking the corresponding runtime exception handling task based on runtime exception information to perform exception recovery processing on the container, targeted recovery of container runtime exception problems can be achieved, avoiding the simple recovery method of restarting the container, improving the accuracy and efficiency of container runtime exception recovery, and reducing service interruption time caused by container runtime exceptions. At the same time, since there is a correspondence between runtime exception handling tasks and runtime exception information, the targeting and effectiveness of exception recovery processing are ensured, reducing the cost of manual intervention and improving the stability and service availability of the service platform.
[0124] In one optional embodiment of this specification, container operation detection is performed based on a preset detection cycle to obtain operation detection results, including: Based on a preset detection cycle, container operation detection is performed to obtain multiple operation detection results; When a container is determined to be malfunctioning based on operational monitoring results, the malfunction information is obtained, including: For multiple operational detection results, count the number of corresponding container operational anomalies; When the number of anomalies reaches a preset threshold, runtime anomaly information is obtained.
[0125] Multiple operational detection results are generated by periodically checking the container's operational status at preset time intervals. Each operational status check generates an independent result, forming a set of multiple operational detection results. These multiple results collectively reflect the container's operational stability over a period of time, avoiding misjudgments caused by the randomness of a single check. For example, with a preset detection period of 5 minutes, the service platform will perform a container operational check every 5 minutes, and after 5 consecutive executions, it will obtain 5 operational detection results, each of which can include the container's operational status.
[0126] The number of container runtime anomalies is counted among multiple runtime monitoring results. Specifically, when the service platform receives multiple runtime monitoring results, it performs anomaly assessment on the container's runtime status based on each result and records whether the container is running abnormally. The results of those deemed abnormal are then accumulated. For example, if 3 out of 5 consecutive runtime monitoring results show service response timeouts and the interface returns error code 500, the number of container runtime anomalies is 3. The number of container runtime anomalies can be counted continuously or using a sliding window approach.
[0127] A preset threshold is a pre-defined acceptable number of anomalies in container operation anomaly detection, used to determine whether a container operation anomaly has actually occurred. The preset threshold can be set based on various factors. For example, for services with high availability requirements, a lower preset threshold can be set, such as two records indicating a container operation anomaly; for services with high stability, a higher preset threshold can be set, such as four records indicating a container operation anomaly; alternatively, a reasonable preset threshold can be set based on historical operation data, by analyzing the frequency of historical anomalies. For example, based on analysis of historical container anomaly data, it was found that 95% of anomalies occur within three detections, so the preset threshold can be set to 3; furthermore, the preset threshold can be differentiated according to factors such as service type and service importance, for example, a lower threshold can be set for critical services, and a higher threshold for non-critical services. This specification does not specifically limit this in the embodiments.
[0128] In practical applications, for multiple operational detection results, the service platform will perform anomaly judgment on each operational detection result separately, and accumulate the operational detection results judged as abnormal to form the number of abnormalities in container operation. When the number of abnormalities reaches a preset threshold, the container operation is determined to be abnormal and operational abnormality information is generated.
[0129] Optionally, the preset threshold can be set as the number of consecutive container operation failures or the number of container operation failures within a sliding window. For example, if the preset threshold for the number of consecutive failures is set to 3, it means that if three consecutive container operation checks show container operation failures (e.g., displaying service response timeout, interface returning error code 500), the current container is judged to be operating abnormally, and corresponding operation failure information is generated. If the sliding window is set to 5 and the preset threshold for the number of failures within the window is 3, it means that if three or more of the five consecutive container operation checks show container operation failures, the current container is judged to be operating abnormally, and corresponding operation failure information is generated.
[0130] In the embodiments of this specification, multiple operation detection results are obtained by performing container operation detection based on a preset detection cycle. The number of abnormalities in container operation is counted for multiple operation detection results. When the number of abnormalities reaches a preset threshold, operation abnormality information is obtained. This can effectively avoid misjudgment of abnormalities caused by the randomness of a single detection, and improve the accuracy and reliability of container operation abnormality judgment. By flexibly setting the preset threshold, the strictness of abnormality judgment can be customized according to the needs of different containerized services. While ensuring the timely response of highly available containerized services, excessive alarms for low-priority containerized services are avoided, which improves the stability and service availability of the service platform and makes abnormality recovery processing more intelligent and accurate.
[0131] In an optional embodiment of this specification, after performing exception recovery processing on the container based on the startup exception information, the method further includes: An exception handling log is generated based on the startup exception information and the exception recovery process for the container. Visualize the exception handling logs to obtain a visual report of the exception handling process. Send the visualization report to the target client.
[0132] The anomaly recovery process is a complete flow of targeted repair and recovery operations performed on a container based on startup anomaly information after a container startup failure. It may include obtaining startup anomaly information, determining the corresponding anomaly recovery task, and executing anomaly recovery operations. Optionally, anomaly recovery processing may also include anomaly recovery processing of containerized services based on service anomaly information after service anomaly detection, and anomaly recovery processing of containers based on runtime anomaly information after container runtime detection.
[0133] An exception handling log is a structured data file that records key information and operational details during the exception recovery process. It may include the content of the startup exception information, the execution time of the exception recovery task, and the specific steps of the recovery operation. For example, the exception handling log may record "Startup exception information: Exit code 3; Exception recovery task: Re-pull container image; Execution time: 20XX-10-29 15:30:22", etc.
[0134] Visualization processing is the process of transforming structured data in anomaly handling logs into intuitive and easy-to-understand graphs or charts through statistical analysis, data collection, and extraction methods. This can include extracting key indicators from the logs, performing statistical analysis on the extracted data, and generating corresponding visualization charts. For example, visualization processing can statistically analyze the frequency of different anomaly types over a period of time (e.g., a week) and generate a pie chart showing the distribution of anomaly types; it can also extract the average processing time for anomaly recovery and generate a line graph showing the time-varying changes.
[0135] A visualization report is a comprehensive information display that presents the results of visualization processing in graphical or tabular form. It may include anomaly type distribution charts, recovery success rate statistics tables, and anomaly recovery time trend charts. The combination of charts and tables clearly displays key indicators and trends in anomaly handling. For example, a visualization report may include a pie chart of "anomaly type distribution" to show the proportion of different anomaly types (such as environment variable configuration errors, unavailable dependent services, container image issues, etc.); it may also include a line chart of "anomaly recovery time trend" to show the average time change of anomaly recovery over different time periods; and it may include a table of "anomaly recovery success rate" to list the recovery success rate corresponding to different anomaly types.
[0136] The target client is the terminal device or application that receives and views the anomaly handling visualization report. It is used to receive and display relevant information about anomaly handling and can be a service platform's management console, mobile application, web interface, etc. The target client is the receiver of the visualization report, enabling quick and accurate anomaly location based on the report and providing improvements for subsequent anomaly detection and recovery.
[0137] In practical applications, after performing exception recovery processing on containers or containerized services, exception handling logs can be generated based on the corresponding exception information and exception recovery process. Specifically, the generation process of exception handling logs can include the content of exception information recorded during the exception recovery process, the name and parameters of the exception recovery task, the start and end times of execution, and other key data, and integrate this data into structured log entries according to a preset log format.
[0138] Optionally, exception handling logs can be generated synchronously or asynchronously. Synchronous logging generates logs in real-time during each step of the exception recovery process, ensuring log integrity and timeliness. Asynchronous logging generates logs in batches after exception recovery is complete, reducing the performance impact on real-time processing. Furthermore, exception handling log generation can be tiered based on the severity of the exception in the corresponding container or containerized service. For example, for severe exceptions, more detailed execution steps are recorded; for minor exceptions, only key information is recorded to balance log detail and storage overhead.
[0139] Once an exception handling log has been generated, it can be further visualized to obtain a visual report of the exception handling process.
[0140] Specifically, the visualization process can include extracting key metrics from anomaly handling logs, performing statistical analysis on the extracted data, and generating corresponding visualization charts. The specific execution steps can include: collecting key data such as anomaly type and recovery time from the anomaly handling logs; performing statistical calculations on the key data to obtain statistical results, such as calculating the occurrence frequency of various anomalies, average recovery time, and recovery success rate; and converting the statistical results into a target visualization format, such as using a bar chart to display the distribution of anomaly types or using a line chart to display the recovery time trend.
[0141] In practical applications, visualization processing can employ a variety of different technical solutions. For example, it can include interactive charts based on open-source data visualization libraries, static reports generated based on reporting engines, or real-time monitoring and analysis based on data dashboards. The specific visualization processing adopted can be flexibly configured based on the actual service platform and the log content to be presented. This specification does not impose specific limitations on this aspect in the embodiments.
[0142] In the embodiments described in this specification, after performing anomaly recovery processing on containers based on startup anomaly information, anomaly handling logs are further generated. These logs are then visualized to generate a visual report, which is sent to the target client. This improves the traceability and analyzability of anomaly handling for containers or containerized services within the service platform, enabling the service platform to gain a more comprehensive understanding of the overall anomaly handling situation. The anomaly handling logs provide detailed data support for subsequent anomaly analysis and optimization, avoiding the inaccuracies of manual memorization and recording. Visualization transforms the raw log data into intuitive charts and tables, allowing platform administrators to quickly identify anomaly patterns and trends without in-depth analysis of the raw log data. Sending the visual report to the target client allows service platform administrators to promptly understand the anomaly handling situation, facilitating decision-making and intervention. This improves the efficiency and accuracy of anomaly handling, reduces manual intervention costs, enhances the stability of the service platform and the availability of containerized services, and makes anomaly handling more transparent and manageable, providing a data foundation for the continuous optimization of the service platform.
[0143] In one optional embodiment of this specification, the exception handling log includes the processing results of exception recovery processing for the container; After generating exception handling logs based on startup exception information and the exception recovery process for containers, the following is also included: If the processing result is an abnormal recovery failure record, an abnormal alarm signal is generated based on the abnormal recovery failure record and sent to the target client; In response to the exception recovery instructions returned by the target client, perform exception recovery processing for the container.
[0144] The results of anomaly recovery processing are the execution status and effectiveness evaluation information obtained after performing anomaly recovery processing on a container. This information may include whether the anomaly recovery operation was successful, execution time, and scope of impact. The results directly reflect the actual effectiveness of the anomaly recovery processing and can be used to evaluate the effectiveness of the anomaly recovery strategy. Specifically, the results may include status indicators such as "success," "failure," or "unable to verify," and may also include more detailed processing data, such as "Three consecutive executions of the anomaly recovery task: re-pulling the container image, all failed."
[0145] Anomaly recovery failure logs are detailed records generated during anomaly recovery processes when the recovery operation fails to achieve the expected results, i.e., fails to eliminate the anomaly. These logs may include data such as the failed recovery task, the reason for the failure, the time of failure, and the number of recovery attempts. Anomaly recovery failure logs are a concrete manifestation of recovery failures and provide clear evidence for subsequent manual intervention and strategy optimization.
[0146] Anomaly alarm signals are specific signals automatically generated and sent by the service platform to indicate failure in anomaly recovery processing. These signals can take various forms, including text messages, sound alerts, and system notifications. Anomaly alarm signals can be used to trigger subsequent processes after anomaly recovery failure, allowing relevant personnel to be notified in a timely manner for manual intervention.
[0147] Specifically, abnormal alarm signals can be presented in various ways, such as a system notification message that says "Container abnormal: Container ID 0x123456, abnormal recovery processing failed"; or an email alarm that says "Container abnormal: Container ID 0x123456, abnormal recovery processing failed, email sent to administrator".
[0148] Anomaly recovery instructions are manual commands generated and returned by administrators or other personnel to guide the service platform in performing specific anomaly recovery operations after the target client receives an anomaly alarm signal. Specifically, anomaly recovery instructions can include newly written executable code, anomaly handling tasks, service files, or environment configuration parameters. For example, anomaly recovery instructions might include a configuration command such as "Reconfigure container network parameters: set IP address to 192.168.1.100, port to 8080"; or an execution command such as "Execute a specific repair script: / scripts / fix_network.sh".
[0149] In practical applications, if the processing results in the exception handling log include records of exception recovery failures, an exception alarm signal can be generated based on the exception recovery failure records and sent to the target client.
[0150] Specifically, the generation and transmission of abnormal alarm signals can be achieved in various ways. The service platform can automatically generate alarm content containing key information such as the abnormality type, container ID, and failure reason based on abnormal recovery processing failure records, using a preset alarm signal format. The alarm signal is then sent through different channels, including but not limited to system notifications, email notifications, SMS alerts, and telephone notifications. System notifications can be displayed in real-time on the service platform's management console interface for easy viewing by administrators; email notifications can be sent to a preset administrator email address to ensure alarms are received even when the service platform is not in use; SMS alerts can be sent to the administrator's mobile device to ensure timely notification in urgent situations; and telephone notifications can automatically call the administrator in cases of severe abnormalities to ensure high-priority handling of the abnormality. The specific method used to send alarm signals can be flexibly determined according to actual abnormality handling needs, and this specification does not impose specific limitations on this.
[0151] Upon receiving an exception recovery handling instruction from the target client, the system can respond to the instruction and perform exception recovery handling.
[0152] Specifically, when the service platform receives the exception recovery handling instruction returned by the target client, it can further parse the exception recovery handling instruction to obtain the specific execution operation, which may include loading newly written code, executing a specific exception recovery task, applying new service files or environment configuration parameters, etc.
[0153] For example, when the exception recovery handling instructions include newly written code, the service platform can automatically compile and deploy the code to the container's runtime environment; when it includes specific exception recovery tasks, the service platform can start the corresponding task execution thread to perform predefined exception recovery operations; when it includes service files, the service platform can upload the service files to the specified location in the container and reload them; when it includes environment configuration parameters, the service platform can update the container's environment configuration parameters and restart the corresponding containerized service.
[0154] In practical applications, depending on the content of the exception recovery processing instructions, they can be executed in parallel. For example, the service platform can process multiple exception recovery processing instructions at the same time; they can also be executed according to the priority order or the order in which they are received.
[0155] In the embodiments described in this specification, when the processing result of the exception handling log is a record of exception recovery failure, an exception alarm signal is generated and sent to the target client. The exception recovery is then performed in response to the exception recovery instructions returned by the target client. This improves the comprehensiveness and accuracy of exception handling, avoids prolonged interruptions to containerized services due to automatic recovery failures, and ensures that problems can be quickly detected and handled through a timely alarm mechanism. Furthermore, targeted recovery based on manually returned exception recovery instructions improves the success rate and efficiency of exception recovery, reduces service interruption time caused by container or containerized service exceptions, and provides the service platform with more flexible and targeted exception handling capabilities, making the exception handling process more automated and intelligent.
[0156] It should be noted that, in one or more embodiments of this specification, "service anomaly detection" primarily focuses on checking the basic operational capabilities of containerized services, which may include environment configuration, connectivity and basic response of directly dependent services, and integrity of critical files. The design of the " / health" interface also follows this principle, emphasizing the return of the basic health status of the service. Of course, in specific implementation, the availability checks of each critical component can be included in the "dependent service" detection scope according to actual detection and processing needs.
[0157] In one optional embodiment of this specification, a timing description of the execution process of an exception handling method is provided. Specifically, see [link to documentation]. Figure 2 , Figure 2 This specification illustrates a timing diagram of the execution process of an exception handling method according to an embodiment of the present invention, as shown below. Figure 2 As shown.
[0158] The startup detection phase can include service anomaly detection for containerized services, as well as container startup detection during the container startup process.
[0159] During the service anomaly detection process, the service platform triggers the service anomaly detection process through a detection script (such as ready_check.sh). This process may include environment variable detection, and return exit code 1 if the environment variable is abnormal; it may also include dependency service detection, and return exit code 2 if the dependency service is abnormal, etc., and then return the service detection result.
[0160] During the container startup detection process, the service platform obtains the container startup results through command-line tools (such as Algo cmd).
[0161] The periodic inspection phase may include periodically inspecting the operational status of the container.
[0162] During periodic testing, the container can be started and tested repeatedly, and the results of the container's operation can be obtained through the testing interface.
[0163] Corresponding to the above method embodiments, this specification also provides service platform embodiments. Figure 3 A schematic diagram of the structure of a service platform provided in one embodiment of this specification is shown. Figure 3 As shown, the service platform includes a container 302, a service anomaly detection module 304, a container startup module 306, and a container anomaly recovery module 308. Container 302 is used to run containerized services; The service anomaly detection module 304 is used to respond to the service anomaly detection command, perform service anomaly detection on the containerized service, obtain the service detection result, and generate a container startup command if the containerized service is determined to be normal based on the service detection result, and send the container startup command to the container startup module 306. The container startup module 306 is used to start the container in response to the container startup command, and generate a container exception recovery command in the event of container startup failure, and send the container exception recovery command to the container exception recovery module 308. The container exception recovery module 308 is used to respond to the container exception recovery command, obtain the startup exception information, and perform exception recovery processing on the container based on the startup exception information.
[0164] A container is a lightweight, portable unit of software execution that encapsulates an application and its dependencies in an isolated environment, ensuring that the application or service runs consistently across different scenarios. Containers can include Docker containers, as well as operating system containers, security-enhanced containers, and more.
[0165] The service anomaly detection module is a component within the service platform used to perform service anomaly detection functions. Responding to service anomaly detection commands, it performs multi-dimensional anomaly detection on containerized services, including environment variable configuration checks, dependency service availability checks, and file integrity checks, thereby obtaining service detection results. The service anomaly detection module can perform anomaly detection on containerized services before container startup, ensuring the correctness and availability of containerized services. If the service anomaly detection module detects an anomaly in the containerized service, it generates service anomaly information and triggers the anomaly recovery process for the containerized service; when the containerized service is detected as normal, it generates a container startup command and notifies the container startup module to start the container, ensuring the correct configuration and availability of the containerized services provided by the container.
[0166] The container startup module is a component in the service platform used to perform container startup operations. Responding to container startup commands, it calls the container's runtime interface to start the container based on preset container configuration parameters (such as container image address, resource limits, network configuration, etc.). During startup, it loads preset environment variables, connects dependent services, and initializes model files, ensuring the corresponding containerized service can run correctly within the container. The container startup module can start the container after the service anomaly detection module confirms the containerized service is normal. In the event of a container startup failure, it generates a container anomaly recovery command, notifying the container anomaly recovery module to handle the anomaly, avoiding the inefficient method of simply restarting the container and improving the accuracy and efficiency of anomaly handling.
[0167] The container failure recovery module is a component in the service platform used to perform container failure recovery operations. It responds to container failure recovery commands, obtains startup failure information, and performs recovery processing based on this information. This may include, but is not limited to, restarting the container, re-pulling the container image, reloading container resource configurations, and switching to a backup network configuration. The container failure recovery module can perform targeted recovery processing based on startup failure information when a container fails to start, avoiding simply restarting the container.
[0168] Service anomaly detection commands are instructions used in the service platform to trigger service anomaly detection functions. They can be received and sent by the client or automatically generated by the service platform to notify the service anomaly detection module to perform service anomaly detection on containerized services and obtain service detection results.
[0169] The container startup command is a command in the service platform used to trigger the container startup operation. It can be generated by the service anomaly detection module when the containerized service is determined to be normal based on the service detection results. The container startup command is then sent to the container startup module to notify the container startup module to start the container.
[0170] The container exception recovery command is a command in the service platform used to trigger the container exception recovery operation. It can be generated by the container startup module in the event of a container startup failure. The container exception recovery command is sent to the container exception recovery module to notify the container exception recovery module to perform exception recovery processing.
[0171] The service platform provided in the embodiments of this specification, through its included containers, service anomaly detection module, container startup module, and container anomaly recovery module, and the interaction between them via service anomaly detection commands, container startup commands, and container anomaly recovery commands, achieves service anomaly detection for containerized services before container startup, ensuring the correct configuration and availability of containerized services. Simultaneously, when a container startup fails, it directly obtains startup anomaly information and performs targeted anomaly recovery processing based on this information. This improves the comprehensiveness and accuracy of anomaly identification, enhances the targeted processing efficiency of anomaly recovery, reduces manual intervention costs, and guarantees the stability and availability of the containerized services provided by the service platform. It avoids service anomalies caused by service configuration errors or dependency issues, thereby improving the stability and service availability of the service platform.
[0172] The above is an illustrative scheme of a service platform according to this embodiment. It should be noted that the technical solution of this service platform and the technical solution of the above-described exception handling method belong to the same concept. For details not described in detail in the technical solution of the service platform, please refer to the description of the technical solution of the above-described exception handling method.
[0173] In one optional embodiment of this specification, the service platform further includes a containerized service anomaly recovery module; The service anomaly detection module is also used to respond to the service anomaly detection command, perform service anomaly detection on the containerized service, obtain the service detection result, and generate a containerized service anomaly recovery command if the containerized service is determined to be abnormal based on the service detection result, and send the containerized service anomaly recovery command to the containerized service anomaly recovery module.
[0174] The containerized service exception recovery module is used to respond to the containerized service exception recovery command, obtain service exception information, and perform exception recovery processing for the containerized service based on the service exception information.
[0175] The containerized service exception recovery module is a component in the service platform used to perform containerized service exception recovery operations. It can respond to containerized service exception recovery instructions, obtain service exception information, and perform exception recovery processing for containerized services based on the service exception information. This may include, but is not limited to, reloading configuration files, re-pulling services, or switching to backup dependent services.
[0176] The containerized service exception recovery instruction is a command in the service platform used to trigger the containerized service exception recovery operation. It can be generated by the service exception detection module when it detects an exception in the containerized service, and the containerized service exception recovery instruction is sent to the containerized service exception recovery module to notify the containerized service exception recovery module to perform exception recovery processing for the containerized service.
[0177] The service platform provided in the embodiments of this specification, through its containerized service anomaly recovery module and containerized service anomaly recovery instructions, enables targeted anomaly recovery processing for containerized services in the event of anomalies. This avoids the limitations of traditional methods that only recover for containers, ensuring that targeted recovery processing can be performed directly on containerized services in the event of anomalies. This improves the accuracy and efficiency of anomaly recovery, reduces the service interruption time provided by the service platform, reduces manual intervention costs, and improves the stability and service availability of the service platform.
[0178] See Figure 4 , Figure 4 A structural block diagram of a computing device 400 according to one embodiment of this specification is shown. The components of the computing device 400 include, but are not limited to, a memory 410 and a processor 420. The processor 420 is connected to the memory 410 via a bus 430, and a database 450 is used to store data.
[0179] The computing device 400 also includes an access device 440, which enables the computing device 400 to communicate via one or more networks 460. Examples of these networks include a Public Switched Telephone Network (PSTN), a Local Area Network (LAN), a Wide Area Network (WAN), a Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 440 may include one or more of any type of wired or wireless network interface (e.g., a network interface controller (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC) interface.
[0180] In one embodiment of this specification, the aforementioned components of the computing device 400 and Figure 4 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 4 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0181] The computing device 400 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 400 can also be a mobile or stationary server.
[0182] The processor 420 is used to execute the following computer program / instruction, which, when executed by the processor, implements the steps of the above-described exception handling method.
[0183] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the above-described exception handling method belong to the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the above-described exception handling method.
[0184] An embodiment of this specification also provides a computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the above-described exception handling method.
[0185] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the above-described exception handling method belong to the same concept. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the above-described exception handling method.
[0186] An embodiment of this specification also provides a computer program product, including a computer program / instruction that, when executed by a processor, implements the steps of the above-described exception handling method.
[0187] The above is an illustrative example of a computer program according to this embodiment. It should be noted that the technical solution of this computer program and the technical solution of the above-described exception handling method belong to the same concept. Details not described in detail in the technical solution of the computer program can be found in the description of the technical solution of the above-described exception handling method.
[0188] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0189] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.
[0190] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments in this specification are not limited to the described order of actions, because according to the embodiments in this specification, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments in this specification.
[0191] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0192] The preferred embodiments disclosed above are merely illustrative of this specification. Optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. An exception handling method applied to a service platform, the service platform including a container, the container being used to run containerized services, the method comprising: Perform service anomaly detection on the containerized service and obtain service detection results; If the containerized service is determined to be normal based on the service detection results, the container is started. If startup fails, startup exception information is obtained, and based on the startup exception information, exception recovery processing is performed on the container.
2. The method according to claim 1, wherein the service anomaly detection includes service dependency anomaly detection; The process of performing service anomaly detection on the containerized service and obtaining service detection results includes: The containerized service is subjected to service dependency anomaly detection, and the service detection result is determined based on the response status of other services that the containerized service depends on.
3. The method according to claim 1, after performing service anomaly detection on the containerized service and obtaining the service detection result, further comprising: If the containerized service is determined to be abnormal based on the service detection results, service abnormality information is obtained, and abnormal recovery processing is performed on the containerized service based on the service abnormality information.
4. The method according to claim 3, wherein the step of performing anomaly recovery processing on the containerized service based on the service anomaly information includes: Based on the service exception information, the service exception handling task corresponding to the service exception information is invoked to perform exception recovery processing on the containerized service.
5. The method according to claim 1, wherein the abnormal recovery processing of the container based on the startup abnormality information includes: Based on the startup exception information, the startup exception handling task corresponding to the startup exception information is invoked to perform exception recovery processing on the container.
6. The method according to claim 1, further comprising, after starting the container when the containerized service is determined to be normal based on the service detection result: If startup is successful, container operation checks will be performed based on the preset detection cycle to obtain the operation check results; If the container is determined to be malfunctioning based on the operational detection results, operational anomaly information is obtained, and based on the operational anomaly information, anomaly recovery processing is performed on the container.
7. The method according to claim 6, wherein the abnormal recovery processing of the container based on the operational abnormality information includes: Based on the aforementioned operational anomaly information, the corresponding operational anomaly handling task is invoked to perform anomaly recovery processing on the container.
8. The method according to claim 6, wherein performing container operation detection based on a preset detection period and obtaining operation detection results includes: Based on a preset detection cycle, container operation detection is performed to obtain multiple operation detection results; The step of obtaining operational anomaly information when the container is determined to be operationally abnormal based on the operational detection results includes: For the multiple operational detection results, count the number of abnormalities corresponding to the container operation anomalies; When the number of anomalies reaches a preset threshold, operational anomaly information is obtained.
9. A service platform, comprising a container, a service anomaly detection module, a container startup module, and a container anomaly recovery module; in, The container is used to run containerized services; The service anomaly detection module is used to respond to the service anomaly detection command, perform service anomaly detection on the containerized service, obtain the service detection result, and, if it is determined that the containerized service is normal based on the service detection result, generate a container startup command and send the container startup command to the container startup module. The container startup module is used to start the container in response to the container startup command, and generate a container exception recovery command in the event that the container startup fails, and send the container exception recovery command to the container exception recovery module. The container anomaly recovery module is used to respond to the container anomaly recovery command, obtain startup anomaly information, and perform anomaly recovery processing on the container based on the startup anomaly information.
10. The service platform according to claim 9 further includes a containerized service anomaly recovery module; The service anomaly detection module is further configured to respond to the service anomaly detection command, perform service anomaly detection on the containerized service, obtain service detection results, and, if the containerized service is determined to be abnormal based on the service detection results, generate a containerized service anomaly recovery command and send the containerized service anomaly recovery command to the containerized service anomaly recovery module. The containerized service exception recovery module is used to respond to the containerized service exception recovery instruction, obtain service exception information, and perform exception recovery processing on the containerized service based on the service exception information.
11. A computing device, comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the exception handling method according to any one of claims 1 to 8.
12. A computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the exception handling method according to any one of claims 1 to 8.
13. A computer program product comprising a computer program / instructions that, when executed by a processor, implement the steps of the exception handling method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Service orchestration and dependency relationship management method and system based on container cloud technology
CN110333932A
Container security protection method and system for biological information container cloud
CN117040797A
Abnormality processing method
CN119201551A
Intelligent service starting arrangement system based on container life cycle
CN121116486A