Container exception processing method and device, processor and electronic equipment

By receiving and comparing the heartbeat messages of containers, stuck containers can be identified and processed, solving the problem of difficulty in timely determining the status of containers in existing technologies and realizing timely handling of container anomalies.

CN116366508BActive Publication Date: 2026-05-19INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INDUSTRIAL AND COMMERCIAL BANK OF CHINA
Filing Date
2023-03-30
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing technologies make it difficult to determine the status of containers in a timely manner, resulting in the inability to handle container anomalies promptly.

Method used

By receiving heartbeat messages from multiple containers in the target node, the target set of containers sending heartbeat messages is determined and compared with a predetermined set of containers to identify sluggish containers that have not sent heartbeat messages. If no heartbeat message is received within a predetermined time period, sluggish information is sent to a predetermined terminal, carrying the identifier code of the sluggish container.

Benefits of technology

It enables timely identification and handling of stuck containers, solves the problem of difficulty in determining the status of containers in a timely manner, and improves the efficiency of container anomaly handling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116366508B_ABST
    Figure CN116366508B_ABST
Patent Text Reader

Abstract

The application discloses a container abnormality processing method and device, a processor and electronic equipment. It relates to the technical field of containers. The method comprises the following steps: receiving target heartbeat messages sent by multiple containers in a target node; determining a target container set that sends heartbeat messages according to the target heartbeat messages sent by the multiple containers in the target node; comparing the target container set with a predetermined container set to determine a stuck container in the target node that does not send heartbeat messages; and sending stuck information to a predetermined terminal under the condition that no heartbeat message sent by the stuck container is received within a predetermined time period, wherein the stuck information carries an identification code corresponding to the stuck container. The application solves the technical problem that the state of a container cannot be determined in time in the related art, so that the abnormal state cannot be determined in time for processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of containers, and more specifically, to a method, apparatus, processor, and electronic device for handling container anomalies. Background Technology

[0002] The container alarm technology in related technologies has the technical problem of being unable to determine the container status in a timely manner, thus failing to identify and handle abnormal states promptly.

[0003] There is currently no effective solution to the above problems. Summary of the Invention

[0004] This invention provides a container anomaly handling method, apparatus, processor, and electronic device to at least solve the technical problem in related technologies that it is difficult to determine the container status in a timely manner, thus making it impossible to determine and handle abnormal states in a timely manner.

[0005] According to one aspect of the present invention, a container anomaly handling method is provided, comprising: receiving target heartbeat messages sent by multiple containers in a target node respectively; determining a target container set that sent the heartbeat messages based on the target heartbeat messages sent by the multiple containers in the target node respectively; comparing the target container set with a predetermined container set to determine a sluggish container in the target node that has not sent a heartbeat message; and sending sluggish information to a predetermined terminal if no heartbeat message is received from the sluggish container within a predetermined time period, wherein the sluggish information carries an identification code corresponding to the sluggish container.

[0006] Optionally, before comparing the target container set with the predetermined container set to determine the sluggish container in the target node that has not sent a heartbeat message, the method further includes: receiving a registration request for adding a new container sent by a monitoring container in the target node, wherein the registration request for adding a new container is generated when the monitoring container detects a new container event in the target node, and the registration request for adding a new container carries an identifier code corresponding to the new container; in response to the registration request for adding a new container, updating the initial container set to obtain a predetermined container set including the new container.

[0007] Optionally, before comparing the target container set with the predetermined container set to determine the sluggish container in the target node that has not sent a heartbeat message, the method further includes: receiving a request to unregister a disappeared container sent by a monitoring container in the target node, wherein the request to unregister a disappeared container is generated when the monitoring container detects a disappeared container event in the target node, and the request to unregister a disappeared container carries an identifier code corresponding to the disappeared container; in response to the request to unregister a disappeared container, updating the initial container set to obtain a predetermined container set to remove the disappeared container.

[0008] Optionally, before comparing the target container set with the predetermined container set to determine the sluggish container in the target node that has not sent a heartbeat message, the method further includes: receiving a registration supplementary container request sent by a monitoring container in the target node, wherein the registration supplementary container request is generated by the monitoring container obtaining a database container set and a real-time container set stored in the database, comparing the database container set and the real-time container set, and if the database container set is a proper subset of the real-time container set, the registration supplementary container request carries an identifier code corresponding to the supplementary container, and the supplementary container is a container included in the real-time container set but not included in the database container set; in response to the registration supplementary container request, updating the initial container set to obtain a predetermined container set including the supplementary container.

[0009] Optionally, before comparing the target container set with the predetermined container set to determine the sluggish container in the target node that has not sent a heartbeat message, the method further includes: receiving a deregistration and release container request sent by the monitoring container in the target node, wherein the deregistration and release container request is generated by the monitoring container obtaining a database container set and a real-time container set stored in the database, comparing the database container set and the real-time container set, and if the real-time container set is a proper subset of the database container set, the deregistration and release container request carries an identifier code corresponding to the release container, and the release container is a container that is not included in the real-time container set but is included in the database container set; in response to the deregistration and release container request, updating the initial container set to obtain a predetermined container set to remove the release container.

[0010] Optionally, after determining the target container set for sending heartbeat messages based on the target heartbeat messages sent by multiple containers in the target node, the method further includes: determining the container status corresponding to each container in the target container set based on the target heartbeat messages sent by multiple containers in the target node.

[0011] Optionally, after sending the lag information to the predetermined terminal if no heartbeat message is received from the lag container within a predetermined time period, the method further includes: receiving a request to deregister the lag container from the predetermined terminal, wherein the request to deregister the lag container is generated when the predetermined terminal determines that the lag container is a faulty container, and the request to deregister the lag container carries an identification code corresponding to the lag container; and updating the predetermined container set in response to the request to deregister the lag container to obtain a container set in which the lag container is removed.

[0012] According to one aspect of the present invention, a container anomaly handling apparatus is provided, comprising: a receiving module, configured to receive target heartbeat messages sent by multiple containers in a target node respectively; a determining module, configured to determine a target container set that sent the heartbeat messages based on the target heartbeat messages sent by the multiple containers in the target node respectively; a comparison module, configured to compare the target container set with a predetermined container set to determine a stalled container in the target node that has not sent a heartbeat message; and a sending module, configured to send stall information to a predetermined terminal if no heartbeat message is received from the stalled container within a predetermined time period, wherein the stall information carries an identification code corresponding to the stalled container.

[0013] According to one aspect of the present invention, a processor is provided for running a program, wherein the program executes the container exception handling method described in any one of the preceding embodiments.

[0014] According to one aspect of the present invention, an electronic device is provided, including one or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the container exception handling method described in any of the preceding claims.

[0015] In this embodiment of the invention, by receiving target heartbeat messages sent by multiple containers in a target node, a target container set for sending heartbeat messages is determined based on these messages. The target container set is then compared with a predetermined container set to identify the sluggish containers in the target node that have not sent heartbeat messages. The sluggish containers are then monitored closely. If no heartbeat message is received from a sluggish container within a predetermined time period, sluggish information, including the corresponding identifier of the sluggish container, is sent to a predetermined terminal. This allows maintenance personnel using the predetermined terminal to promptly handle sluggish anomalies, check the status of sluggish containers, and address the sluggish containers in a timely manner. This solves the technical problem in related technologies where it is difficult to determine the container status in a timely manner, leading to the inability to promptly identify and handle abnormal states. Attached Figure Description

[0016] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:

[0017] Figure 1 This is a flowchart of a container exception handling method provided according to an embodiment of this application;

[0018] Figure 2 This is a structural block diagram of a container anomaly handling device provided according to an embodiment of this application;

[0019] Figure 3 This is a schematic diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0020] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0021] It should be noted that all information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data used for analysis, etc.) involved in this disclosure are information and data authorized by the user or fully authorized by all parties.

[0022] The present invention will now be described in conjunction with preferred implementation steps. Figure 1 This is a flowchart of a container exception handling method provided according to an embodiment of this application, such as... Figure 1 As shown.

[0023] First, it should be noted that the method provided in this application is applied to the following scenario: a cluster includes multiple nodes, and each node has one or more containers. Among these containers, at least one monitoring container is included to monitor the status of the node corresponding to that monitoring container. For example, it can monitor events such as adding or deleting containers. The cluster is connected to the Application Management Controller (AMC), and the status of containers on the nodes within the cluster is monitored using the method described below.

[0024] The method provided in the embodiments of this application is described below, and the method includes the following steps:

[0025] Step S101: Receive target heartbeat messages sent by multiple containers in the target node;

[0026] In step S101 provided in this application, the target node refers to a node in a cluster. A node includes multiple containers, and each container sends a target heartbeat message. The heartbeat message may contain basic information about the container, such as the container ID, IP address, resource utilization, etc. By receiving heartbeat messages sent by multiple containers, the operating status of the containers can be monitored in real time, so that timely processing and repair can be carried out in case of failures or anomalies.

[0027] Step S102: Determine the target container set for sending heartbeat messages based on the target heartbeat messages sent by multiple containers in the target node;

[0028] In step S102 provided in this application, based on the target heartbeat message, it is possible to determine which containers on the target node have sent heartbeat messages, and these containers are referred to as the target container set.

[0029] It should be noted that obtaining heartbeat messages can be achieved by accessing the HTTP interface within the container or executing a specific script. There are no restrictions here, and you can customize the settings according to the actual application and scenario, as long as you have the ability to obtain heartbeat messages.

[0030] Step S103: Compare the target container set with the predetermined container set to determine the stalled containers in the target node that have not sent heartbeat messages;

[0031] In step S103 provided in this application, the predetermined container set is the set that the Application Management Controller (AMC) has pre-stored. The predetermined container set represents all containers on the target node. Therefore, by comparing the target container set with the predetermined container set, it is possible to identify containers on the target node that have not sent heartbeat messages. These containers are called "stuck containers." This indicates that the container may be abnormal, and its state needs to be determined to ascertain whether it is in a stuck state or another state, so as to perform corresponding processing.

[0032] Step S104: If no heartbeat message is received from the lag container within the predetermined time period, lag information is sent to the predetermined terminal, wherein the lag information carries the identification code corresponding to the lag container.

[0033] In step S104 provided in this application, after identifying the lag container, the lag container is monitored closely. If no heartbeat message is received from the lag container within a predetermined time period, it indicates that the lag container is in an abnormal state and needs intervention. Therefore, lag information carrying the identification code corresponding to the lag container is sent to the predetermined terminal so that the predetermined terminal can find the lag container based on the identification code and process the lag container.

[0034] Through the above steps, by receiving target heartbeat messages sent by multiple containers in the target node, the target container set sending heartbeat messages is determined based on these messages. The target container set is then compared with a predetermined container set to identify the sluggish containers in the target node that have not sent heartbeat messages. These sluggish containers are then monitored closely. If no heartbeat message is received from a sluggish container within a predetermined time period, sluggish information, including the corresponding identifier of the sluggish container, is sent to a predetermined terminal. This allows maintenance personnel using the predetermined terminal to promptly handle sluggish anomalies, check the status of sluggish containers, and address the issues in a timely manner. This solves the technical problem in related technologies where it is difficult to determine the container status in a timely manner, thus hindering timely identification and handling of abnormal states.

[0035] As an optional embodiment, before comparing the target container set with the predetermined container set and determining the sluggish container in the target node that has not sent a heartbeat message, the method further includes: receiving a registration request for adding a new container sent by the monitoring container in the target node, wherein the registration request for adding a new container is generated when the monitoring container detects a new container event in the target node, and the registration request for adding a new container carries an identifier code corresponding to the new container; in response to the registration request for adding a new container, updating the initial container set to obtain a predetermined container set including the new container.

[0036] In this embodiment, when the monitoring container detects a new container event in the target node, the monitoring container sends a registration request for the new container to the AMC. The AMC receives the registration request for the new container sent by the monitoring container, updates the initial container set according to the identifier code corresponding to the new container carried in the registration request for the new container, adds the new container to the initial container set, and obtains a predetermined container set including the new container. This ensures that the predetermined container set is updated in real time, and accurate results of the stuck container can be obtained after comparison.

[0037] As an optional embodiment, before comparing the target container set with the predetermined container set and determining the sluggish container in the target node that has not sent a heartbeat message, the method further includes: receiving a request to unregister a disappeared container sent by a monitoring container in the target node, wherein the request to unregister a disappeared container is generated when the monitoring container detects a disappeared container event in the target node, and the request to unregister a disappeared container carries an identifier code corresponding to the disappeared container; in response to the request to unregister a disappeared container, updating the initial container set to obtain the predetermined container set to remove the disappeared container.

[0038] In this embodiment, when the monitoring container detects a missing container event in the target node, the monitoring container sends an unregistered container request to the AMC. The AMC receives the unregistered missing container request from the monitoring container, updates the initial container set according to the identifier code corresponding to the missing container carried in the unregistered missing container request, deletes the missing container from the initial container set, and obtains a predetermined container set with the missing container removed. This ensures that the predetermined container set is updated in real time, and accurate results of the stuck container can be obtained after comparison.

[0039] As an optional embodiment, before comparing the target container set with the predetermined container set to determine the sluggish container in the target node that has not sent a heartbeat message, the method further includes: receiving a registration supplementary container request sent by the monitoring container in the target node, wherein the registration supplementary container request is generated by the monitoring container obtaining the database container set and the real-time container set stored in the database, comparing the database container set and the real-time container set, and if the database container set is a proper subset of the real-time container set, the registration supplementary container request carries an identifier code corresponding to the supplementary container, and the supplementary container is a container included in the real-time container set but not included in the database container set; in response to the registration supplementary container request, updating the initial container set to obtain a predetermined container set including the supplementary container.

[0040] In this embodiment, the monitoring container obtains a database container set and a real-time container set stored in the database. It should be noted that the containers in the database container set are consistent with the containers in the predetermined container set. Therefore, by comparing the database container set and the real-time container set, it can be determined whether the predetermined container set is incorrect. The real-time container set can be obtained by calling the corresponding interface. If the database container set is a proper subset of the real-time container set, it means that no actually existing containers are recorded in the AMC. Therefore, the monitoring container sends a registration request for supplementary containers to the AMC. The AMC receives the registration request, updates the initial container set according to the identifier code corresponding to the supplementary container carried in the registration request, and adds the supplementary container to the initial container set, resulting in a predetermined container set including the supplementary container. This ensures that the predetermined container set is calibrated, and the comparison can yield accurate results for the stuck containers.

[0041] As an optional embodiment, before comparing the target container set with the predetermined container set to determine the sluggish container in the target node that has not sent a heartbeat message, the method further includes: receiving a deregistration and release container request sent by the monitoring container in the target node, wherein the deregistration and release container request is generated by the monitoring container obtaining the database container set and the real-time container set stored in the database, comparing the database container set and the real-time container set, and finding that the real-time container set is a proper subset of the database container set; the deregistration and release container request carries an identifier code corresponding to the release container; and the release container is a container that is not included in the real-time container set but is included in the database container set; in response to the deregistration and release container request, the initial container set is updated to obtain the predetermined container set for removing the release container.

[0042] In this embodiment, the monitoring container obtains the database container set and the real-time container set stored in the database. It should be noted that the containers in the database container set are consistent with the predetermined container set. Therefore, by comparing the database container set and the real-time container set, it can be determined whether the predetermined container set is incorrect. The real-time container set can be obtained by calling the corresponding interface. If the real-time container set is a proper subset of the database container set, it indicates that the AMC records containers that do not actually exist. The monitoring container sends an unregistered container release request to the AMC. The AMC receives the unregistered container release request from the monitoring container, updates the initial container set according to the identifier code corresponding to the released container carried in the unregistered container release request, and deletes the released container from the initial container set, obtaining a predetermined container set with the released container removed. This ensures that the predetermined container set is corrected, and the comparison can yield accurate results for the stuck containers.

[0043] As an optional embodiment, after determining the target container set for sending heartbeat messages based on the target heartbeat messages sent by multiple containers in the target node, the method further includes: determining the container status corresponding to each container in the target container set based on the target heartbeat messages sent by multiple containers in the target node.

[0044] In this embodiment, the container status of a container can also be determined by the target heartbeat messages sent by multiple containers respectively. Because the target heartbeat message includes a variety of information, the container status of each container can be determined by the series of information included in the target heartbeat message.

[0045] As an optional embodiment, if no heartbeat message is received from the lag container within a predetermined time period, after sending lag information to the predetermined terminal, the method further includes: receiving a request to deregister the lag container sent by the predetermined terminal, wherein the request to deregister the lag container is generated when the predetermined terminal determines that the lag container is a faulty container, and the request to deregister the lag container carries an identification code corresponding to the lag container; in response to the request to deregister the lag container, updating the predetermined container set to obtain a container set in which the lag container is removed.

[0046] In this embodiment, after sending the lag information to the predetermined terminal, feedback is also received from the predetermined terminal. If the predetermined terminal determines that the lag container is faulty, it determines that this container is a faulty container and sends a deregistration request to the AMC. The AMC receives the deregistration request, updates the predetermined container set, and obtains a container set with the lag container removed. This container set is then used for comparison operations when the method of this application is executed again.

[0047] Based on the above embodiments and optional embodiments, an optional implementation method is provided, which is described in detail below.

[0048] An optional embodiment of the present invention provides a container failure handling method. The method provided by the present invention not only monitors the startup and death status of the container, but also tracks the heartbeat messages of the container. Generally, the heartbeat messages are sent from the deep health check inside the container, which can effectively detect whether the container is in a normal state. Moreover, the method provided by the present invention can avoid the situation where the container is alive but in an abnormal state by receiving the heartbeat messages sent by the container. The optional embodiments of the present invention are described in detail below.

[0049] When a new container is added to a target node in the cluster, a new container event, also known as a container startup event, is triggered. After the event monitor (the same monitoring node mentioned above) detects the container startup event, it sends a registration request to the AMC, which can be specifically called a new container registration request. The new container registration request contains the new container ID (the same identifier code mentioned above), which is the container's unique identifier. After this, the AMC will begin monitoring the health status of this container.

[0050] Containers monitored by the AMC in the cluster will continuously send heartbeat messages to the AMC. If the AMC does not receive a heartbeat message within a predetermined time period, such as ten minutes, it will send a stall message to a predetermined terminal. For example, it can send a disconnection alarm to the operations and maintenance personnel using the predetermined terminal, so that the operations and maintenance personnel can intervene in time and handle the container conveniently and quickly.

[0051] When a container on a target node in the cluster dies (or disappears), the cluster emits a container death event. After the event monitor detects this event, it issues an unregistration request, specifically an unregistered disappearing container request. Afterward, the AMC will no longer monitor this container. It's important to note that after detecting a container death event, it determines whether the death was caused by cluster-wide scheduling, i.e., whether it was due to manual operation or system-defined procedures. Manual operation could include operations personnel upgrading or taking containers offline. If the death was not caused by cluster-wide scheduling, a container restart operation can be performed, and a container restart alert can be sent to a designated terminal, allowing operations personnel using that terminal to intervene. Because there are many possible causes for deaths not caused by cluster-wide scheduling, such as internal application errors (bugs or exceptions in the application code causing the container process to crash or stop running), or resource limitations (the container's resource requirements exceeding allocated limits causing forced termination by the system), each case requires specific analysis and operations personnel intervention.

[0052] Eventmonitor can also interface with the persistent file storage system HDFS. After a container restarts abnormally, it will upload the container process logs stored on the target node to HDFS for easy download and analysis by operations and maintenance personnel. Without downloading, the system will periodically clean up container logs, which can be inconvenient during problem analysis.

[0053] Eventmonitor also records the container's historical state, startup time, death time, assigned node, cluster, and cause of death in the database, making it easier for operations and maintenance personnel to analyze container restart records and summarize the reasons for restarts.

[0054] Eventmonitor compares the containers already recorded in the database (the same set of database containers mentioned above) with the containers that exist in real-time on the host machine (the same set of real-time containers mentioned above). If a container that exists in real-time on the host machine does not exist in the database, it will be added to the database and a request to register and add the container will be sent to the AMC. If a container marked as alive in the database is not on the current host machine, the status of this container in the database will be marked as dead, and a request to unregister and release the container will be sent to the AMC.

[0055] The above optional implementation methods can achieve at least the following beneficial effects:

[0056] (1) It solves the problem in related technologies that it is impossible to promptly alert by simply querying the container status in scenarios such as containers being alive but processes being stuck;

[0057] (2) This solves the problem that related technologies cannot record container history information and exception logs, which makes it inconvenient to troubleshoot problems;

[0058] (3) No special rules need to be configured; the container automatically connects to the alarm mechanism.

[0059] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0060] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of the present invention.

[0061] According to embodiments of the present invention, an apparatus for implementing the above-described container anomaly handling method is also provided. Figure 2 This is a structural block diagram of a container fault handling device provided according to an embodiment of this application, such as... Figure 2 As shown, the device includes: a receiving module 201, a determining module 202, a comparing module 203, and a sending module 204. The device will be described in detail below.

[0062] The receiving module 201 is used to receive target heartbeat messages sent by multiple containers in the target node; the determining module 202, connected to the receiving module 201, is used to determine the target container set for sending heartbeat messages based on the target heartbeat messages sent by multiple containers in the target node; the comparing module 203, connected to the determining module 202, is used to compare the target container set with a predetermined container set to determine the stalled container in the target node that has not sent a heartbeat message; the sending module 204, connected to the comparing module 203, is used to send stall information to a predetermined terminal if no heartbeat message is received from the stalled container within a predetermined time period, wherein the stall information carries an identification code corresponding to the stalled container.

[0063] It should be noted that the receiving module 201, determining module 202, comparing module 203 and sending module 204 mentioned above correspond to steps S101 to S106 in the container exception handling method. The instances and application scenarios implemented by multiple modules and corresponding steps are the same, but are not limited to the content disclosed in the above embodiments.

[0064] The container anomaly handling device provided in this application embodiment receives target heartbeat messages sent by multiple containers in a target node, determines the target container set that sent the heartbeat messages based on the target heartbeat messages sent by multiple containers in the target node, compares the target container set with a predetermined container set, identifies the sluggish containers in the target node that have not sent heartbeat messages, and then focuses on monitoring the sluggish containers. If no heartbeat message is received from the sluggish containers within a predetermined time period, sluggish information including the corresponding identifier code of the sluggish containers is sent to a predetermined terminal. This allows maintenance personnel using the predetermined terminal to handle sluggish anomalies in a timely manner, check the status of sluggish containers, and handle sluggish containers promptly. This solves the technical problem in related technologies where it is difficult to determine the container status in a timely manner, thus making it impossible to determine and handle abnormal states in a timely manner.

[0065] The container exception handling device includes a processor and a memory. The aforementioned modules are all stored in the memory as program units, and the processor executes the aforementioned program units stored in the memory to implement the corresponding functions.

[0066] The processor contains a kernel, which retrieves the corresponding program units from memory. One or more kernels can be configured; by adjusting kernel parameters, the technical problem in related technologies of difficulty in timely determining the container state, thus hindering timely handling of abnormal states, can be solved.

[0067] The memory may include non-permanent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.

[0068] This invention provides a computer-readable storage medium storing a program that, when executed by a processor, implements a container exception handling method.

[0069] This invention provides a processor for running a program, wherein the program executes a container exception handling method during runtime.

[0070] Figure 3 This is a schematic diagram of an electronic device provided in an embodiment of the present invention, such as... Figure 3As shown, this embodiment of the invention provides an electronic device, which includes a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, it performs the following steps: receiving target heartbeat messages sent by multiple containers in a target node; determining a target container set that sent the heartbeat messages based on the target heartbeat messages sent by the multiple containers in the target node; comparing the target container set with a predetermined container set to determine the sluggish container in the target node that did not send a heartbeat message; and sending sluggish information to a predetermined terminal if no heartbeat message is received from the sluggish container within a predetermined time period, wherein the sluggish information carries an identification code corresponding to the sluggish container.

[0071] Optionally, before comparing the target container set with the predetermined container set to determine the stalled container in the target node that has not sent a heartbeat message, the method further includes: receiving a registration request for adding a new container sent by the monitoring container in the target node, wherein the registration request for adding a new container is generated when the monitoring container detects a new container event in the target node, and the registration request for adding a new container carries the identification code corresponding to the new container; in response to the registration request for adding a new container, updating the initial container set to obtain a predetermined container set including the new container.

[0072] Optionally, before comparing the target container set with the predetermined container set to determine the sluggish container in the target node that has not sent a heartbeat message, the method further includes: receiving a request to unregister a disappeared container sent by a monitoring container in the target node, wherein the request to unregister a disappeared container is generated when the monitoring container detects a disappeared container event in the target node, and the request to unregister a disappeared container carries an identifier code corresponding to the disappeared container; in response to the request to unregister a disappeared container, updating the initial container set to obtain the predetermined container set to remove the disappeared container.

[0073] Optionally, before comparing the target container set with the predetermined container set to determine the sluggish container in the target node that has not sent a heartbeat message, the method further includes: receiving a registration supplementary container request sent by the monitoring container in the target node, wherein the registration supplementary container request is generated by the monitoring container obtaining the database container set and the real-time container set stored in the database, comparing the database container set and the real-time container set, and if the database container set is a proper subset of the real-time container set, the registration supplementary container request carries an identifier code corresponding to the supplementary container, and the supplementary container is a container included in the real-time container set but not included in the database container set; in response to the registration supplementary container request, updating the initial container set to obtain a predetermined container set including the supplementary container.

[0074] Optionally, before comparing the target container set with the predetermined container set to determine the sluggish container in the target node that has not sent a heartbeat message, the method further includes: receiving a deregistration and release container request sent by the monitoring container in the target node, wherein the deregistration and release container request is generated by the monitoring container obtaining the database container set and the real-time container set stored in the database, comparing the database container set and the real-time container set, and if the real-time container set is a proper subset of the database container set, the deregistration and release container request carries an identifier code corresponding to the release container, and the release container is a container that is not included in the real-time container set but is included in the database container set; in response to the deregistration and release container request, updating the initial container set to obtain the predetermined container set for removing the release container.

[0075] Optionally, after determining the target container set for sending heartbeat messages based on the target heartbeat messages sent by multiple containers in the target node, the method further includes: determining the container status corresponding to each container in the target container set based on the target heartbeat messages sent by multiple containers in the target node.

[0076] Optionally, if no heartbeat message is received from the lag container within the predetermined time period, after sending the lag information to the predetermined terminal, the method further includes: receiving a request to deregister the lag container from the predetermined terminal, wherein the request to deregister the lag container is generated when the predetermined terminal determines that the lag container is a faulty container, and the request to deregister the lag container carries the identification code corresponding to the lag container; in response to the request to deregister the lag container, updating the predetermined container set to obtain a container set in which the lag container is removed.

[0077] The devices mentioned in this article can be servers, PCs, tablets, mobile phones, etc.

[0078] This application also provides a computer program product, which, when executed on a data processing device, is suitable for executing an initialization program with the following method steps: receiving target heartbeat messages sent by multiple containers in a target node; determining a target container set for sending heartbeat messages based on the target heartbeat messages sent by the multiple containers in the target node; comparing the target container set with a predetermined container set to determine the stalled containers in the target node that have not sent heartbeat messages; and, if no heartbeat message is received from the stalled containers within a predetermined time period, sending stall information to a predetermined terminal, wherein the stall information carries an identification code corresponding to the stalled container.

[0079] Optionally, before comparing the target container set with the predetermined container set to determine the stalled container in the target node that has not sent a heartbeat message, the method further includes: receiving a registration request for adding a new container sent by the monitoring container in the target node, wherein the registration request for adding a new container is generated when the monitoring container detects a new container event in the target node, and the registration request for adding a new container carries the identification code corresponding to the new container; in response to the registration request for adding a new container, updating the initial container set to obtain a predetermined container set including the new container.

[0080] Optionally, before comparing the target container set with the predetermined container set to determine the sluggish container in the target node that has not sent a heartbeat message, the method further includes: receiving a request to unregister a disappeared container sent by a monitoring container in the target node, wherein the request to unregister a disappeared container is generated when the monitoring container detects a disappeared container event in the target node, and the request to unregister a disappeared container carries an identifier code corresponding to the disappeared container; in response to the request to unregister a disappeared container, updating the initial container set to obtain the predetermined container set to remove the disappeared container.

[0081] Optionally, before comparing the target container set with the predetermined container set to determine the sluggish container in the target node that has not sent a heartbeat message, the method further includes: receiving a registration supplementary container request sent by the monitoring container in the target node, wherein the registration supplementary container request is generated by the monitoring container obtaining the database container set and the real-time container set stored in the database, comparing the database container set and the real-time container set, and if the database container set is a proper subset of the real-time container set, the registration supplementary container request carries an identifier code corresponding to the supplementary container, and the supplementary container is a container included in the real-time container set but not included in the database container set; in response to the registration supplementary container request, updating the initial container set to obtain a predetermined container set including the supplementary container.

[0082] Optionally, before comparing the target container set with the predetermined container set to determine the sluggish container in the target node that has not sent a heartbeat message, the method further includes: receiving a deregistration and release container request sent by the monitoring container in the target node, wherein the deregistration and release container request is generated by the monitoring container obtaining the database container set and the real-time container set stored in the database, comparing the database container set and the real-time container set, and if the real-time container set is a proper subset of the database container set, the deregistration and release container request carries an identifier code corresponding to the release container, and the release container is a container that is not included in the real-time container set but is included in the database container set; in response to the deregistration and release container request, updating the initial container set to obtain the predetermined container set for removing the release container.

[0083] Optionally, after determining the target container set for sending heartbeat messages based on the target heartbeat messages sent by multiple containers in the target node, the method further includes: determining the container status corresponding to each container in the target container set based on the target heartbeat messages sent by multiple containers in the target node.

[0084] Optionally, if no heartbeat message is received from the lag container within the predetermined time period, after sending the lag information to the predetermined terminal, the method further includes: receiving a request to deregister the lag container from the predetermined terminal, wherein the request to deregister the lag container is generated when the predetermined terminal determines that the lag container is a faulty container, and the request to deregister the lag container carries the identification code corresponding to the lag container; in response to the request to deregister the lag container, updating the predetermined container set to obtain a container set in which the lag container is removed.

[0085] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0086] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0087] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0088] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0089] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0090] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0091] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0092] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0093] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0094] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A method for handling container anomalies, characterized in that, include: Receive target heartbeat messages sent by multiple containers in the target node; Based on the target heartbeat messages sent by multiple containers in the target node, determine the target container set for sending heartbeat messages; By comparing the target container set with the predetermined container set, the sluggish containers in the target node that have not sent heartbeat messages are identified. If no heartbeat message is received from the lag container within a predetermined time period, lag information is sent to a predetermined terminal, wherein the lag information carries the identification code corresponding to the lag container; Before comparing the target container set with the predetermined container set to determine the sluggish containers in the target node that have not sent heartbeat messages, the process further includes: updating the initial container set based on the registration request for adding a new container sent by the monitoring container in the target node, the unregistration request for disappearing a container sent by the monitoring container in the target node, the registration request for supplementing a container sent by the monitoring container in the target node, and the unregistration request for releasing a container sent by the monitoring container in the target node, to obtain the predetermined container set. The registration request for supplementing a container is a container included in the real-time container set but not included in the database container set, and the unregistration request for releasing a container is a container not included in the real-time container set but included in the database container set.

2. The method according to claim 1, characterized in that, Before comparing the target container set with the predetermined container set to determine the sluggish container in the target node that has not sent a heartbeat message, the method further includes: Receive a registration request for a new container sent by the monitoring container in the target node, wherein the registration request for a new container is generated when the monitoring container detects a new container event in the target node, and the registration request for a new container carries an identifier code corresponding to the new container; In response to the registration request for a new container, the initial container set is updated to obtain a predetermined container set that includes the new container.

3. The method according to claim 1, characterized in that, Before comparing the target container set with the predetermined container set to determine the sluggish container in the target node that has not sent a heartbeat message, the method further includes: Receive a request to unregister a vanished container sent by the monitoring container in the target node, wherein the request to unregister a vanished container is generated when the monitoring container detects a vanished container event in the target node, and the request to unregister a vanished container carries an identifier code corresponding to the vanished container; In response to the request to unregister the disappeared container, the initial container set is updated to obtain a predetermined container set from which the disappeared container is removed.

4. The method according to claim 1, characterized in that, Before comparing the target container set with the predetermined container set to determine the sluggish container in the target node that has not sent a heartbeat message, the method further includes: The system receives a registration request for a supplementary container sent by the monitoring container in the target node. The registration request for a supplementary container is generated by the monitoring container obtaining a database container set and a real-time container set stored in the database, comparing the database container set and the real-time container set, and if the database container set is a proper subset of the real-time container set. The registration request for a supplementary container carries an identifier code corresponding to the supplementary container. The supplementary container is a container included in the real-time container set but not included in the database container set. In response to the registration supplementary container request, the initial container set is updated to obtain a predetermined container set including the supplementary container.

5. The method according to claim 1, characterized in that, Before comparing the target container set with the predetermined container set to determine the sluggish container in the target node that has not sent a heartbeat message, the method further includes: The system receives a deregistration and release container request sent by the monitoring container in the target node. The deregistration and release container request is generated by the monitoring container obtaining a database container set and a real-time container set stored in the database, comparing the database container set and the real-time container set, and if the real-time container set is a proper subset of the database container set. The deregistration and release container request carries an identifier code corresponding to the release container. The release container is a container that is not included in the real-time container set but is included in the database container set. In response to the unregistered release container request, the initial container set is updated to obtain a predetermined container set from which the release container is removed.

6. The method according to claim 1, characterized in that, After determining the target container set for sending heartbeat messages based on the target heartbeat messages sent by multiple containers in the target node, the process further includes: Based on the target heartbeat messages sent by multiple containers in the target node, the container status corresponding to each container in the target container set is determined.

7. The method according to any one of claims 1 to 6, characterized in that, After sending the lag information to the predetermined terminal when no heartbeat message is received from the lag container within the predetermined time period, the process further includes: The system receives a request to deregister a lag container sent by the predetermined terminal, wherein the request to deregister a lag container is generated when the predetermined terminal determines that the lag container is a faulty container, and the request to deregister a lag container carries an identification code corresponding to the lag container. In response to the request to unregister the lag container, the predetermined container set is updated to obtain the container set to remove the lag container.

8. A container anomaly handling device, characterized in that, include: The receiving module is used to receive target heartbeat messages sent by multiple containers in the target node. The determination module is used to determine the target container set that sent the heartbeat message based on the target heartbeat messages sent by multiple containers in the target node. The comparison module is used to compare the target container set with the predetermined container set to determine the sluggish containers in the target node that have not sent heartbeat messages; The sending module is used to send lag information to a predetermined terminal if no heartbeat message is received from the lag container within a predetermined time period. The lag information carries an identification code corresponding to the lag container. The device is further configured to, before comparing the target container set with the predetermined container set and determining the sluggish container in the target node that has not sent a heartbeat message, update the initial container set based on the registration request for adding a new container sent by the monitoring container in the target node, the deregistration request for disappearing a container sent by the monitoring container in the target node, the registration request for supplementing a container sent by the monitoring container in the target node, and the deregistration request for releasing a container sent by the monitoring container in the target node, to obtain the predetermined container set. The registration request for supplementing a container is a container included in the real-time container set but not included in the database container set, and the deregistration request for releasing a container is a container not included in the real-time container set but included in the database container set.

9. A processor, characterized in that, The processor is used to run a program, wherein the program executes the container exception handling method according to any one of claims 1 to 7 when it runs.

10. An electronic device, characterized in that, It includes one or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the container exception handling method according to any one of claims 1 to 7.