Virtualization System Failure Isolation Device and Virtualization System Failure Isolation Method
The virtualization system failure isolation device addresses the challenge of delayed recovery in virtualization systems by using an abnormality detection and response unit to quickly identify and isolate failing containers, achieving faster recovery than traditional methods.
Patent Information
- Application Number
- JP2023531191
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-06-29
- Publication Date
- 2025-06-18
- Estimated Expiration
- 2041-06-29
AI Technical Summary
In virtualization systems, manual recovery operations are typically required for failures, leading to delays in normalizing system operations, and existing automated recovery methods, such as those using Kubernetes' probe functions, are limited in their ability to rapidly recover from failures.
A virtualization system failure isolation device that includes a computing resource cluster, a cluster management unit, an abnormality detection unit, and an abnormality response unit. This device detects abnormalities outside the virtualization configuration and transmits commands to stop abnormal containers, allowing for earlier recovery than traditional methods.
Enables faster failure recovery in virtualization systems by allowing for immediate detection and isolation of abnormal containers, independent of predetermined monitoring cycles, thereby reducing the time to normalize system operations.
Smart Images

Figure 0007694659000001 
Figure 0007694659000002 
Figure 0007694659000003
Abstract
Description
Technical Field
[0001] The present invention relates to a virtualization system failure isolation device and a virtualization system failure isolation method for realizing abnormal detection and failure recovery of containers and applications operating on containers in a computing infrastructure based on virtual machines and containers.
Background Art
[0002] The above-described virtual machine is a computer that realizes the same functions as a physical computer by software. A container is a virtualization technology created by packaging an application in an environment called a "container" and operating on a container engine. In conventional container-based technologies, abnormal detection and failure recovery of containers and applications operating on containers are mainly realized by the Liveness / Readiness Probe function (also referred to as the probe function) of Kubernetes, which will be described later.
[0003] Kubernetes is container virtualization software that creates and clusters containers such as Docker and is open-source software. The Liveness Probe function controls operations such as restarting a container, and the Readiness Probe function controls whether a container accepts requests or not. There is a technique described in Non-Patent Document 1 as this type of prior art.
Prior Art Documents
Non-Patent Documents
[0004]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] By the way, not only in the above-described containers but also in a virtualization system as a virtualization technology area, for a failure in the virtualization system, a manual recovery operation or the like is performed based on the reported alert. However, since the recovery operation is performed manually after the alert is reported, it is difficult to shorten the time from the occurrence of the failure to normalization.
[0006] When Kubernetes recovers the failure using its probe function for failure recovery, the monitoring period for the failure can only be set to a slow period such as 1 second determined in advance. For this reason, when it is necessary to recover as quickly as possible, there is a problem that it is impossible to recover faster than the recovery by the default failure recovery function of Kubernetes.
[0007] The present invention has been made in view of such circumstances, and an object thereof is to recover a failure that has occurred in a virtualization system earlier than the recovery by the failure recovery function of container virtualization software.
Means for Solving the Problems
[0008] To solve the above problems, the virtualization system failure isolation device of the present invention includes a computing resource cluster that is virtually created by container virtualization software on a physical machine and clusters and arranges the virtually created containers, a cluster management unit that manages the control related to the arrangement and operation of the virtually created and clustered containers, an abnormality detection unit that is created outside the virtually created computing resource cluster and the cluster management unit and detects abnormalities of the containers, and an abnormality response unit that transmits a command for instructing to stop the abnormal containers detected by the abnormality detection unit to the cluster management unit. Obtain service information related to communication from a counter device that communicates via a network, and include an endpoint setting unit shared by a plurality of containers in the computing resource cluster. In a configuration where the cluster management unit accesses requests to a plurality of containers via the endpoint setting unit, a label with either an accessible label value that enables access to requests from the cluster management unit or an inaccessible label value that disables access is assigned to the clustered containers. The abnormality handling unit transmits a command to the cluster management unit that instructs to set the label value of the abnormal container detected by the abnormality detection unit to inaccessible. The cluster management unit changes the label value of the abnormal container instructed by the command to inaccessible, thereby It is characterized by stopping the abnormal containers.
Effects of the Invention
[0009] According to the present invention, a failure occurring in a virtualization system can be recovered earlier than recovery by a failure recovery function of container virtualization software.
Brief Description of the Drawings
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Embodiments for Carrying Out the Invention
[0011] Hereinafter, embodiments of the present invention will be described with reference to the drawings. However, in all the drawings of this specification, components having corresponding functions are denoted by the same reference numerals, and the description thereof will be omitted as appropriate. <Configuration of the Embodiment> FIG. 1 is a block diagram showing the configuration of a virtualization system failure isolation device according to an embodiment of the present invention.
[0012] The virtualization system failure isolation device (also referred to as the failure isolation device) 10 shown in FIG. 1 stops or deletes and isolates the containers in which a failure has occurred in the container system 20 described later. This failure isolation device 10 includes a cluster management unit 14, a computing resource cluster 15, an abnormality detection unit 17, and an abnormality handling unit 18. The cluster 12 is configured by the cluster management unit 14 and the computing resource cluster 15. The abnormality detection unit 17 and the abnormality handling unit 18 are provided outside this cluster 12.
[0013] The computing resource cluster 15 is configured to include a plurality of applications 15a, 15b. The applications 15a, 15b are, in other words, Pods as management units of an aggregate of one or more containers. A Pod is the smallest unit of an application that can be executed by Kubernetes (container virtualization software). That is, containers are created and clustered by the applications 15a, 15b as Pods, and this cluster is operated on a container engine. This computing resource cluster 15 is virtually created by container virtualization software on a physical machine, and clusters and arranges the virtually created containers.
[0014] The container system 20 is a virtualization system composed of one or more clusters 12. When there are two clusters 12, each cluster 12 is configured to include a cluster management unit 14 and a computing resource cluster 15.
[0015] The cluster management unit 14 manages the control related to the arrangement and operation of the virtually created and clustered containers. This cluster management unit 14 includes a communication distribution unit 14a, a computing resource operation unit 14b, a computing resource management unit 14c, a container configuration reception unit 14d, a container placement destination determination unit 14e, and a container management unit 14f.
[0016] In the failure isolation device 10 configured as described above, the abnormality detection unit 17 detects an abnormality in one or more containers, i.e., Pods (applications) 15a and 15b, within the container system 20. The abnormality handling unit 18 stops the Pod (e.g., Pod 15a) in which an abnormality has been detected by the abnormality detection unit 17, and notifies the communication distribution unit 14a of a command as request information for operating only the normal Pod 15b.
[0017] The communication distribution unit 14a is a router, and distributes and notifies a command from the abnormality handling unit 18 or a request corresponding to the command to the corresponding units 14b to 14f.
[0018] The container configuration reception unit (also referred to as the reception unit) 14d receives configuration information for deploying (placing) a container in the computing resource cluster 15 from an external server or the like.
[0019] The container placement destination determination unit (also referred to as the placement destination determination unit) 14e determines, based on the configuration information received by the reception unit 14d, which container is to be placed on which worker node (computing resource cluster 15).
[0020] The container management unit 14f checks whether a container is operating normally or not.
[0021] The computing resource management unit 14c grasps and manages whether a worker node is operable, the usage amount of the computing resources of the server constituting the worker node, the remaining amount of the CPU (Central Processing Unit), and the like.
[0022] The computing resource operation unit 14b performs an operation of allocating a certain amount of computing resources such as the CPU to a certain container, in other words, an operation of allocating a storage capacity, CPU time, the memory capacity available to the container, and the like.
[0023] Next, various abnormality detection processes (first to sixth abnormality detection processes) related to the containers of the container system 20 by the abnormality detection unit 17 of the failure isolation device 10 will be described with reference to FIGS. 2 to 7.
[0024] <First Abnormality Detection Process> FIG. 2 is a block diagram for explaining the first abnormality detection process of containers by Pods (applications) 15a and 15b of the virtualization system failure isolation device 10 according to the present embodiment. However, each of the Pods 15a and 15b constitutes one or more containers.
[0025] In FIG. 2, within the container system 20, a master node 14J, an infrastructure node 14K, and worker nodes 15J and 15K are configured by virtual machines, and each is connected by a virtual switch {OVS (Open vSwitch)} 30. The master node 14J and the infrastructure node 14K correspond to the cluster management unit 14 (FIG. 1), and the worker nodes 15J and 15K correspond to the computing resource cluster 15 (FIG. 1).
[0026] Furthermore, a first cluster 12 is configured by the master node 14J and the worker node 15J, and a second cluster 12 is configured by the infrastructure node 14K and the worker node 15K. It is assumed that the container system 20 is configured by these clusters 12.
[0027] Outside the container system 20, an abnormality detection unit 17 is arranged in the same configuration as in FIG. 1. In FIG. 2, a total of two abnormality detection units 17 are shown for each of the worker nodes 15J and 15K, but one may also be sufficient. The master node 14J, the infrastructure node 14K, the worker nodes 15J and 15K, and the abnormality detection unit 17 are connected to a counterpart device 24 by a network 22. The counterpart device 24 is a communication device such as an external server that transmits a request signal or the like to the container system 20.
[0028] The abnormality detection unit 17 sends a predetermined command (for example, "sudo crictl ps") to the Pods 15a and 15b of the worker nodes 15J and 15K by polling indicated by the reciprocating arrows Y1 and Y2, and determines whether it is normal or abnormal based on the response results returned from the Pods 15a and 15b according to the command. In this polling test, the average value of the round-trip time when polling was executed 10 times was 0.06 seconds.
[0029] The abnormality determination in the abnormality detection unit 17 is performed by reading a character string indicating normal or abnormal described in the command response results returned from the Pods 15a and 15b by polling. For example, the character string "Running" indicates that the operation of the container (Pods 15a and 15b) is normal, and a character string other than "Running" indicates abnormality. Therefore, when "Running" is described in the command response results, the abnormality detection unit 17 determines that the operation of the container (Pods 15a and 15b) is normal, and when a character string other than "Running" is described, it determines that it is abnormal.
[0030] <Second Abnormality Detection Process> Next, FIG. 3 is a block diagram for explaining a second abnormality detection process by the routing table 15c provided for each of the worker nodes 15J and 15K of the virtualization system failure isolation device 10 of the present embodiment.
[0031] The routing table (also referred to as a table) 15c manages the destination container of the packet transmitted to the Pods 15a and 15b of the worker nodes 15J and 15K via the network 22 from the opposing device 24 with route information indicating the destination. If the destination management of this table 15c is incorrect, the packet will not reach the appropriate container. Therefore, the abnormality detection unit 17 is configured to detect the normality or abnormality of the destination management of the table 15c.
[0032] However, the routing table 15c is composed of a pair of tables of "iptables" and "nftables".
[0033] The abnormality detection unit 17 sends a predetermined command to each table 15c of the worker nodes 15J and 15K by polling indicated by the reciprocating arrows Y3 and Y4, and determines whether it is normal or abnormal based on the response results returned from each table 15c according to the command.
[0034] The above-mentioned predetermined commands are a pair of "sudo iptables -L│wc-│" and "sudo nft list ruleset". The command "sudo iptables -L│wc-│" is notified to the "iptables" of the table 15c, and the command "sudo nft list ruleset" is notified to the "nftables". Then, each table of "iptables" and "nftables" returns a response according to the command to the abnormality detection unit 17.
[0035] In the polling experiment using a pair of commands, the average value of the round-trip time when polling was executed 10 times was 0.03 seconds in the case of the command "sudo iptables -L│wc-│" and 0.08 seconds in the case of the command "sudo nft list ruleset".
[0036] The abnormality determination in the abnormality detection unit 17 determines that it is normal if the transmission destination path information is described in the command response result returned from each table 15c, and determines that it is abnormal if nothing is described.
[0037] <The Third Abnormality Detection Process> Next, FIG. 4 is a block diagram for explaining the third abnormality detection process by monitoring the daemon of the virtual switch 30 provided for each of the worker nodes 15J and 15K of the virtualization system failure isolation device 10 of the present embodiment. Note that the daemon of the virtual switch 30 is also referred to as the OVS daemon.
[0038] The daemon is a program that manages the packet destination in the virtual switch 30. The abnormality detection unit 17 monitors the OVS daemon and detects it as normal if the packet is properly transmitted, and as abnormal if it is not transmitted.
[0039] The abnormality detection unit 17 transmits a predetermined command (for example, "ps aux|grep ovs-vswitchd|grep \"db.sock\"|wc -│") to each virtual switch 30 of the worker nodes 15J and 15K by polling indicated by the reciprocating arrows Y5 and Y6, and determines whether it is normal or abnormal based on the response result returned from the virtual switch 30 according to the command.
[0040] In this polling experiment, the average value of the round-trip time when polling was executed 10 times was 0.03 seconds.
[0041] The abnormality determination in the abnormality detection unit 17 determines that it is normal if, for example, the "db.sock process" related to the destination is described in the command response result returned from each virtual switch 30, and determines that it is abnormal if it is not described.
[0042] <Fourth Abnormality Detection Process> Next, FIG. 5 is a block diagram for explaining a fourth abnormality detection process by monitoring the daemon of the container runtime 15d provided for each of the worker nodes 15J and 15K of the virtualization system failure isolation apparatus 10 of the present embodiment. Note that the daemon of the container runtime 15d is also referred to as a crio daemon. Crio (cri-o) is an open-source community-led container engine used in container-type virtualization technology.
[0043] Since the container runtime 15d is responsible for starting the containers of Pods 15a and 15b, it is possible to detect whether the containers are normally started by monitoring the container runtime 15d. Therefore, the abnormality detection unit 17 monitors the crio daemon and detects it as normal if the container is started, and as abnormal if it is not started.
[0044] The abnormality detection unit 17 sends a predetermined command (e.g., "systemctl status crio | grep Active") to the container runtime 15d for each of the worker nodes 15J and 15K by polling indicated by the reciprocating arrows Y7 and Y8, and determines whether it is normal or abnormal based on the response results returned from each container runtime 15d according to the command.
[0045] In this polling experiment, the average value of the round-trip time when polling was executed 10 times was 0.03 seconds.
[0046] The abnormality determination in the abnormality detection unit 17 determines that it is normal if "active (running)", which indicates the startup state of the crio daemon, is described in the command response result returned from each virtual switch 30, and determines that it is abnormal if it is described otherwise than "active (running)".
[0047] <Fifth Abnormality Detection Process> Next, FIG. 6 is a block diagram for explaining the fifth abnormality detection process by monitoring each of the worker nodes 15J and 15K of the virtualization system failure isolation device 10 of the present embodiment.
[0048] However, it is premised on a configuration in which the worker nodes 15J and 15K are created by a virtualization technology (virtual machine) using the physical machine 32. In the case of this configuration, the abnormality detection unit 17 exists on the physical machine 32 outside the virtual machine, and if the virtual machine is started by this abnormality detection unit 17, it detects that the container is normal, and if it is not started, it detects that the container is abnormal.
[0049] The abnormality detection unit 17 sends a predetermined command (e.g., "sudo virsh list") to each of the worker nodes 15J and 15K by polling indicated by the reciprocating arrows Y9 and Y10, and determines whether it is normal or abnormal based on the response results returned from each worker node 15J and 15K according to the command.
[0050] In this polling experiment, the average value of the round-trip time when polling was executed 10 times was 0.03 seconds.
[0051] The abnormality determination in the abnormality detection unit 17 is as follows: in the command response results returned from each worker node 15J, 15K, if "running" indicating the startup state of the target worker node 15J, 15K is described, it is determined as normal, and if it is described other than "running", it is determined as abnormal.
[0052] <Sixth Abnormality Detection Process> Next, FIG. 7 is a block diagram for explaining the sixth abnormality detection process by monitoring the DB (Data Base) 26a, 26b externally attached to the cluster 12 of the container system 20 of the virtualization system failure isolation device 10 of the present embodiment.
[0053] As an externally attached device of the cluster 12 (FIG. 1), there is a configuration in which DBs (also referred to as external DBs) 26a, 26b for storing data related to containers are connected to the worker nodes 15J, 15K via the network 22. At this time, the abnormality detection unit 17 is also connected to the worker nodes 15J, 15K via the network 22.
[0054] Here, since there is also a configuration in which a plurality of clusters 12 are connected to each other via the network 22, as shown in FIG. 7, even if the abnormality detection unit 17 is connected to the cluster 12 via the network 22, it is positioned as the abnormality detection unit 17 in the failure isolation device 10 in the same manner as shown in FIG. 1.
[0055] The abnormality detection unit 17 transmits a predetermined command to the external DBs 26a, 26b via the network 22 by polling indicated by the round-trip arrows Y11, Y12, and determines whether it is normal or abnormal based on the response results returned from each external DB 26a, 26b according to the command. The command in this case depends on the type of the external DBs 26a, 26b.
[0056] As response results, there are results related to response and liveness monitoring, and results related to exceeding the upper limit of the number of connections. Response and liveness monitoring monitors whether the external DBs 26a, 26b are started normally. That is, if the response result describes that the external DBs 26a, 26b are not started normally, the abnormality detection unit 17 determines it as an abnormality.
[0057] Exceeding the upper limit of the number of connections indicates that the number of containers to which the external DBs 26a, 26b are connected exceeds a predetermined threshold. That is, if the response result describes that the number of connected containers of the external DBs 26a, 26b exceeds the threshold, the abnormality detection unit 17 determines it as an abnormality.
[0058] In this polling experiment, the polling round-trip time depends on the types of the external DBs 26a, 26b.
[0059] Next, the abnormality handling processes at the time of the above-described first to fourth and sixth abnormality detections will be described with reference to FIGS. 8 to 11.
[0060] <First Abnormality Handling Process> FIG. 8 is a block diagram for explaining the first abnormality handling process of the virtualization system failure isolation apparatus 10 according to the present embodiment. It is assumed that the abnormality detection that requires the first abnormality handling process is any one of the first to fourth and sixth abnormality detections.
[0061] When the abnormality detection unit 17 shown in FIG. 8 detects an abnormality, the abnormality handling unit 18 performs the first abnormality handling process as follows. In this abnormality handling process, labels are attached to the Pods 15a, 15b in advance. The label describes information (label value) that enables or disables (disallows) access requests to the Pods 15a, 15b. The label value can be rewritten to be accessible or inaccessible by the router 14a of the master node 14J and thus can be changed.
[0062] The request is such that the router (communication oscillator section) 14a of the infrastructure node 14K notifies each of the Pods 15a, 15b of the worker nodes 15J, 15K via the endpoint setting section 14h. The endpoint setting section 14h receives service information related to the communication indicated by the arrow Y20 from the opposing device 24 and is shared by a plurality of Pods 15a, 15b of each worker node 15J, 15K.
[0063] When the label values of the Pods 15a, 15b are accessible, the Pods 15a, 15b can receive the request and communication etc. becomes possible. On the other hand, when the label values are inaccessible, the Pods 15a, 15b cannot receive the request and communication etc. becomes impossible.
[0064] In this way, in the 1-to-N configuration of the endpoint setting section 14h and the Pods 15a, 15b, when an abnormality is detected by the abnormality detection section 17, by changing the label value of the abnormal Pod 15a to a value different from the access setting (for example, accessible) to the Pods 15a, 15b via the endpoint setting section 14h (inaccessible), communication to the abnormal Pod 15a can be suppressed.
[0065] Next, the operation of the first abnormality handling process will be described with reference to the flowcharts shown in FIGS. 8 and 9. As a prerequisite, it is assumed that the label values of the Pods 15a, 15b for each of the worker nodes 15J, 15K are accessible.
[0066] In step S1 shown in FIG. 9, it is assumed that an abnormality of the Pod 15a of the worker node 15J is detected by the abnormality detection section 17 shown in FIG. 8.
[0067] At the time of this abnormality detection, in step S2, the abnormality handling section 18 transmits a command (for example, "oc label pod Pod name lb=label value --overwrite=true") that makes the label value of the abnormal Pod 15a inaccessible, as indicated by the arrow Y14, to the router 14a of the master node 14J.
[0068] In step S3, the router 14a that has received the command changes the label value of the abnormal Pod15a to inaccessible, as indicated by the arrow Y15.
[0069] Here, the execution time from when the command is sent from the above abnormal handling unit 18 until the label value of the abnormal Pod15a is changed is such that the average value of the execution times when the command is executed 10 times is 0.38 seconds.
[0070] In step S4, after the label change in step S3 above, the router 14a of the infrastructure node 14K notifies a communication request to each Pod15a of the worker nodes 15J and 15K via the endpoint setting unit 14h, as indicated by the arrows Y16 and Y17.
[0071] Here, in step S5, it is determined whether the label value of each Pod15a of the worker nodes 15J and 15K is accessible.
[0072] Suppose the request in step S4 above is notified to the abnormal Pod15a of the worker node 15J as indicated by the arrow Y16. In this case, since its label value is inaccessible, it is determined in step S5 that it is not accessible (inaccessible) (No).
[0073] In this case, in step S6, as indicated by the cross mark, the abnormal Pod15a becomes unable to receive requests, so communication and the like of the abnormal Pod15a become impossible. In other words, the abnormal Pod15a enters a stopped state and is separated from the normal Pod15b.
[0074] On the other hand, suppose the request in step S4 above is notified to the abnormal Pod15a of the worker node 15K as indicated by the arrow Y17. In this case, since its label value is accessible, it is determined in step S5 that it is accessible (Yes).
[0075] In this case, in step S7, Pod15a of worker node 15K receives the request and communication becomes possible.
[0076] <Second Abnormality Handling Process> FIG. 10 is a block diagram for explaining the second abnormality handling process of the virtualization system failure isolation apparatus 10 of the present embodiment. It is assumed that the abnormality detection requiring the second abnormality handling process is any one of the first to fourth and sixth abnormality detections. Further, it is assumed that the second abnormality handling process is an abnormality handling process at the time of detecting an abnormality of a container in an Istio environment (described later).
[0077] Here, a plurality of applications 15a and 15b in the computing resource cluster 15 (FIG. 1) are referred to as components. The component has a function of a proxy for performing relay communication. Note that the component is, in other words, also a plurality of Pods 15a and 15b and a plurality of containers.
[0078] The Istio environment refers to an environment in which components are operating, and is realized by a control plane (also referred to as a plane) 14g provided in the infrastructure node 14K shown in FIG. 10. The plane 14g has a virtual endpoint setting unit 4g1. The virtual endpoint setting unit 4g1 virtually configures an endpoint setting unit within the plane 14g, has the same function as the endpoint setting unit 14h (FIG. 8), and is shared by a plurality of Pods 15a and 15b of each worker node 15J and 15K.
[0079] Also, in the Istio environment, various functions are realized in combination with a proxy called "Envoy" that operates on a container. In the Istio environment, processing of Envoy, management of collected data, monitoring of services, etc. are performed. It is Envoy that actually mediates communication between microservices in the Istio environment. Istio performs processing for setting and managing the operation of Envoy and acquiring information on network traffic indicated by the arrow Y20 from the counter device 24 via Envoy.
[0080] When the abnormality detection unit 17 detects an abnormality in the container, the abnormality handling unit 18 performs a second abnormality handling process. That is, assuming that the abnormality detection unit 17 detects an abnormality in Pod15a of the worker node 15J, the abnormality handling unit 18 sends a command (for example, "oc label pod Pod name lb=label value --overwrite=true") that makes the label of the abnormal Pod15a inaccessible, as indicated by the arrow Y14, to the router 14a of the master node 14J. The router 14a that receives the command changes the label of the abnormal Pod15a to be inaccessible, as indicated by the arrow Y15.
[0081] Here, the execution time from when the command is sent from the abnormality handling unit 18 until the label of the abnormal Pod15a is changed was 0.38 seconds, which is the average value of the execution times when the command was sent 10 times.
[0082] After the label is changed, the plane 14g of the infrastructure node 14K notifies a communication request to each Pod15a of the worker nodes 15J and 15K, as indicated by the arrows Y18 and Y19, via the virtual endpoint setting unit 4g1. At this time, the abnormal Pod15a whose label value of the worker node 15J has been made inaccessible cannot be accessed, as indicated by the cross mark. For this reason, the abnormal Pod15a will be in a stopped state and separated from the normal Pod15b.
[0083] In this way, using the Istio environment of the plane 14g, when the abnormality detection unit 17 detects an abnormality in the 1-to-N configuration of the virtual endpoint setting unit 4g1 and the Pods 15a and 15b, the label value of the abnormal Pod15a is changed to a value (inaccessible) different from the access setting (for example, accessible) to the Pods 15a and 15b via the virtual endpoint setting unit 4g1. As a result, communication to the abnormal Pod15a can be suppressed.
[0084] <Third Abnormality Handling Process> FIG. 11 is a block diagram for explaining the third abnormality handling process of the virtualization system failure isolation apparatus 10 according to the present embodiment. Assume that the abnormality detection that requires the third abnormality handling process is the fifth abnormality detection.
[0085] In the third abnormality handling process, as a prerequisite, the container management unit 14f attaches labels indicating the nodes to which the Pods 15a and 15b belong to each of the worker nodes 15J and 15K by polling.
[0086] Assume that the abnormality detection unit 17 on the physical machine 32 detects an abnormality of a worker node 15J, which is a virtual machine, as indicated by an arrow Y9. At this time, the label of the Pod 15a is also detected.
[0087] Next, as indicated by an arrow Y16, the abnormality handling unit 18 transmits a command (for example, "oc delete pod -l node label specification --grace-period 0 --force") for forcibly deleting the Pod 15a with the label of the worker node 15J for which abnormality detection is to be handled, to the router 14a of the master node 14J. When transmitting the command, the label information of the abnormal Pod 15a is also transmitted.
[0088] The router 14a notifies the received command to the container management unit 14f. The container management unit 14f forcibly deletes (cross mark) the abnormal Pod 15a of the worker node 15J indicated by the label information of the command, as indicated by an arrow Y15a. Since this forced deletion is executed immediately, it is possible to accelerate the restart of the Pod 15a corresponding to the abnormal Pod 15a.
[0089] <Fourth Abnormality Handling Process> FIG. 12 is a block diagram for explaining the fourth abnormality handling process of the virtualization system failure isolation apparatus 10 according to the present embodiment. Assume that the abnormality detection that requires the fourth abnormality handling process is the fifth abnormality detection.
[0090] Suppose that the abnormality detection unit 17 on the physical machine 32 detects an abnormality in, for example, the worker node 15J which is a virtual machine, as indicated by the arrow Y9. In this case, as indicated by the arrow Y17, the abnormality handling unit 18 sends a command (for example, "oc adm drain node name --delete-local-data --ignore-daemonsets --grace-period 0") for forcibly deleting all the Pods 15a of the worker node 15J where the abnormality was detected to the router 14a of the master node 14J.
[0091] The router 14a notifies the container management unit 14f of the received command. As indicated by the arrow Y15b, the container management unit 14f forcibly deletes (marked with an 'x') all the Pods 15a of the worker node 15J indicated by the command. After this forced deletion, the corresponding Pods 15a for all the deleted Pods 15a are restarted on another node (for example, the worker node 15K).
[0092] <Hardware Configuration> The virtualization system failure isolation device 10 according to the above-described embodiment is realized by, for example, a computer 100 having a configuration as shown in FIG. 13. The computer 100 includes a CPU (Central Processing Unit) 101, a ROM (Read Only Memory) 102, a RAM (Random Access Memory) 103, an HDD (Hard Disk Drive) 104, an input / output I / F (Interface) 105, a communication I / F 106, and a media I / F 107.
[0093] The CPU 101 operates based on a program stored in the ROM 102 or the HDD 104 and controls each functional unit. The ROM 102 stores a boot program executed by the CPU 101 when the computer 100 is started up, a program related to the hardware of the computer 100, and the like.
[0094] The CPU 101 controls an output device 111 such as a printer or a display, and an input device 110 such as a mouse or a keyboard via the input / output I / F 105. The CPU 101 acquires data from the input device 110 or outputs the generated data to the output device 111 via the input / output I / F 105.
[0095] The HDD 104 stores programs executed by the CPU 101 and data used by the programs. The communication I / F 106 receives data from another device (not shown) via the communication network 112 and outputs it to the CPU 101, and also transmits the data generated by the CPU 101 to another device via the communication network 112.
[0096] The media I / F 107 reads a program or data stored in the recording medium 113 and outputs it to the CPU 101 via the RAM 103. The CPU 101 loads a program related to the target process from the recording medium 113 onto the RAM 103 via the media I / F 107 and executes the loaded program. The recording medium 113 is an optical recording medium such as a DVD (Digital Versatile Disc) or a PD (Phase change rewritable Disk), a magneto-optical recording medium such as an MO (Magneto Optical disk), a magnetic recording medium, a conductor memory tape medium, or a semiconductor memory, etc.
[0097] For example, when the computer 100 functions as the virtualization system failure isolation device 10 according to the embodiment, the CPU 101 of the computer 100 realizes the functions of the virtualization system failure isolation device 10 by executing the program loaded on the RAM 103. Also, the data in the RAM 103 is stored in the HDD 104. The CPU 101 reads and executes a program related to the target process from the recording medium 113. In addition, the CPU 101 may read a program related to the target process from another device via the communication network 112.
[0098] <Effects of the Embodiment> The effects of the virtualization system failure isolation device 10 according to the embodiments of the present invention will be described.
[0099] (1a) The failure isolation device 10 includes a computing resource cluster 15 that is virtually created by container virtualization software on a physical machine and clusters and arranges the virtually created containers, and a cluster management unit 14 that manages the control related to the arrangement and operation of the virtually created and clustered containers (Pod15a, 15b). Further, an abnormality detection unit 17 that is created outside the virtually created computing resource cluster 15 and the cluster management unit 14 and detects abnormalities in the containers, and an abnormality response unit 18 that transmits a command for instructing the cluster management unit 14 to stop the abnormal containers detected by the abnormality detection unit 17 are provided. The cluster management unit 14 is configured to stop the abnormal containers according to the command.
[0100] According to this configuration, when a failure occurs in a container, the abnormality detection unit 17 and the abnormality response unit 18 are arranged outside the virtualization configuration part in the container virtualization software that creates the container. The abnormality detection unit 17 detects the abnormality of the container, and the abnormality response unit 18 transmits a command for stopping the detected abnormal container to the cluster management unit 14. The cluster management unit 14 stops the abnormal container according to the command.
[0101] Since the abnormality detection unit 17 and the abnormality response unit 18 do not participate in the container virtualization software, recovery can be achieved earlier than recovery by the container failure recovery function of the container virtualization software. To further explain the basis for this earlier recovery, in the above failure recovery function, the cycle for performing failure monitoring can only be set to a predetermined cycle, but in the present invention, regardless of the monitoring cycle, a container failure can be detected and an abnormal container can be stopped. Therefore, recovery can be achieved earlier than recovery by the above failure recovery function.
[0102] (2a) It acquires service information related to communication from the counterpart device 24 that communicates via the network 22, and includes an endpoint setting unit 14h shared by a plurality of containers in the computing resource cluster 15. The cluster management unit 14 is configured to access requests to a plurality of containers via the endpoint setting unit 14H. In this configuration, a label is assigned to the clustered containers, on which either an accessible label value that enables access to requests from the cluster management unit 14 or an inaccessible label value that disables access is described. The abnormality handling unit 18 transmits a command to the cluster management unit 14 that gives an instruction to make the label value of the abnormal container detected by the abnormality detection unit 17 inaccessible. The cluster management unit 14 is configured to change the label value of the abnormal container instructed by the command to be inaccessible.
[0103] According to this configuration, a label with an accessible or inaccessible label value is assigned to the container, and the label value of the abnormal container detected by the abnormality detection unit 17 is changed to be inaccessible. For this reason, when the cluster management unit 14 accesses a request to the abnormal container via the endpoint setting unit 14H, the access becomes impossible. In other words, the abnormal container enters a stopped state and is separated from the normal containers.
[0104] (3a) The cluster management unit 14 includes a virtual endpoint setting unit that is virtually created and integrates data of a plurality of containers in the computing resource cluster 15 to monitor the containers. The cluster management unit 14 is configured to access requests to a plurality of containers via the virtual endpoint setting unit. In this configuration, a label is assigned to the clustered containers, on which either an accessible label value that enables access to requests from the cluster management unit 14 or an inaccessible label value that disables access is described. The abnormality handling unit 18 transmits a command to the cluster management unit 14 that gives an instruction to make the label value of the abnormal container detected by the abnormality detection unit 17 inaccessible. The cluster management unit 14 is configured to change the label value of the abnormal container instructed by the command to be inaccessible.
[0105] According to this configuration, a label with a label value indicating whether a container is accessible or not is assigned, and the label value of the abnormal container detected by the abnormality detection unit 17 is changed to inaccessible. Therefore, when the cluster management unit 14 accesses an abnormal container via the virtual endpoint setting unit, the access becomes impossible. In other words, the abnormal container becomes in a stopped state and is separated from normal containers.
[0106] (4a) The abnormality detection unit 17 and the abnormality response unit 18 are created on a physical machine. A label indicating the computing resource cluster 15 to which the container belongs is assigned to the clustered containers. The abnormality detection unit 17 detects that a container is normal if the virtually created computing resource cluster 15 is running, and detects that a container is abnormal if it is stopped. The abnormality response unit 18 transmits a command for instructing forced deletion of the detected abnormal container and the label information of the abnormal container to the cluster management unit 14. The cluster management unit 14 is configured to forcibly delete the abnormal container indicated by the label information according to the instruction of the command.
[0107] According to this configuration, after detection of an abnormality in the computing resource cluster 15, the containers with the label of the detected abnormal computing resource cluster 15 are searched by the label information and immediately and forcibly deleted. Therefore, it is possible to accelerate the restart of the abnormal container.
[0108] (5a) The abnormality detection unit 17 and the abnormality response unit 18 are created on a physical machine. The abnormality detection unit 17 detects that a container is normal if the virtually created computing resource cluster 15 is running, and detects that a container is abnormal if it is stopped. The abnormality response unit 18 transmits a command for instructing forced deletion of all containers of the detected computing resource cluster 15 to the cluster management unit 14. The cluster management unit 14 forcibly deletes all containers of the computing resource cluster 15 according to the command, and restarts the containers corresponding to all the deleted containers in another computing resource cluster 15.
[0109] According to this configuration, after the detection of an abnormality in the computing resource cluster 15, all the containers in the detected abnormal computing resource cluster 15 are immediately and forcibly deleted, and all the deleted containers are restarted in another computing resource cluster 15. Therefore, it is possible to accelerate the restart of all the containers deleted after the detection of the abnormality.
[0110] <Effect> (1) A computing resource cluster that is virtually created by container virtualization software on a physical machine and clusters and arranges the virtually created containers, a cluster management unit that manages the control related to the arrangement and operation of the virtually created and clustered containers, an abnormality detection unit that is created outside the virtually created computing resource cluster and the cluster management unit and detects abnormalities of the containers, and an abnormality response unit that is created outside and transmits a command for instructing the cluster management unit to stop the abnormal containers detected by the abnormality detection unit, wherein the cluster management unit stops the abnormal containers according to the command, and the virtualization system failure isolation device is characterized in that.
[0111] According to this configuration, when a container fails, the abnormality detection unit and the abnormality response unit are arranged outside the virtualization configuration part in the container virtualization software that creates the container. The abnormality detection unit detects the abnormality of the container, and the abnormality response unit transmits a command for stopping the detected abnormal container to the container management unit. The container management unit stops the abnormal container according to the command. Since the abnormality detection unit and the abnormality response unit do not participate in the container virtualization software, it is possible to recover faster than the recovery by the container failure recovery function of the container virtualization software. To further explain the basis for this faster recovery, in the above failure recovery function, the cycle for performing failure monitoring can only be set to a predetermined cycle, but in the present invention, regardless of the monitoring cycle, it is possible to detect the failure of the container and stop the abnormal container. Therefore, it is possible to recover faster than the recovery by the above failure recovery function.
[0112] (2) Obtain service information related to communication from a counterpart device that communicates via a network, and include an endpoint setting unit shared by a plurality of containers in the computing resource cluster. In a configuration where the cluster management unit accesses requests to a plurality of containers via the endpoint setting unit, a label indicating either an accessible label value that enables access to requests from the cluster management unit or an inaccessible label value that disables access is assigned to the clustered containers. The abnormality handling unit transmits a command to the cluster management unit to instruct to set the label value of the abnormal container detected by the abnormality detection unit to inaccessible. The cluster management unit changes the label value of the abnormal container instructed by the command to inaccessible. The virtualization system failure isolation device according to claim 1, characterized in that.
[0113] According to this configuration, a label with an accessible or inaccessible label value is assigned to the container, and the label value of the abnormal container detected by the abnormality detection unit is changed to inaccessible. Therefore, when the cluster management unit accesses a request to the abnormal container via the endpoint setting unit, access becomes impossible. In other words, the abnormal container will be in a stopped state and separated from the normal containers.
[0114] (3) The cluster management unit includes a virtual endpoint setting unit that integrates data of a plurality of containers of the computationally virtualized resource cluster that is virtually created, and monitors the containers. In a configuration where the cluster management unit accesses requests to a plurality of containers via the virtual endpoint setting unit, a label is attached to the clustered containers, which indicates either an accessible or inaccessible label value that allows or disallows access to requests from the cluster management unit. The abnormality handling unit transmits a command to the cluster management unit to instruct the label value of the abnormal container detected by the abnormality detection unit to be inaccessible. The cluster management unit changes the label value of the abnormal container instructed by the command to be inaccessible. The virtualization system failure isolation apparatus according to claim 1 is characterized in that.
[0115] According to this configuration, a label with an accessible or inaccessible label value is attached to the container, and the label value of the abnormal container detected by the abnormality detection unit is changed to be inaccessible. Therefore, when the cluster management unit accesses a request to the abnormal container via the virtual endpoint setting unit, access becomes impossible. In other words, the abnormal container becomes a stopped state and is separated from the normal containers.
[0116] (4) The abnormality detection unit and the abnormality handling unit are created on the physical machine, and a label indicating the computationally virtualized resource cluster to which the container belongs is attached to the clustered containers. The abnormality detection unit detects that the container is normal if the computationally virtualized resource cluster that is virtually created is running, and detects that the container is abnormal if it is stopped. The abnormality handling unit transmits a command instructing the forced deletion of the detected abnormal container and the label information of the abnormal container to the cluster management unit. The cluster management unit forcibly deletes the abnormal container indicated by the label information according to the instruction of the command. The virtualization system failure isolation apparatus according to claim 1 is characterized in that.
[0117] According to this configuration, after detecting an abnormality in the computing resource cluster, the containers labeled with the detected abnormal computing resource cluster are searched for by label information and immediately and forcibly deleted. Therefore, it is possible to accelerate the restart of the abnormal containers.
[0118] (5) The abnormality detection unit and the abnormality handling unit are created on the physical machine. The abnormality detection unit detects that the container is normal if the virtually created computing resource cluster is running, and detects that the container is abnormal if it is stopped. The abnormality handling unit sends a command to the cluster management unit to instruct the forced deletion of all containers in the detected computing resource cluster. The cluster management unit forcibly deletes all containers in the computing resource cluster corresponding to the command, and restarts the containers corresponding to all the deleted containers in another computing resource cluster. The virtualization system failure isolation device according to claim 1, characterized in that.
[0119] According to this configuration, after detecting an abnormality in the computing resource cluster, all containers in the detected abnormal computing resource cluster are immediately and forcibly deleted, and all the deleted containers are restarted in another computing resource cluster. Therefore, it is possible to accelerate the restart of all the containers deleted after the abnormality detection.
[0120] In addition, regarding the specific configuration, appropriate changes can be made without departing from the gist of the present invention.
Explanation of Signs
[0121] 10 Virtualization system failure isolation device 14 Cluster management unit 14a Communication distribution unit 14b Computing resource operation unit 14c Computing resource management unit 14d Container configuration reception unit 14e Container placement destination determination unit 14f Container management unit 15 Computing resource cluster 15a, 15b Application 17 Abnormality detection unit 18 Abnormality response unit
Claims
1. A computing resource cluster that is virtually created by container virtualization software on a physical machine and clusters and arranges the virtually created containers, A cluster management unit that manages the control related to the arrangement and operation of the virtually created and clustered containers, An abnormality detection unit that is created outside the virtually created computing resource cluster and the cluster management unit and detects abnormalities of the containers, An abnormality response unit that is created outside and sends a command for instructing to stop an abnormal container detected by the abnormality detection unit to the cluster management unit comprising obtaining service information related to communication from a counterpart device that communicates via a network, and comprising an endpoint setting unit shared by a plurality of containers in the computing resource cluster, and in a configuration where the cluster management unit accesses requests to a plurality of containers via the endpoint setting unit, attaching a label to the clustered containers, on which either an accessible label value that enables access to requests from the cluster management unit or an inaccessible label value that disables access is described, the abnormality response unit sends a command for instructing to set the label value of the abnormal container detected by the abnormality detection unit to inaccessible to the cluster management unit, the cluster management unit stops the abnormal container by changing the label value of the abnormal container instructed by the command to inaccessible A virtualization system failure isolation device characterized by the above.
2. A computing resource cluster that is virtually created by container virtualization software on a physical machine and clusters and arranges the virtually created containers, A cluster management unit that manages the control related to the arrangement and operation of the virtually created and clustered containers, An abnormality detection unit that is created outside the virtually created computing resource cluster and the cluster management unit and detects an abnormality of the container; An abnormality response unit that is created outside and transmits a command for instructing the cluster management unit to stop the abnormal container detected by the abnormality detection unit; Comprising; The cluster management unit acquires service information related to communication from a counterpart device that is virtually created and communicates via a network, and includes a virtual endpoint setting unit that is shared by a plurality of containers in the computing resource cluster and monitors the containers. In a configuration where the cluster management unit accesses a plurality of containers via the virtual endpoint setting unit, A label is attached to the clustered container, and the label describes either an accessible or inaccessible label value that enables or disables access to a request from the cluster management unit; The abnormality response unit transmits a command for instructing the cluster management unit to make the label value of the abnormal container detected by the abnormality detection unit inaccessible; The cluster management unit stops the abnormal container by changing the label value of the abnormal container instructed by the command to be inaccessible; A virtualization system failure isolation device characterized by the above.
3. A computing resource cluster that is virtually created by container virtualization software on a physical machine and clusters and arranges the virtually created containers; A cluster management unit that manages the control related to the arrangement and operation of the virtually created and clustered containers; An abnormality detection unit that is created outside the virtually created computing resource cluster and the cluster management unit and detects an abnormality of the container; An abnormality response unit that is created outside and transmits a command for instructing the cluster management unit to stop the abnormal container detected by the abnormality detection unit; Comprising; The abnormal detection unit and the abnormal response unit are created on the physical machine, A label indicating the computing resource cluster to which the container belongs is assigned to the clustered container, If the virtualized computing resource cluster is running, the abnormal detection unit detects that the container is normal; if it is stopped, the abnormal detection unit detects that the container is abnormal. The abnormal response unit transmits a command for instructing forced deletion of the detected abnormal container and the label information of the abnormal container to the cluster management unit, The cluster management unit forcibly deletes the abnormal container indicated by the label information according to the instruction of the command. A virtualization system failure isolation device characterized by the above.
4. A computing resource cluster that is virtually created by container virtualization software on a physical machine and clusters and arranges the virtually created containers, A cluster management unit that manages the control related to the arrangement and operation of the virtually created and clustered containers, An abnormal detection unit that is created outside the virtually created computing resource cluster and the cluster management unit and detects abnormalities of the containers, An abnormal response unit that is created outside and transmits a command for instructing to stop the abnormal containers detected by the abnormal detection unit to the cluster management unit are provided, The abnormal detection unit and the abnormal response unit are created on the physical machine, If the virtualized computing resource cluster is running, the abnormal detection unit detects that the container is normal; if it is stopped, the abnormal detection unit detects that the container is abnormal. The abnormal response unit transmits a command for instructing forced deletion of all containers in the detected computing resource cluster to the cluster management unit, The cluster management unit forcibly deletes all containers of the computing resource cluster according to the command, and restarts the containers corresponding to all the deleted containers in another computing resource cluster. A virtualization system failure isolation device characterized by the above.
5. A virtualization system failure isolation method by a virtualization system failure isolation device, The virtualization system failure isolation device is On a computing resource cluster virtually created by container virtualization software on a physical machine, clustering and arranging the virtually created containers; Creating a cluster management unit that manages control related to the arrangement and operation of the clustered containers virtually; Assigning a label to the clustered containers, on which either an accessible label value that enables access to requests from the cluster management unit or an inaccessible label value that disables access is described; Obtaining service information related to communication from a counterpart device that communicates via a network, and having an endpoint setting unit shared by a plurality of containers in the computing resource cluster, wherein the cluster management unit accesses requests to a plurality of containers via the endpoint setting unit; Detecting an abnormality of the container outside the virtually created computing resource cluster and the cluster management unit; Sending a command to the cluster management unit to give an instruction to make the label value of the detected abnormal container inaccessible outside; The cluster management unit stops the abnormal container by changing the label value of the abnormal container instructed by the command to be inaccessible; A virtualization system failure isolation method characterized by executing the above.
Citation Information
Patent Citations
Application container exception handling method and device
CN111414229A
Method and system for centralized networking and storage
JP2017512350A
Container daemon, information processing device, container-type virtualization system, packet distribution method, and program
JP2021027398A
Sensor information processing system using container orchestration
WO2020184362A1