System, method and program for managing virtualization infrastructure

The system addresses service continuity issues in virtualization infrastructures by prioritizing and reallocating workloads to maintain high-priority services during device failures, enhancing resilience and resource efficiency.

WO2025177500A1PCT designated stage Publication Date: 2025-08-28SOFTBANK CORPORATION
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/006379
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-21
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

Existing virtualization infrastructures face challenges in maintaining service continuity when failures occur in infrastructure component devices, such as racks or container-type data centers, leading to the shutdown of all physical servers and disruption of services.

Method used

A system that classifies services into priority groups and reallocates workloads to ensure high-priority services continue by stopping lower-priority services on non-failed devices, utilizing surplus resources to restart high-priority services on failed devices.

Benefits of technology

Ensures high-priority services can be continuously provided even in the event of failures, reducing excess resources and minimizing service disruptions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024006379_28082025_PF_FP_ABST
    Figure JP2024006379_28082025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention continuously provides a high-priority service on a virtualization infrastructure even when a failure occurs in any of a plurality of infrastructure constituent devices constituting the virtualization infrastructure, and all physical resources in the infrastructure constituent device in which the failure has occurred stop operation. This system comprises: an information storage unit that stores information relating to physical resources of the entire cluster in which a plurality of workloads constructed on a virtualization infrastructure can operate, and information relating to the respective physical resources of a plurality of infrastructure constituent devices; and an information processing unit that classifies a plurality of services provided by the cluster into a plurality of service groups having different priorities, and for each of the plurality of respective service groups, arranges the plurality of workloads to be executed in the cluster and manages the physical resources required for the respective service groups on the basis of the information relating to the physical resources of the entire cluster and the information relating to the respective physical resources of the plurality of infrastructure constituent devices.
Need to check novelty before this filing date? Find Prior Art

Description

System, method and program for managing virtualization infrastructure

[0001] The present invention relates to management of a virtualization infrastructure on which virtual machines can be constructed.

[0002] 2. Description of the Related Art Conventionally, a virtualization platform for building virtual machines, which is configured with a plurality of racks each having a plurality of physical servers (physical resources), has been known (see, for example, Patent Document 1).

[0003] International Publication No. 2011 / 043317

[0004] If a failure occurs in one of the racks that are multiple infrastructure-component devices that make up the virtualization infrastructure (for example, a failure in the power supply installed in each rack), all physical servers (physical resources) in the failed rack may stop operating, potentially rendering it impossible to perform processing (for example, provide services) using virtual machines built on the virtualization infrastructure. Note that a similar issue can also arise when a virtualization infrastructure is configured with multiple container-type data centers (infrastructure-component devices).

[0005] A system according to one aspect of the present invention is a system for managing a virtualization infrastructure including a plurality of infrastructure-constituting devices each having a plurality of physical resources. The system includes an information storage unit that stores information about the overall physical resources of a cluster on which a plurality of workloads can run built on the virtualization infrastructure and information about the physical resources of each of the plurality of infrastructure-constituting devices, and an information processing unit that classifies a plurality of services provided by the cluster on which the plurality of workloads can run built on the virtualization infrastructure into a plurality of service groups having different priorities, and that, for each of the plurality of service groups, allocates the plurality of workloads to be executed on the cluster and manages the physical resources required by the service group based on the information about the overall physical resources of the cluster and the information about the physical resources of each of the plurality of infrastructure-constituting devices.

[0006] In the system, when a failure occurs in one of the multiple infrastructure component devices, the system may be provided with a control unit that checks the priority of the service groups to which each of the multiple workloads that were running on the failed infrastructure component device belongs, and if it is confirmed that a workload of a service group with a higher priority than a predetermined priority was running on the failed infrastructure component device, stops the workload of the low-priority service group that was running on the non-failed infrastructure component device, and controls the system to restart the workload of the high-priority service group that was running on the failed infrastructure component device using the physical resources freed up by the stopping of the workload on the non-failed infrastructure component device.

[0007] In the system, the information processing unit may calculate, for each resource component, the physical resources that can be used in the event of a failure in an infrastructure component device with the largest physical resources among the plurality of infrastructure component devices, based on information on the physical resources of each of the plurality of infrastructure component devices, and determine the specified priority based on the calculation results for each resource component of the available physical resources, information for each resource component of the physical resources corresponding to each of the plurality of priorities, and competitive allocation setting conditions for each resource component.

[0008] In the system, the information processing unit may provide the user with impact information in the event that a failure occurs in an infrastructure component device with the greatest physical resources among the plurality of infrastructure component devices.

[0009] A method according to another aspect of the present invention is a method for managing a virtualization infrastructure including a plurality of infrastructure-constituting devices each having a plurality of physical resources, the method including: storing information about the overall physical resources of a cluster on which a plurality of workloads can run built on the virtualization infrastructure and information about the physical resources of each of the plurality of infrastructure-constituting devices; classifying a plurality of services provided by the cluster on which the plurality of workloads can run built on the virtualization infrastructure into a plurality of service groups having different priorities; and, for each of the plurality of service groups, allocating a plurality of workloads to be executed in the cluster and managing the physical resources required by the service group based on the information about the overall physical resources of the cluster and the information about the physical resources of each of the plurality of infrastructure-constituting devices.

[0010] The method may include, when a failure occurs in one of the plurality of infrastructure component devices, checking the priority of the service groups to which each of the plurality of workloads running on the infrastructure component device where the failure belongs; and, if it is confirmed that a workload of a service group with a higher priority than a predetermined priority was running on the infrastructure component device where the failure occurred, stopping the workload of a lower priority service group running on the infrastructure component device where the failure did not occur; and restarting the workload of the higher priority service group that was running on the infrastructure component device where the failure occurred using the physical resources freed up by stopping the workload on the infrastructure component device where the failure did not occur.

[0011] The method may include calculating, for each resource component, the physical resources that can be used in the event of a failure in an infrastructure component device with the largest physical resources among the plurality of infrastructure component devices, based on information about the physical resources of each of the plurality of infrastructure component devices; and determining the predetermined priority based on the calculation results for each resource component of the available physical resources, information for each resource component of the physical resources corresponding to each of the plurality of priorities, and competitive allocation setting conditions for each resource component.

[0012] The method may include providing a user with impact information in the event that a failure occurs in an infrastructure component device with the greatest physical resources among the plurality of infrastructure components.

[0013] According to yet another aspect of the present invention, there is provided a program executed by a computer or processor provided in a system for managing a virtualization infrastructure including a plurality of infrastructure-constituting devices each having a plurality of physical resources, the program including: program code for storing information about the overall physical resources of a cluster on which a plurality of workloads can run constructed on the virtualization infrastructure and information about the physical resources of each of the plurality of infrastructure-constituting devices; program code for classifying a plurality of services provided by the cluster on which a plurality of workloads can run constructed on the virtualization infrastructure into a plurality of service groups having different priorities; and program code for allocating a plurality of workloads to be executed in the cluster and managing the physical resources required by the service group for each of the plurality of service groups based on the information about the overall physical resources of the cluster and the information about the physical resources of each of the plurality of infrastructure-constituting devices.

[0014] The program may include, when a failure occurs in one of the multiple infrastructure component devices, program code for checking the priority of the service groups to which each of the multiple workloads running on the failed infrastructure component device belongs; program code for stopping workloads of low priority service groups running on the infrastructure component device where the failure did not occur when it is confirmed that workloads of high priority service groups with a priority equal to or higher than a predetermined priority were running on the failed infrastructure component device; and program code for restarting workloads of the high priority service groups running on the failed infrastructure component device using physical resources freed up by stopping the workloads on the infrastructure component device where the failure did not occur.

[0015] The program may include program code for calculating, for each resource component, the physical resources available for use in the event of a failure in an infrastructure component device with the largest physical resources among the plurality of infrastructure component devices, based on information about the physical resources of each of the plurality of infrastructure component devices; and program code for determining the predetermined priority based on the calculation results for each resource component of the available physical resources, information about each resource component of the physical resources corresponding to each of the plurality of priorities, and competitive allocation setting conditions for each resource component.

[0016] The program may include program code for providing a user with impact information in the event that a failure occurs in an infrastructure component device with the greatest physical resources among the plurality of infrastructure components.

[0017] In the system, the method, and the program, the physical resource may be a physical server, and the infrastructure component device may be a rack having a plurality of physical servers and a power supply.

[0018] In the system, the method, and the program, the physical resource may be a rack having a plurality of physical servers, and the infrastructure component device may be a container-type data center having a plurality of racks and a power source.

[0019] The program may include a machine-learned model.

[0020] The virtualization platform may be used, for example, in a data center or in a RAN Intelligent Controller (RIC) used in a RAN (Radio Access Network) of a mobile communication system.

[0021] According to the present invention, even if a failure occurs in one of the multiple infrastructure component devices that make up the virtualization infrastructure and all physical resources within the failed infrastructure component device stop operating, high-priority services can be continuously provided on the virtualization infrastructure.

[0022] FIG. 1 is a schematic diagram showing an example of the overall configuration of a virtualization platform to which a system according to an embodiment can be applied. FIG. 2 is an explanatory diagram showing an example of a failure occurring in a single physical server in a virtualization platform according to a reference example. FIG. 3 is an explanatory diagram showing an example of a rack-scale failure occurring in a virtualization platform according to an embodiment. FIG. 4 is an explanatory diagram showing an example of the placement of workloads of multiple priorities in a virtualization platform according to an embodiment. FIG. 5 is an explanatory diagram showing an example of the placement of workloads of multiple priorities when a rack-scale failure occurs in a virtualization platform according to an embodiment. FIG. 6 is an explanatory diagram showing an example of the configuration of a management system that manages a virtualization platform according to an embodiment. FIG. 7 is a flowchart showing an example of generating a virtual environment resource-related table in a virtualization platform according to an embodiment. FIG. 8 is a flowchart showing an example of generating a physical resource-related table in a virtualization platform according to an embodiment. FIG. 9 is a flowchart showing an example of calculating the priority of a service group that restarts a workload when a rack-scale failure occurs in a virtualization platform according to an embodiment.

[0023]

[0016] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. An example of a system according to the embodiment described herein defines service group units and priorities for each service group for resources of workloads running in a virtual environment on a virtualization platform. When a resource situation makes it difficult to start all workloads (service groups) due to a failure of a platform component (rack), for example, the system stops workloads of low-priority service groups and controls so that workloads of high-priority service groups can be started. In addition, the control involves defining and managing resource components (CPU (Central Processing Unit), memory, GPU (Graphics Processing Unit)) of physical resources consumed by each service group and conflict allocation settings for physical resources (e.g., allocating four virtual CPUs to one physical CPU actually installed in a physical server).

[0024] In particular, the virtualization infrastructure managed by the system according to this embodiment can be used for a RAN Intelligent Controller (RIC) used in a RAN (Radio Access Network) of a mobile communication system.

[0025] FIG. 1 is a schematic diagram showing an example of the overall configuration of a virtualization platform to which a system according to an embodiment can be applied. In FIG. 1, the virtualization platform 10 includes racks 110, 120, and 130 as multiple platform-constituting devices. The racks 110, 120, and 130 each include physical servers (also referred to as "physical hosts") 111, 121, and 131 as multiple physical resources. In the example of FIG. 1, the multiple physical servers 111, 121, and 131 (24 in the illustrated example) included in the racks 110, 120, and 130 form a cluster, which is a unit of a group in which workloads such as virtual machines (VMs) or containers can operate.

[0026] 1, the virtualization infrastructure 10 is configured with three racks 110, 120, and 130, but the virtualization infrastructure 10 may be configured with two racks or four or more racks. Furthermore, the racks 110, 120, and 130 each have eight physical servers 111, 121, and 131, respectively, but the number of physical servers each of the racks 110, 120, and 130 may be one to seven, or nine or more. Furthermore, the number of physical servers each of the racks 110, 120, and 130 may differ from one another.

[0027] A cluster capable of running multiple workloads is constructed on the virtualization platform 10 using multiple racks 110, 120, and 130, and workloads such as virtual machines (VMs) or containers run on multiple physical servers 111, 121, and 131. Each of the multiple racks 110, 120, and 130 also has other devices such as a switch 112. Each of the multiple racks 110, 120, and 130 may have a power supply that supplies power to the multiple physical servers 111, 121, and 131 for each rack.

[0028] In the virtualization platform 10 configured as described above, if a failure occurs in a single physical server 111' in one rack 110 (in the example of FIG. 2, the physical server at the bottom of the rack 110) as shown in FIG. 2, the workload that was running on the failed physical server 111' can be restarted on another physical server in the cluster. Therefore, it is sufficient for the entire cluster to have surplus resources equivalent to one physical server.

[0029] However, rack-scale failures may occur in the virtualization platform 10 configured as described above. For example, as shown in FIG. 3, if a failure occurs in the power supply system of one rack 110, all physical servers 111 in the rack 110 go down. When a rack-scale failure occurs, there is no destination to which the large number of workloads running on all physical servers 111' in the rack 110 can be moved, and these workloads cannot be started, thereby affecting the services provided. While it is possible to start all workloads running on all physical servers 111' in the rack 110 on physical servers in other racks, this would require a huge amount of surplus resources and would be impractical. For example, the entire cluster requires surplus resources equivalent to one rack. For this reason, rack failures are generally not included in failure cases.

[0030] The most likely scenario for a rack failure is a power outage or power system failure. Because the power systems of typical data centers are highly stable due to multiple power lines, the installation of uninterruptible power supplies (UPSs), and in-house power generation, power system failures are often considered negligible. However, with the recent emergence of container-type data centers, compact data centers have become practical, and power system failures are no longer negligible. Even in the event of a power system failure, it is desirable to continue providing services as much as possible (by restarting the necessary workloads on the virtualization platform). While one option for preparing for such cases is to provide redundant workloads using load balancers, a single configuration without redundancy offers advantages, such as the ability to provide services with minimal resources and simplified data sharing.

[0031] In this embodiment, multiple services provided by a cluster constructed on the virtualization infrastructure 10 in which multiple workloads can run are classified into multiple service groups with different priorities, and for each of the multiple service groups, the placement of multiple workloads executed in the cluster and management of the physical servers (physical resources) 111, 121, and 131 required for the service group are performed based on information on the physical servers (physical resources) 111, 121, and 131 of the entire cluster and information on the physical servers (physical resources) 111, 121, and 131 of multiple racks (infrastructure components) 110, 120, and 130. This workload placement and management of the physical servers (physical resources) for each of the multiple service groups makes it possible to continue providing high-priority services on the virtualization infrastructure 10 even if a failure occurs in one of the multiple racks 110, 120, and 130 that constitute the virtualization infrastructure 10 and all physical servers in the failed rack stop operating.

[0032] FIG. 4 is an explanatory diagram showing an example of the arrangement of workloads with multiple priorities in the virtualization platform 10 according to the embodiment. In FIG. 4, the workloads running on the physical servers 111(1) to 111(3), 121(1) to 121(3), and 131(1) to 131(3) of the racks 110, 120, and 130 are shown as square blocks 201A to 201C, 202A to 202C, and 203A to 203C. The symbols A, B, and C in each workload block indicate the priority of each service group. Priority A is the highest priority, Priority B is the second highest priority, and Priority C is the lowest priority. In the example of FIG. 4, four workloads can be started per physical server 111(1) to 111(3), 121(1) to 121(3), and 131(1) to 131(3).

[0033] In the virtualization platform 10 shown in Figure 4, the upper limit for the number of workloads in a cluster is 36, and there are currently 30 workloads running. Even if any single physical server fails, there are surplus resources that allow up to one physical server to be restarted on another physical server. However, if a rack 110 experiences a rack-wide failure, the remaining racks 120 and 130 do not have the capacity to start all 10 workloads 201A to 201C running on physical servers 111(1) to 111(3) of rack 110.

[0034] 5 is an explanatory diagram showing an example of the allocation of workloads of multiple priorities when a rack-wide failure occurs in the virtualization platform 10 according to the embodiment. In FIG. 5, when a failure occurs in the rack 110, four workloads 201A in the service group with the highest priority A can be restarted using the available resources of other physical servers 121 and 131. However, for three workloads 201B in the service group with the second highest priority B that were running in the rack 110, there are insufficient available resources to restart them on the other physical servers 121 and 131.

[0035] Therefore, in this embodiment, when a failure occurs in one of multiple racks (infrastructure configuration devices), the priority of the service groups to which each of the multiple workloads running on the rack (infrastructure configuration device) 110 where the failure occurred belongs is confirmed. Then, if it is confirmed that a workload of a service group with a high priority (e.g., "medium" priority) or higher is running on the rack (infrastructure configuration device) 110 where the failure occurred, the workload of a service group with a low priority (e.g., "low" priority) running on the other racks (infrastructure configuration devices) 120, 130 where the failure did not occur is stopped. Furthermore, the workloads of the service groups with high priorities A and B that were running on the rack (infrastructure configuration device) 110 where the failure occurred are restarted using the physical resources freed up by stopping the workloads on the other racks (infrastructure configuration devices) 120, 130 where the failure did not occur.

[0036] In the example of Figure 5, some of the workloads 202C and 203C of the service group with low priority C that were running on physical servers 121 and 131 in racks 120 and 130 are stopped, and the resources freed up by stopping these workloads are used to restart three workloads 201B of the service group with priority B that were running on rack 110.

[0037] 5 , depending on the resource status of each rack, all of the workloads of the service group with priority C that were running on the physical servers 121 and 131 of the racks 120 and 130 may be stopped, and the resources freed up by stopping the workloads may be used to restart the workload of the service group with priority B that was running on the rack 110. Also, depending on the resource status of each rack, all of the workloads of the service group with priority C and some or all of the workloads of the service group with priority B that were running on the physical servers 121 and 131 of the racks 120 and 130 may be stopped, and the resources freed up by stopping the workloads may be used to restart workload 201A of the service group with priority A, workload 201B of the service group with priority B, or both of these workloads that were running on the rack 110.

[0038] As described above, in this embodiment, as a countermeasure against rack failures, by accepting the abandonment of services of low priority service groups, it becomes possible to reduce excess resources in the virtualization infrastructure 10.

[0039] 6 is an explanatory diagram showing an example of the configuration of a management system 30 that manages the virtualization infrastructure 10 according to the embodiment. In FIG. 6, the management system 30 includes an information storage unit (DB) 310, an information processing unit 320, and a control unit 330. The management system 30 may further include a communication unit 340. The management system 30 may be installed on a server or cloud computer system separate from the virtualization infrastructure 10, or may be installed as a management server on a physical server in the virtualization infrastructure 10.

[0040] The information storage unit (DB) 310 stores information on all physical servers (physical resources) 111, 121, and 131 of a cluster constructed on the virtualization platform 10 and capable of running multiple workloads, as well as information on the physical servers (physical resources) 111, 121, and 131 of each of multiple racks (platform configuration devices) 110, 120, and 130. The information storage unit (DB) 310 also stores information on various tables in the configuration of the virtualization platform 10 described above, such as the service group table in Table 1, the priority-based resource table in Table 2, the rack resource table in Table 3, the cluster resource table in Table 4, and the condition table in Table 5.

[0041]

[0042]

[0043]

[0044]

[0045]

[0046] In the example condition table of Table 5, the following conditions are shown: four virtual CPUs can be assigned to one physical CPU installed in the physical server for use; memory can be used up to the upper limit of the physical memory installed in the physical server; and two virtual GPUs can be assigned to one physical GPU installed in the physical server for use.

[0047] The information processing unit 320 classifies multiple services provided by a cluster that is constructed on the virtualization infrastructure 10 and that can run multiple workloads into multiple service groups with different priorities. Furthermore, the information processing unit 320 allocates multiple workloads to be executed in the cluster and manages the physical servers (physical resources) required for each of the multiple service groups, based on information on the physical servers (physical resources) 111, 121, and 131 of the entire cluster and information on the physical servers (physical resources) 111, 121, and 131 of each of the multiple racks (infrastructure components) 110, 120, and 130.

[0048] As described above, when a failure occurs in one of the multiple racks (infrastructure configuration devices) 110, 120, and 130, the control unit 330 checks the priority of the service groups to which each of the multiple workloads running on the failed rack (infrastructure configuration device) 110 belongs. Furthermore, if the control unit 330 checks that a workload from a service group with a high priority (e.g., a "medium" priority) or higher was running on the failed rack (infrastructure configuration device) 110, the control unit 330 stops the workloads of a service group with a low priority (e.g., a "low" priority) running on the other racks (infrastructure configuration devices) 120 and 130 where the failure did not occur. Furthermore, the control unit 330 controls the restart of the workloads of the high-priority service groups A and B running on the failed rack (infrastructure configuration device) 110 using the physical resources freed up by stopping the workloads on the other racks (infrastructure configuration devices) 120 and 130 where the failure did not occur.

[0049] The communication unit 340 communicates with the terminal device of the user (or operator, etc.) 40 via a communication line, etc., and provides the user with impact information in the event of a failure in the rack (infrastructure component device) 110 with the greatest physical resources among multiple racks (infrastructure component devices) 110, 120, and 130.

[0050] 7 is a flowchart showing an example of generating a resource-related table for a virtual environment in the virtualization platform 10 according to the embodiment. In FIG. 7, the information processing unit 320 generates a workload table related to a workload when deploying (e.g., constructing or creating) the workload in the virtualization platform 10 (S101). Next, the information processing unit 320 compiles resource components (CPU, memory, GPU) for each service group with different priorities to create a service group table as shown in Table 1 (S102). Next, the information processing unit 320 compiles resource components (CPU, memory, GPU) for each priority level for the service group table in Table 1 to create a priority-based resource table as shown in Table 3 (S103).

[0051] 8 is a flowchart showing an example of generating a physical resource-related table in the virtualization platform according to the embodiment. In FIG. 8, the information processing unit 320 generates a physical server resource table from information about the physical servers in each rack that constitutes a cluster in the virtualization platform 10 (S201). Next, the information processing unit 320 aggregates the resource components (CPU, memory, GPU) in the physical server resource table to generate a cluster resource table as shown in Table 4 (S202). Next, the information processing unit 320 generates a rack resource table as shown in Table 3 from the physical server resource table (S203).

[0052] 9 is a flowchart showing an example of calculation of priorities of service groups that restart workloads when a rack-wide failure occurs in the virtualization platform 10 according to the embodiment. In Fig. 9, the information processing unit 320 calculates the physical resources that can be used for each resource component (CPU, memory, GPU) when the rack 110 with the largest number of resource components (CPU, memory, GPU) goes down (fails), i.e., when a failure occurs in the rack 110, based on information on the physical servers 111, 121, and 131 of the multiple racks 110, 120, and 130 (S301). Next, the information processing unit 320 calculates and determines how many priorities can be covered in descending order of priority, i.e., the predetermined priority, based on the calculation results for each resource component (CPU, memory, GPU) of the physical resources of the available racks 120 and 130 (rack resource table in Table 7), information for each resource component (CPU, memory, GPU) of the physical resources corresponding to each of multiple priorities (high, medium, low) (resource table by priority in Table 6), and the competitive allocation setting conditions for each resource component (CPU, memory, GPU) in Table 5 (S302).

[0053]

[0054]

[0055] Referring to Table 7, it can be seen that if rack 110 (rack ID=1) goes down (fails), the number of CPUs available in the entire virtualization platform 10 will be 16 (64 cores), memory will be 256 GB, and GPUs will be 8 (16 cores). Also, referring to the row in Table 7 that takes into account the conditions calculated in consideration of the conflict allocation setting conditions for each resource component (CPU, memory, GPU) in Table 5,

[0056] Next, the information processing unit 320 transmits and provides impact information regarding the impact that would occur if the rack 110 with the largest resource components (CPU, memory, GPU) were to go down (fail), i.e., if a failure occurs in the rack 110, to the terminal device of the user (or operator, etc.) 40 via the communication unit 340 (S303).

[0057] In the virtualization infrastructure 10 of this embodiment, in normal times, if all of the racks 110, 120, and 130 are operating normally, the workloads of all service groups with high, medium, and low priorities can be run.

[0058] Furthermore, in this embodiment, in a pre-calculation performed in preparation for a rack failure in which one of the racks goes down (falls), it can be determined that the impact of a rack failure in which the rack 110 goes down (falls) will be large because the resource components (CPU, memory, GPU) of the rack 110 are the largest among the three racks 110, 120, and 130. If the rack 110 goes down (falls), the number of CPU cores in the high and medium priority service groups will be 60, which will not be able to cover the workload of the low priority service group.

[0059] Furthermore, since the above pre-calculation shows that the workloads of low-priority service groups cannot be covered, when a rack failure occurs in which rack 110 goes down (falls), the workloads of low-priority service groups are stopped on physical servers 121 and 131 of other racks 120 and 130, and the workloads of high-priority and medium-priority service groups that were running on rack 110 before the failure are restarted using the resources of physical servers 121 and 131 that have become available on racks 120 and 130.

[0060] In this embodiment, all or part of the generation of the virtual environment resource-related table in Figure 7, the generation of the physical resource-related table in Figure 8, and the calculation of the service group priority in Figure 9 may be pre-calculated before the actual operation of the virtualization infrastructure 10 begins, or may be recalculated when the number or configuration of racks in the virtualization infrastructure 10 changes, when the workload of virtual machines (VMs) or containers running on the virtualization infrastructure 10 changes, when the services provided by the virtualization infrastructure 10 change, etc.

[0061] As described above, according to this embodiment, even if a failure occurs in one of the multiple racks 110, 120, and 130 that make up the virtualization infrastructure 10 and all physical servers 111 in the failed rack stop operating, high-priority services can be continuously provided on the virtualization infrastructure 10.

[0062] Furthermore, according to this embodiment, excess resources of physical servers in a rack in the virtualization platform 10 can be reduced.

[0063] Note that in this embodiment, the case where the multiple infrastructure-constituting devices constituting the virtualization infrastructure 10 are racks 110, 120, and 130 each having multiple physical servers and power supplies, and the physical resources are physical servers 111, 121, and 131, has been mainly described; however, the types of the multiple infrastructure-constituting devices and physical resources constituting the virtualization infrastructure 10 are not limited to these. For example, the multiple infrastructure-constituting devices constituting the virtualization infrastructure 10 may each be a container-type data center having multiple physical servers and power supplies, and the physical resources may be racks having multiple physical servers. In this case, even if a failure occurs in one of the multiple container-type data centers constituting the virtualization infrastructure 10 and all physical servers in the container-type data center where the failure occurred stop operating, it is possible to continue providing high-priority services on the virtualization infrastructure 10.

[0064] Furthermore, according to this embodiment, surplus resources in the container-type data center in the virtualization platform 10 can be suppressed.

[0065] The present invention can reduce excess resources in the virtualization infrastructure 10, while allowing high-priority services to continue to be provided on the virtualization infrastructure 10 even if a failure occurs in any of the multiple racks 110, 120, and 130 that make up the virtualization infrastructure 10, thereby contributing to the achievement of Goal 9 of the Sustainable Development Goals (SDGs), which is to "build resilient infrastructure, promote inclusive and sustainable industrialization and foster innovation."

[0066] The processing steps and components of the management system (storage unit, information processing unit, control unit, communication unit) described herein can be implemented by various means. For example, these steps and components may be implemented by hardware, firmware, software, or a combination thereof.

[0067] For hardware implementation, the means, such as processing units, used to implement the above steps and components in an entity (e.g., a physical server, rack, container-type data center, hard disk drive device, or optical disk drive device) may be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, electronic devices, other electronic units designed to perform the functions described herein, computers, or combinations thereof.

[0068] Furthermore, with regard to firmware and / or software implementations, the means, such as a processing unit, used to realize the components may be implemented with a program (e.g., code, such as procedures, functions, modules, instructions, etc.) that performs the functions described herein. In general, any computer / processor-readable medium tangibly embodying firmware and / or software code may be used to implement the means, such as a processing unit, used to realize the steps and components described herein. For example, the firmware and / or software code may be stored in a memory and executed by a computer or processor, such as in a controller. The memory may be implemented within the computer or processor, or external to the processor. The firmware and / or software code may also be stored on a computer or processor readable medium such as, for example, random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), flash memory, floppy disk, compact disk (CD), digital versatile disk (DVD), magnetic or optical data storage device, etc. The code may be executed by one or more computers or processors and may cause the computers or processors to perform certain aspects of the functionality described herein.

[0069] The medium may be a non-transitory recording medium. The program code may be in any format as long as it can be read and executed by a computer, processor, or other device or machine. For example, the program code may be in any of source code, object code, and binary code, or may be a mixture of two or more of these codes.

[0070] Moreover, the description of the embodiments disclosed herein is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to the present disclosure will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other variations without departing from the spirit or scope of the present disclosure. Thus, the present disclosure is not intended to be limited to the examples and designs described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0071] 10: Virtualization platform 30: Management system 110: Rack 111: Physical server 111(1) to 111(3): Physical server 111': Physical server 112: Switch 120: Rack 121: Physical server 121(1) to 121(3): Physical server 130: Rack 131: Physical server 131(1) to 131(3): Physical server 201A to 201C: Workloads 202A to 202C: Workloads 203A to 203C: Workloads 320: Information processing unit 330: Control unit 340: Communication unit

Claims

1. A system for managing a virtualization infrastructure having multiple infrastructure component devices each having multiple physical resources, comprising: an information storage unit that stores information on the overall physical resources of a cluster on which multiple workloads constructed on the virtualization infrastructure can operate, and information on the physical resources of each of the multiple infrastructure component devices; and an information processing unit that classifies multiple services provided by the cluster on which multiple workloads constructed on the virtualization infrastructure can operate into multiple service groups with different priorities, and manages the allocation of multiple workloads to be executed on the cluster and the physical resources required by each of the multiple service groups based on the information on the overall physical resources of the cluster and the information on the physical resources of each of the multiple infrastructure component devices.

2. A system according to claim 1, characterized in that it comprises a control unit that, when a failure occurs in one of the plurality of infrastructure component devices, checks the priority of the service groups to which each of the plurality of workloads running on the infrastructure component device where the failure belongs, and if it is confirmed that a workload of a service group with a priority equal to or higher than a predetermined priority was running on the infrastructure component device where the failure occurred, stops the workload of the low-priority service group running on the infrastructure component device where the failure did not occur, and controls the system to restart the workload of the high-priority service group that was running on the infrastructure component device where the failure occurred using the physical resources freed up by the stopping of the workload on the infrastructure component device where the failure did not occur.

3. A system according to claim 2, wherein the information processing unit calculates, for each resource component, the physical resources that can be used in the event of a failure in the infrastructure component device with the largest physical resources among the plurality of infrastructure component devices, based on information on the physical resources of each of the plurality of infrastructure component devices; and determines the predetermined priority based on the calculation results for each resource component of the available physical resources, information for each resource component of the physical resources corresponding to each of the plurality of priorities, and the competitive allocation setting conditions for each resource component.

4. A system according to claim 3, wherein the information processing unit provides users with impact information in the event that a failure occurs in the infrastructure component device with the greatest physical resources among the plurality of infrastructure components.

5. A system according to any one of claims 1 to 4, wherein the physical resource is a physical server, and the infrastructure component device is a rack having a plurality of physical servers and a power supply.

6. A system according to any one of claims 1 to 4, wherein the physical resource is a rack having a plurality of physical servers, and the infrastructure component device is a container-type data center having a plurality of racks and a power source.

7. A method for managing a virtualization infrastructure having a plurality of infrastructure component devices each having a plurality of physical resources, comprising: storing information about the overall physical resources of a cluster on which a plurality of workloads constructed on the virtualization infrastructure can operate, and information about the physical resources of each of the plurality of infrastructure component devices; classifying a plurality of services provided by the cluster on which a plurality of workloads constructed on the virtualization infrastructure can operate into a plurality of service groups having different priorities; and, for each of the plurality of service groups, allocating a plurality of workloads to be executed on the cluster and managing the physical resources required by the service group based on the information about the overall physical resources of the cluster and the information about the physical resources of each of the plurality of infrastructure component devices.

8. The method of claim 7, comprising, when a failure occurs in one of the plurality of infrastructure component devices, checking the priority of the service groups to which each of the plurality of workloads running on the failed infrastructure component device belongs; if it is confirmed that a workload of a service group with a priority equal to or higher than a predetermined priority was running on the failed infrastructure component device, stopping the workload of a low priority service group running on the non-failed infrastructure component device; and restarting the workload of the high priority service group running on the failed infrastructure component device using the physical resources freed up by stopping the workload on the non-failed infrastructure component device.

9. A program executed by a computer or processor provided in a system that manages a virtualization infrastructure having multiple infrastructure component devices each having multiple physical resources, comprising: program code for storing information on the overall physical resources of a cluster on which multiple workloads constructed on the virtualization infrastructure can operate, and information on the physical resources of each of the multiple infrastructure component devices; program code for classifying multiple services provided by the cluster on which multiple workloads constructed on the virtualization infrastructure can operate, into multiple service groups with different priorities; and program code for allocating multiple workloads to be executed in the cluster and managing the physical resources required by the service group for each of the multiple service groups, based on the information on the overall physical resources of the cluster and the information on the physical resources of each of the multiple infrastructure component devices.

10. A program according to claim 9, comprising: program code for, when a failure occurs in one of the plurality of infrastructure component devices, checking the priority of the service groups to which each of the plurality of workloads running on the infrastructure component device where the failure occurs belongs; program code for, when it is confirmed that a workload of a service group with a priority equal to or higher than a predetermined priority was running on the infrastructure component device where the failure occurred, stopping the workload of a low priority service group running on the infrastructure component device where the failure did not occur; and program code for restarting the workload of the high priority service group that was running on the infrastructure component device where the failure occurred using physical resources freed up by stopping the workload on the infrastructure component device where the failure did not occur.

Citation Information

Patent Citations

  • Power saving system and power saving method

    WO2011043317A1

  • Information processing device, information processing system, and program

    JP2019087033A

  • Secure enclosure systems in a provider network

    US10121026B1