Resource management device, and resource management method
The resource management device and method address the issue of unaccounted error budget consumption by predicting and adjusting resource allocation between subsystems, ensuring stable IT system operation by managing computing resources proactively.
Patent Information
- Application Number
- JP2024079530
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-15
- Publication Date
- 2025-11-28
AI Technical Summary
Existing methods for managing computing resources in IT systems do not adequately account for error budget consumption during resource allocation, leading to potential reliability losses.
A resource management device and method that predicts resource shortages and adjusts resource allocation between subsystems based on error budget changes, allowing for proactive management of computing resources to prevent reliability loss.
Enables effective management of computing resources in anticipation of error budget consumption, ensuring stable operation by sharing resources between subsystems.
Smart Images

Figure 2025173777000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to techniques for managing computing resources in information systems. [Background technology]
[0002] A technology called autoscaling is used in IT (Information Technology) systems. Autoscaling is a technology that automatically increases or decreases computing resources such as VMs (Virtual Machines) according to factors such as the computing load. By appropriately controlling the amount of computing resources, it is possible to reduce the cost required for computing resources while ensuring performance such as the reliability of the IT system. In autoscaling, increasing computing resources is called "scaling out," and decreasing computing resources is called "scaling in."
[0003] Patent Document 1 discloses a method for appropriately controlling the amount of allocated computing resources. The method in Patent Document 1 calculates the dependency of the allocated amount of computing resources on performance, and increases or decreases only the amount of dependent computing resources. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Publication No. 2020-123849 Summary of the Invention [Problem to be solved by the invention]
[0005] In IT systems, the amount of error tolerance is determined by an error budget. The error budget is an index that indicates the degree of acceptable reliability loss. Examples of error budget indexes include the amount of time that can be down, the percentage of response errors that can occur, and the percentage of delayed responses that can occur. However, the method in Patent Document 1 does not provide control that ensures computational resources while taking into account cases where the error budget is expected to be consumed due to a shortage of computational resources.
[0006] One objective of the present disclosure is to provide a technology that enables appropriate management of computing resources in anticipation of error budget consumption. [Means for solving the problem]
[0007] A resource management device according to one aspect included in the present disclosure is a resource management device that manages computational resources of an information system that includes a plurality of subsystems, and includes: a resource shortage prospect detection unit that predicts, for each of the plurality of subsystems, a future required resource quantity that indicates the quantity of computational resources that the subsystem will require; if the future required resource quantity exceeds a current resource quantity, which is the quantity of computational resources currently allocated to the subsystem, the subsystem is designated as a shortage subsystem and detects an expected resource shortage of the shortage subsystem; an error budget change estimation unit that estimates a change in the remaining amount of an error budget, which is an index indicating the degree of tolerable reliability loss in a non-deficit subsystem, which is a subsystem for which an expected resource shortage has not been detected, when computational resources equal to the expected resource shortage quantity, which is the difference between the future required resource quantity and the current resource quantity, are removed from the non-deficit subsystem; a resource accommodation method identification unit that identifies an accommodation subsystem, which is a non-deficit subsystem that should transfer computational resources to the shortage subsystem, based on the change in the remaining amount of the error budget of each of the non-deficit subsystems; and a resource accommodation control unit that changes the allocation of computational resources from the accommodation subsystem to the shortage subsystem by the expected resource shortage quantity.
[0008] A resource management method according to one aspect included in the present disclosure is a resource management method in which a resource management device manages the computational resources of an information system including a plurality of subsystems, the method comprising: predicting, for each of the plurality of subsystems, a future required resource quantity indicating the quantity of computational resources that the subsystem will require; if the future required resource quantity exceeds a current resource quantity, which is the quantity of computational resources currently allocated to the subsystem, designating the subsystem as a shortage subsystem; detecting an expected resource shortage for the shortage subsystem; and calculating a change in the remaining error budget, which is an index indicating the degree of tolerable reliability loss, in each of the non-deficit subsystems, which are subsystems for which an expected resource shortage has not been detected, if computational resources equal to the expected resource shortage quantity, which is the difference between the future required resource quantity and the current resource quantity, are removed; identifying a flexible subsystem, which is a non-deficit subsystem that should transfer computational resources to the deficit subsystem, based on the change in the remaining error budget for each of the non-deficit subsystems; and changing the allocation of computational resources from the flexible subsystem to the deficit subsystem by the expected resource shortage quantity. [Effects of the Invention]
[0009] According to one aspect of the present disclosure, it is possible to appropriately manage computing resources in anticipation of error budget consumption. [Brief explanation of the drawings]
[0010] [Figure 1] 1 is a block diagram showing the overall configuration of the present embodiment; [Figure 2] FIG. 2 is a block diagram showing a hardware configuration of a management server. [Figure 3] FIG. 2 is a block diagram showing the functional configuration of an IT system and a management server. [Figure 4] FIG. 1 is a conceptual diagram illustrating the successful addition of computing resources to an IT system. [Figure 5]FIG. 1 is a conceptual diagram illustrating a situation in which adding computing resources to an IT system fails due to a lack of resources. [Figure 6] This is a conceptual diagram showing how computing resources are shared between subsystems of an IT system with the involvement of a management server. [Figure 7] FIG. 10 is a diagram illustrating an example of an additional resource information temporary management table. [Figure 8] FIG. 10 is a diagram illustrating an example of a logical resource specification information temporary management table. [Figure 9] FIG. 10 illustrates an example of an error budget management table. [Figure 10] FIG. 10 is a diagram illustrating an example of an error budget characteristic management table. [Figure 11] FIG. 10 is a diagram illustrating an example of a temporary management table for error budget change trial calculation results. [Figure 12] FIG. 10 is a diagram illustrating an example of a credit standard management table. [Figure 13] FIG. 10 is a diagram illustrating an example of a resource accommodation method temporary management table. [Figure 14] 10 is a flowchart illustrating an example of a resource accommodation process. [Figure 15] FIG. 10 illustrates an example of a GUI of a management server. DETAILED DESCRIPTION OF THE INVENTION
[0011] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.
[0012] FIG. 1 is a block diagram showing the overall configuration of this embodiment.
[0013] An IT system 101, a management server 103, and a console 104 are interconnected so as to be able to communicate with each other via a network 102 such as the Internet. The IT system 101 corresponds to an information system.
[0014] The IT system 101 may be configured with multiple computers. The IT system 101 includes one or more processors (not shown) and one or more memories (not shown). The processors included in the IT system 101 read and execute programs stored in the memories to functionally realize a computing resource group 111, a monitoring unit 112, an error budget management unit 113, a computing resource management unit 114, and a logical resource management unit 115.
[0015] The IT system 101 has one or more computational resource groups. One computational resource group 111 is depicted in Fig. 1. The computational resource group has one or more computational resources. The computational resource group 111 in Fig. 1 has three computational resources, computational resources 121a to 121c. The computational resources include multiple logical resources.
[0016] The monitoring unit 112 monitors the usage status of the computing resources included in the computing resource group.
[0017] The error budget management unit 113 manages the error budget for each computational resource or logical resource in the computational resource group. The error budget is an index indicating the degree to which the reliability of the service can be allowed to be impaired.
[0018] The computational resource manager 114 manages the computational resources in the computational resource group.
[0019] The logical resource manager 115 manages the logical resources included in the computational resources.
[0020] An example configuration of the management server 103 will be described later. The console 104 includes a general input device and an output device. The input device may be, for example, a keyboard, a mouse, or a touch panel. The output device may be, for example, a monitor or a touch panel, and outputs various types of information in the management server 103. A user operates the input device to input information into the IT system while checking the information displayed on the console's output device.
[0021] FIG. 2 is a block diagram showing the hardware configuration of the management server.
[0022] The management server 103 includes a processor 201 , a memory 202 , an auxiliary storage device 203 , a communication interface 204 , a media interface 205 , and an input / output device 206 .
[0023] The processor 201 is configured by, for example, a CPU (Central Processing Unit), an MPU (Micro Processing Unit), a GPU (Graphics Processing Unit), an FPGA (Field-Programmable Gate Array), etc. The processor 201 realizes various functions of the management server 103 by reading and executing programs stored in the memory 202 or the auxiliary storage device 203.
[0024] The auxiliary storage device 203 may be, for example, a hard disk or an SSD. The auxiliary storage device 203 stores the programs executed by the processor 201 and various types of data. An external storage medium 207 may be disposed outside the management server 103. The external storage medium 207 may be communicably connected to the management server 103. The external storage medium 207 stores the programs executed by the processor 201 and various types of data.
[0025] The communication interface 204 communicates information between the management server 103 and various external devices via the network 102 .
[0026] The media interface 205 communicates information with various media such as CDs, USB memories, SD cards, DVDs, and Blu-rays.
[0027] The input / output device 206 inputs information to the management server 103 and outputs information from the management server 103 to an external device. For example, the input / output device 206 may output information to the console 104.
[0028] The external storage medium 207 is a storage medium located outside the management server 103 and is communicably connected to the management server. The external storage medium 207 may be located on a cloud, for example. The external storage medium 207 may also be located on the same premises as the management server 103.
[0029] FIG. 3 is a block diagram showing the functional configuration of the IT system and the management server.
[0030] Of the configuration shown in Fig. 3, the description of the same configuration as in Fig. 1 or 2 will be omitted. The IT system 101 includes multiple subsystems. In the figure, three subsystems, subsystems 301a to 301c, are illustrated as an example. Each subsystem secures a logical resource. For example, subsystem 301a secures three logical resources, logical resources 311a to 311c. One logical resource may be secured across multiple computing resources.
[0031] Error budget management unit 113 has error budget management table 321. Error budget management table 321 will be described later with reference to FIG.
[0032] The CPU 201 of the management server 103 reads and executes the programs stored in the storage device, thereby functionally realizing an IT system cooperation unit 331, a resource accommodation management unit 332, and a management input / output unit 333. The storage device may be any of the memory 202, the auxiliary storage device 203, and the external storage medium 207.
[0033] The IT system cooperation unit 331 has a function of communicating information between the management server 103 and the IT system 101. The resource accommodation management unit 332 will be described later. The management input / output unit 333 has a function of inputting and outputting management information between the management server 103 and a user. The user here refers to, for example, the administrator of the management server 103. The management input / output unit 333 has a setting unit 361, a notification unit 362, and a display unit 363.
[0034] The setting unit 361 has a function of setting various setting values in the management server 103 in response to user input. The notification unit 362 has a function of notifying the user of information. The display unit 363 has a function of displaying information.
[0035] Next, a description will be given of the configuration of the resource accommodation management unit 332. The resource accommodation management unit 332 has a resource shortage prospect detection unit 341, an error budget change estimation unit 343, a resource accommodation method identification unit 347, and a resource accommodation control unit 350.
[0036] The resource shortage prospect detection unit 341 has the function of predicting the future required resource quantity, which indicates the quantity of computing resources that will be required by each of multiple subsystems, and if the future required resource quantity exceeds the current resource quantity, which is the quantity of computing resources currently allocated to the subsystem, designating the subsystem as a deficient subsystem and detecting the prospect of a resource shortage for the deficient subsystem.
[0037] When an attempt is made to add computing resources to a subsystem, the resource shortage likelihood detection unit 341 may monitor the status of the addition of computing resources and determine whether or not a resource shortage is likely based on the status.
[0038] If the addition of computing resources to a subsystem fails, the resource shortage potential detection unit 341 may determine that the subsystem is likely to experience a resource shortage.
[0039] If the addition of computing resources to a subsystem is not completed in time for when it is needed, the resource shortage likelihood detection unit 341 may determine that the subsystem is likely to experience a resource shortage.
[0040] If the current required resource quantity, which indicates the number of computing resources required by a subsystem, exceeds the current resource quantity before the addition of computing resources to the subsystem is completed, the resource shortage prospect detection unit 341 may determine that the addition of computing resources to the subsystem will not be in time for when they are needed.
[0041] The resource shortage probability detection unit 341 may predict changes in the error budget for multiple subsystems, and determine that a subsystem that is expected to consume the error budget is likely to experience a resource shortage.
[0042] The IT system 101 may be a cloud-based system that has a function for increasing or decreasing computing resources by auto-scaling. In this case, the resource shortage prediction detection unit 341 may monitor the status of auto-scaling when auto-scaling occurs, and determine whether or not a resource shortage is predicted based on the status.
[0043] The error budget change estimation unit 343 has a function of estimating the change in the remaining error budget, which is an index showing the degree of allowable reliability loss in each non-deficient subsystem, that is, a subsystem in which a resource shortage has not been detected, when calculation resources for the expected resource shortage quantity, which is the difference between the future required resource quantity and the current resource quantity, are deleted from the non-deficient subsystem.
[0044] The resource accommodation method specifying unit 347 has a function of specifying an accommodation subsystem that is a non-deficient subsystem that should transfer computing resources to a deficit subsystem, based on changes in the remaining amount of the error budget of each non-deficient subsystem.
[0045] The resource accommodation method specifying unit 347 may specify an accommodation subsystem based on the remaining amount of the error budget of each non-deficient subsystem or the amount of change in the remaining amount.
[0046] The resource accommodation control unit 350 has a function of changing the allocation of computing resources from the accommodation subsystem to the shortage subsystem by the estimated amount of resources that are in short supply.
[0047] When the resource accommodation control unit 350 changes the allocation of computing resources from the accommodation subsystem to the shortage subsystem, the resource accommodation control unit 350 may present information that an accommodation of computing resources has occurred between the subsystems. The presentation here includes outputting or displaying the information.
[0048] The storage device stores an additional resource information temporary management table 342, a logical resource specification information temporary management table 344, an error budget characteristic management table 345, an error budget change trial calculation result temporary management table 346, an accommodation standard management table 348, and a resource accommodation method temporary management table 349. The storage device may be the memory 202, the auxiliary storage device 203, or the external storage medium 207.
[0049] FIG. 4 is a conceptual diagram showing how a computing resource is successfully added to an IT system.
[0050] The logical resource manager 115 detects an event that indicates the need to add a logical resource for the subsystem 301a in the computing resource group 111. The computing resource manager 114 receives a notification of the above event from the logical resource manager 115 and determines that the addition of a logical resource is necessary. The computing resource manager 114 performs a process to add a new computing resource 121d to the computing resource group 111. After the addition of the computing resource 121d is successful, the logical resource manager 115 adds a logical resource 311d for the subsystem 301a.
[0051] FIG. 5 is a conceptual diagram showing how adding computing resources to an IT system fails due to a lack of resources.
[0052] The logical resource manager 115 detects an event that indicates the need to add a logical resource to the subsystem 301a in the computing resource group 111. The computing resource manager 114, which has received a notification of the above event from the logical resource manager 115, determines that the addition of a logical resource is necessary. The computing resource manager 114 performs processing to add a new computing resource 121d to the computing resource group 111. However, for some reason, the addition of the computing resource 121d fails. As a result, the addition of the logical resource 311d to the subsystem 301a by the logical resource manager 115 also fails.
[0053] The resource management device of the present disclosure accommodates computing resources between subsystems, for example, in the case where the addition of computing resources fails as described above. The resource management device is, for example, the management server 103.
[0054] FIG. 6 is a conceptual diagram showing how computing resources are shared between subsystems of an IT system with the involvement of a management server.
[0055] For example, in the situation shown in FIG. 5, the addition of the computing resource 121d failed.
[0056] The management server 103, which is a resource management device, is a device that manages the computational resources of the IT system 101, which includes multiple subsystems.
[0057] The resource shortage prospect detection unit 341 predicts the future required resource quantity, which indicates the quantity of computing resources that the subsystem will require, for each of multiple subsystems. If the future required resource quantity exceeds the current resource quantity, which is the quantity of computing resources currently allocated to the subsystem, the subsystem is designated as a deficient subsystem and a resource shortage prospect for the deficient subsystem is detected. In the example of Figure 6, the deficient subsystem is subsystem 301a, which failed to allocate logical resource 311d. The resource shortage prospect is the computing resource that was scheduled to be allocated as logical resource 311d.
[0058] Error budget change estimation unit 343 estimates the change in the remaining error budget, which is an index showing the degree of tolerable reliability loss in each non-deficient subsystem, which is a subsystem for which a resource shortage has not been detected, if computational resources equivalent to the estimated resource shortage quantity, which is the difference between the future required resource quantity and the current resource quantity, are deleted from the non-deficient subsystem. In the example of Figure 6, the subsystem with the shortage is subsystem 301a, and the other subsystems 301b and 301c are non-deficient subsystems, and estimates the change in the remaining error budget if computational resources equivalent to the estimated resource shortage quantity are deleted from each non-deficient subsystem, that is, if computational resources equivalent to the estimated resource shortage quantity are lent from the non-deficient subsystem to the deficit subsystem.
[0059] The resource accommodation method identification unit 347 identifies an accommodation subsystem that is a non-deficient subsystem that should transfer computing resources to the deficit subsystem based on the change in the remaining error budget of each non-deficient subsystem. In the example of Figure 6, the resource accommodation method identification unit 347 identifies subsystem 301b, which requires a smaller reduction in the error budget, as the accommodation subsystem.
[0060] The resource accommodation control unit 350 changes the allocation of computing resources from the accommodation subsystem to the shortage subsystem by the estimated resource shortage quantity. In the example of Figure 6, the logical resources of subsystem 301b, which is the accommodation subsystem, are reduced, and the logical resources of subsystem 301a, which is the shortage subsystem, are added by that amount.
[0061] 7 is a diagram showing an example of an additional resource information temporary management table. The additional resource information temporary management table 342 has, as information items, a subsystem 701 and a logical resource specification 702. The logical resource specification 702 includes, for example, a CPU 703, a memory 704, and a number of logical resources 705. In the example of FIG. 7, a subsystem 301a having four CPUs, 8 GB of memory, and one logical resource is registered in the additional resource information temporary management table 342.
[0062] 8 is a diagram showing an example of a logical resource specification information temporary management table. The logical resource specification information temporary management table 344 has information items such as a subsystem 711 and a logical resource specification 712. The logical resource specification 712 includes, for example, a CPU 713, a memory 714, and a number of logical resources 715. In the example of FIG. 8, subsystems 301a to 301c, each having a different number of CPUs, memory capacity, and number of logical resources, are registered in the logical resource specification information temporary management table 344.
[0063] 9 is a diagram showing an example of an error budget management table. The error budget management table 321 has information items of a subsystem 721 and a remaining error budget 722. The remaining error budget 722 includes one or more indexes, such as index 1, index 2, and index 3.
[0064] 10 is a diagram showing an example of an error budget characteristic management table. Error budget characteristic management table 345 has information items such as subsystem 731 and error budget characteristic 732. Error budget characteristic 732 includes characteristic equations corresponding to the indices shown in FIG. 9. For example, characteristic equation 1-1 is registered in error budget characteristic management table 345 for characteristic 1 of subsystem 301a.
[0065] 11 is a diagram showing an example of a temporary management table for error budget change trial calculation results. The temporary management table for error budget change trial calculation results has, as information items, subsystem 741 and error budget change calculation value 743. Error budget change calculation value 743 has error budget reduction amount 745a and error budget remaining amount 746b for each index such as index 1 to index 3.
[0066] 12 is a diagram showing an example of an accommodation standard management table. The accommodation standard management table 348 has, as information items, a subsystem 751 and a resource accommodation standard value calculation formula 752. The resource accommodation standard value may be a value based on, for example, the remaining amount or reduction amount of the error budget.
[0067] 13 is a diagram showing an example of a resource accommodation method temporary management table. The resource accommodation method temporary management table 349 has, as information items, a subsystem 761 and a reduced logical resource count 762. The reduced logical resource count is a value indicating how many logical resources are reduced when computing resources are accommodated from one subsystem to another.
[0068] FIG. 14 is a flowchart illustrating an example of a resource accommodation process.
[0069] The error budget change estimation unit 343 estimates the error budget change when resources are accommodated for each subsystem based on various information, and stores the estimation result in the error budget change estimation result temporary management table 346 (S801).
[0070] The "various information" in step S801 refers to the following information. Necessary additional logical resource information in the additional resource information temporary management table 342 Logical resource specification information for each subsystem in the logical resource specification information temporary management table 344, remaining error budget information, target period, resource usage rate, etc. Characteristic formula of error budget characteristic management table 345
[0071] The resource accommodation method identification unit 347 identifies a resource accommodation implementation method based on the accommodation standard value calculated using the calculation formula in the accommodation standard management table 348 for the error budget change estimated value stored in the error budget change estimated result temporary management table 346, and stores the method in the resource accommodation method temporary management table 349 (S302).
[0072] Resource accommodation is performed. If resource accommodation has been performed (S803: Y), the process proceeds to step S804. If resource accommodation has not been performed (S803: N), the process shown in FIG. 14 ends.
[0073] In step S804, the resource accommodation control unit 350 instructs the IT system 101 to change the logical resource amount of the resource accommodation target subsystem stored in the resource accommodation method temporary management table 349.
[0074] FIG. 15 is a diagram illustrating an example of the GUI of the management server.
[0075] The GUI is displayed on the display unit 363 shown in Fig. 3. The GUI includes a resource accommodation implementation status 901 and a remaining error budget 902.
[0076] Resource accommodation implementation status 901 displays information indicating the implementation status of resource accommodation being carried out at that time. Resource accommodation implementation status 901 includes, for example, an event time, an event type, and a target. The event time means the time when various events such as resource accommodation occurred. The event type means the type of event, such as the implementation of resource accommodation. The target corresponds to a subsystem whose logical resources have decreased or increased due to resource accommodation. Error budget remaining amount 902 displays the remaining error budget for each subsystem.
[0077] In this embodiment, in an information system with an auto-scaling function, the sharing of computing resources between subsystems is exemplified as a result of auto-scaling. However, this is not limited to auto-scaling. For example, the sharing of computing resources between subsystems may be performed when an attempt is made to add computing resources rather than auto-scaling. Furthermore, if the information system is an on-premise system and the physical computing resources are insufficient, the sharing of computing resources between subsystems may be performed.
[0078] The above-described embodiments of the present invention are examples for explaining the present invention, and are not intended to limit the scope of the present invention to only these embodiments. Those skilled in the art can implement the present invention in various other forms without departing from the scope of the present invention. In addition, the present embodiments include the following features. However, the features included in the present embodiments are not limited to those described below.
[0079] (Item 1) A resource management device for managing computational resources of an information system including a plurality of subsystems, a resource shortage prediction detection unit that predicts a future required resource quantity indicating the quantity of computing resources that will be required by each of the plurality of subsystems, and if the future required resource quantity exceeds a current resource quantity that is the quantity of computing resources currently allocated to the subsystem, determines the subsystem as a shortage subsystem and detects a resource shortage prediction for the shortage subsystem; an error budget change estimation unit that estimates a change in the remaining amount of an error budget, which is an index indicating the degree of allowable reliability loss in each non-deficient subsystem, when a calculation resource corresponding to an estimated resource shortage quantity, which is the difference between the future required resource quantity and the current resource quantity, is deleted from the non-deficient subsystem, which is a subsystem in which a resource shortage expectation has not been detected; a resource accommodation method specifying unit that specifies an accommodation subsystem that is a non-deficient subsystem that should transfer computing resources to the deficit subsystem based on a change in the remaining amount of the error budget of each of the non-deficient subsystems; a resource accommodation control unit that changes the allocation of computing resources from the accommodation subsystem to the shortage subsystem by the estimated resource shortage quantity; It has. This enables appropriate management of computing resources in an information system having multiple subsystems, taking into account the consumption of the error budget.
[0080] (Item 2) When an attempt is made to add a computing resource to a subsystem, the resource shortage prospect detection unit monitors the status of the addition of the computing resource, and determines whether or not there is a prospect of a resource shortage based on the status. Since resource shortages are likely to occur when adding computational resources to a subsystem, it is possible to accurately predict resource shortages by determining the likelihood of resource shortages based on the status of adding computational resources.
[0081] (Item 3) The resource shortage likelihood detection unit determines that a resource shortage is likely to occur in the subsystem if addition of computing resources to the subsystem fails. When an attempt is made to add a computing resource to a subsystem, if the addition of the computing resource fails due to reasons such as a lack of computing resources to be added, the quantity of resources required in the future reaching the upper limit of the computing resources that can be allocated to the subsystem, or a setting error in the process of setting up the computing resource to be added, it is determined that a resource shortage is expected, making it possible to accurately predict a resource shortage.
[0082] (Item 4) The resource shortage detection unit determines that the subsystem is likely to have a resource shortage if the addition of computing resources to the subsystem is not completed in time for when it is needed. This makes it possible to accurately predict resource shortages, since when an attempt is made to add computational resources to a subsystem, if the addition of computational resources is not completed in time for when the addition is required, it is determined that resource shortages are likely.
[0083] (Item 5) If the current required resource quantity, which indicates the number of computing resources required by the subsystem, exceeds the current resource quantity before the addition of computing resources to the subsystem is completed, the resource shortage prospect detection unit determines that the addition of computing resources to the subsystem will not be possible in time for when they are needed. This makes it possible to accurately predict resource shortages when additional computing resources cannot be added in time due to a sudden increase in access or batch retries.
[0084] (Item 6) The resource shortage probability detection unit predicts changes in error budgets for the plurality of subsystems, and determines that a subsystem that is expected to consume the error budget is likely to experience a resource shortage. This allows the predicted consumption of the error budget to allocate computing resources to subsystems where the error budget is expected to run out, ensuring stable operation of the IT system.
[0085] (Item 7) The resource accommodation method specifying unit specifies the accommodation subsystem based on the remaining amount of the error budget of each of the non-deficient subsystems or the amount of change in the remaining amount. This allows subsystems with relatively generous error budgets or relatively stable subsystems to be identified as adaptive subsystems.
[0086] (Item 8) the information system is a cloud-based system that has a function of increasing or decreasing computing resources by auto-scaling, When auto-scaling occurs, the resource shortage potential detection unit monitors the status of the auto-scaling and determines whether or not there is a potential resource shortage based on the status. This makes it possible to accurately predict resource shortages by determining the likelihood of resource shortages based on the auto-scaling status.
[0087] (Item 9) When the resource accommodation control unit changes the allocation of the computing resources from the accommodation subsystem to the shortage subsystem, it presents information that an accommodation of the computing resources between the subsystems has occurred. This notifies the user that a sharing of computing resources has occurred between subsystems within the information system, and helps the user check the computing resources.
[0088] (Item 10) A resource management method for managing computational resources of an information system including a plurality of subsystems by a resource management device, comprising: The computer For each of the plurality of subsystems, a future required resource quantity is predicted, which indicates the quantity of computing resources that the subsystem will require, and if the future required resource quantity exceeds a current resource quantity, which is the quantity of computing resources currently allocated to the subsystem, the subsystem is determined to be a deficient subsystem, and a resource shortage forecast for the deficient subsystem is detected; a trial calculation of a change in the remaining error budget, which is an index showing the degree of allowable reliability loss, in each of the non-deficient subsystems, which are subsystems for which a resource shortage has not been detected, when a calculation resource corresponding to an estimated resource shortage quantity, which is the difference between the future required resource quantity and the current resource quantity, is deleted from the non-deficient subsystems, which are subsystems for which a resource shortage prediction has not been detected; Identifying a flexible subsystem that is a non-deficient subsystem that should transfer computing resources to the deficit subsystem based on a change in the remaining amount of the error budget of each of the non-deficient subsystems; The allocation of computing resources is changed from the accommodation subsystem to the shortage subsystem by the estimated resource shortage quantity. This enables appropriate management of computing resources in an information system having multiple subsystems, taking into account the consumption of the error budget. [Explanation of symbols]
[0089] 101...IT system, 102...network, 103...management server, 104...console, 111...computing resource group, 112...monitoring unit, 113...error budget management unit, 114...computing resource management unit, 115...logical resource management unit, 201...processor, 202...memory, 203...auxiliary storage device, 204...communication interface, 205...media interface, 206...input / output device, 207...external storage medium, 321...error budget management table, 331...IT system collaboration unit, 332...resource Resource accommodation management unit, 333... management input / output unit, 341... detection unit, 342... additional resource information temporary management table, 343... error budget change estimation unit, 344... logical resource specification information temporary management table, 345... error budget characteristic management table, 346... error budget change estimation result temporary management table, 347... resource accommodation method identification unit, 348... accommodation standard management table, 349... resource accommodation method temporary management table, 350... resource accommodation control unit, 361... setting unit, 362... notification unit, 363... display unit
Claims
1. A resource management device for managing computational resources of an information system including a plurality of subsystems, a resource shortage prediction detection unit that predicts a future required resource quantity indicating the quantity of computing resources that will be required by each of the plurality of subsystems, and if the future required resource quantity exceeds a current resource quantity that is the quantity of computing resources currently allocated to the subsystem, determines the subsystem as a shortage subsystem and detects a resource shortage prediction for the shortage subsystem; an error budget change estimation unit that estimates a change in the remaining amount of an error budget, which is an index indicating the degree of allowable reliability loss in each non-deficient subsystem, when a calculation resource corresponding to an estimated resource shortage quantity, which is the difference between the future required resource quantity and the current resource quantity, is deleted from the non-deficient subsystem, which is a subsystem in which a resource shortage expectation has not been detected; a resource accommodation method specifying unit that specifies an accommodation subsystem that is a non-deficient subsystem that should transfer computing resources to the deficit subsystem based on a change in the remaining amount of the error budget of each of the non-deficient subsystems; a resource accommodation control unit that changes the allocation of computing resources from the accommodation subsystem to the shortage subsystem by the estimated resource shortage quantity; A resource management device having:
2. The resource management device according to claim 1 , wherein the resource shortage detection unit monitors the status of the addition of computing resources when an attempt is made to add computing resources to a subsystem, and determines whether or not there is a prospect of a resource shortage based on the status.
3. The resource management device according to claim 2 , wherein the resource shortage likelihood detection unit determines that a resource shortage is likely to occur in the subsystem if an attempt to add computing resources to the subsystem fails.
4. the resource shortage likelihood detection unit determines that a resource shortage is likely to occur in the subsystem if the addition of computing resources to the subsystem is not completed in time for when the resources are needed; The resource management device according to claim 2 .
5. The resource management device of claim 4, wherein the resource shortage potential detection unit determines that the addition of computing resources to the subsystem will not be in time for when it is needed if the current required resource quantity, which indicates the number of computing resources required by the subsystem, exceeds the current resource quantity before the addition of computing resources to the subsystem is completed.
6. The resource management device according to claim 2 , wherein the resource shortage detection unit predicts changes in error budgets for a plurality of the subsystems and determines that a subsystem that is expected to consume the error budget is likely to experience a resource shortage.
7. The resource management device according to claim 1 , wherein the resource accommodation method specifying unit specifies the accommodation subsystem based on the remaining amount of the error budget of each of the non-deficient subsystems or the amount of change in the remaining amount.
8. the information system is a cloud-based system that has a function of increasing or decreasing computing resources by auto-scaling, The resource management device according to claim 1 , wherein the resource shortage potential detection unit monitors a status of the auto-scaling when the auto-scaling occurs, and determines whether or not the resource shortage is likely to occur based on the status.
9. The resource management device according to claim 1 , wherein the resource accommodation control unit, when changing the allocation of computing resources from the accommodation subsystem to the shortage subsystem, presents information that an accommodation of computing resources has occurred between subsystems.
10. A resource management method for managing computational resources of an information system including a plurality of subsystems by a resource management device, comprising: The computer For each of the plurality of subsystems, a future required resource quantity is predicted, which indicates the quantity of computing resources that the subsystem will require, and if the future required resource quantity exceeds a current resource quantity, which is the quantity of computing resources currently allocated to the subsystem, the subsystem is determined to be a deficient subsystem, and a resource shortage forecast for the deficient subsystem is detected; a trial calculation of a change in the remaining error budget, which is an index showing the degree of allowable reliability loss, in each of the non-deficient subsystems, which are subsystems for which a resource shortage has not been detected, when a calculation resource corresponding to an estimated resource shortage quantity, which is the difference between the future required resource quantity and the current resource quantity, is deleted from the non-deficient subsystems, which are subsystems for which a resource shortage prediction has not been detected; Identifying a flexible subsystem that is a non-deficient subsystem that should transfer computing resources to the deficit subsystem based on a change in the remaining amount of the error budget of each of the non-deficient subsystems; changing the allocation of the computing resources from the accommodation subsystem to the shortage subsystem by the estimated resource shortage quantity; Resource management methods.
Citation Information
Patent Citations
Auto scale type performance guarantee system and auto scale type performance guarantee method
JP2020123849A