Management system and method
The management system addresses the challenge of optimizing management operation completion times by using a processor to select countermeasures based on error classification and process status, effectively handling errors and reducing prolonged completion times.
Patent Information
- Application Number
- JP2023202116
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-29
- Publication Date
- 2025-06-10
AI Technical Summary
Existing management systems face challenges in optimizing the time to complete management operations due to complications arising from failed automatic handling processes, which lead to multiple failures and prolonged completion times.
A management system that includes a processor to execute consecutive processes, determine if errors occur, and select countermeasures based on error classification, process execution status, and preset conditions, thereby optimizing the management operation completion time.
The system optimizes the time to complete management operations by effectively handling errors and selecting appropriate countermeasures, thereby reducing the complexity of failure causes and minimizing prolonged completion times.
Smart Images

Figure 2025087453000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a management system and method.
Background Art
[0002] In the operation and management of IT infrastructure, it has become increasingly common to adopt an operation mode in which an administrator who manages the entire IT infrastructure, rather than a storage-exclusive administrator, manages storage. Therefore, there is a demand for storage management that can be easily operated with less labor by general administrators who lack knowledge and skills regarding storage.
[0003] In addition, as customers' use of public clouds progresses, expectations for a service (management service) that centrally manages a service-providing system that provides a storage system in a customer data center and a service-providing system that provides software-defined storage (SDS) operating on a public cloud are increasing.
[0004] For this reason, a Software-as-a-Service (SssS)-type management service managed by a storage management vendor and operating on a public cloud or the like has emerged, providing functions such as the ability to easily introduce without the need for customers to prepare additional hardware such as a management server and perform provisioning (capacity allocation) of storage volumes.
[0005] In management services provided for general administrators, in many cases, a layer that abstracts management APIs (lower-level APIs) that require detailed knowledge of storage is prepared, and a mechanism is provided in which a plurality of lower-level APIs are executed internally by the execution of one abstract API (higher-level API), thereby aiming to streamline management. At this time, when the processing fails during the execution of the higher-level API, in addition to dealing with the cause of the failure, it becomes necessary to deal with the state during the execution of the higher-level API.
[0006] On the other hand, as a technique for the management service to automatically handle processing failures, for example, Patent Document 1 is known. In Patent Document 1, when the creation processing for a plurality of resources executed by one execution request fails midway, an automatic rollback is performed to return the plurality of resources for which the processing has failed to their original state. Also, when the deletion processing for a plurality of resources fails midway, the processing is automatically retried and a roll forward is performed.
Prior Art Documents
Patent Documents
[0007]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0008] In the technique of Patent Document 1, when the automatic handling is successful, it is possible to handle the in-execution state. On the other hand, when the automatic handling fails, in addition to the initially occurring failure, failures resulting from the automatic handling also occur, leading to a state where multiple problems occur and the handling of the failure causes becomes complicated.
[0009] Also, even if the automatic handling by rollback is successful, the time required to return the operations that were successful up to midway to the pre-execution state may become long, and there are cases where it takes a long time until the management operation is completed.
[0010] Therefore, an object of the present invention is to provide a technique capable of optimizing the time until the completion of a management operation.
Means for Solving the Problems
[0011] To solve the above problems, one of the representative management systems of the present invention is a management system for managing one or more bases. The management system includes a processor. The processor executes a plurality of consecutive processes on the base in response to a request, determines whether an error has occurred during the execution of the plurality of consecutive processes, and when an error occurs, selects a countermeasure based on at least one of the classification of the error, the execution status of the process, and a preset countermeasure start condition, and executes the selected countermeasure.
Effect of the Invention
[0012] According to the present invention, the time until the completion of the management operation can be optimized. Problems, configurations, and effects other than those described above will be clarified by the description of the following embodiments.
Brief Description of the Drawings
[0013]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Mode for Carrying Out the Invention
[0014] Hereinafter, embodiments of the present invention will be described with reference to the drawings. However, the present invention is not to be construed as being limited to the description of the embodiments shown below. It will be readily understood by those skilled in the art that the specific configuration can be changed without departing from the spirit or gist of the present invention. In the configuration of the invention described below, the same or similar configurations or functions are denoted by the same reference numerals, and duplicate descriptions are omitted. In this specification and the like, notations such as "first", "second", "third", etc. are attached for identifying components, and do not necessarily limit the number or order.
Embodiment
[0015] FIG. 1 is a block diagram showing an example of the configuration of a system to which the management system of Example 1 is applied.
[0016] The system to which the management system of Example 1 is applied includes a management system 100 and a plurality of bases 101.
[0017] The management system 100 is connected to the plurality of bases 101 via a network 102 such as a WAN (Wide Area Network), a LAN (Local Area Network), and a SAN (Storage Area Network).
[0018] The base 101 provides an environment for constructing a service providing system that provides services.
[0019] The base 101 may be either an on-premises type base or a cloud type base. The service providing system may be composed of physical elements such as computers, or may be composed of virtual elements such as virtual machines.
[0020] The management system 100 centrally manages the service - providing system built on the infrastructure 101.
[0021] The management system 100 has a request control unit 110 and a response determination unit 111, and also holds request - processing correspondence information 120, execution status management information 121, error - handling information 122, and processing time information 123.
[0022] The request - processing correspondence information 120 is information for managing the correspondence between requests to the infrastructure 101 required in the management system 100 and the processes provided by the infrastructure 101.
[0023] The execution status management information 121 is information for managing the execution status of requests to the infrastructure 101 required in the management system 100.
[0024] The error - handling information 122 is information for managing errors that may occur in the infrastructure 101 and the countermeasures against those errors.
[0025] The processing time information 123 is information for managing the processes in the infrastructure 101 and the time that those processes may take.
[0026] Note that the management system 100 may be included in any of the infrastructures 101.
[0027] The management system 100 is composed of, for example, the computer 200 shown in FIG. 2.
[0028] FIG. 2 is a block diagram showing an example of the hardware configuration of the computer that constitutes the management system 100 of Example 1.
[0029] FIG. 1 shows an example of the hardware configuration of the computer 200 in which the request control unit 110 and the response determination unit 111 shown in FIG. 1 operate.
[0030] The computer 200 is a server or a computer configured by connecting a processor 201, a memory device 202, an input device 203, an output device 204, and a communication I / F 205 to each other via a bus 206.
[0031] The processor 201 operates as a functional unit (module) that realizes a specific function by executing processing according to a program stored in the memory device 202. In the following description, when explaining the processing with the functional unit as the subject, it indicates that the processor 201 is executing a program that realizes the functional unit.
[0032] The memory device 202 is a main memory device used when the processor 201 executes processing, and is composed of a volatile memory element such as a RAM (Random Access Memory).
[0033] The input device 203 is an interface that receives input from a user (operator), and is composed of a keyboard, a touch panel, a card reader, a voice input device, or the like.
[0034] The output device 204 is an interface that outputs data to the operator, and is composed of a display, a speaker, a printer, or the like.
[0035] The communication I / F 205 is an interface used for the computer 200 to communicate with an external device, and is composed of a NIC (Network Interface Card) or the like. The communication I / F 205 is connected to the network 102 and communicates with the base 101 via the network 102.
[0036] The bus 206 is an internal communication path of the computer 200.
[0037] In this embodiment, the management system 100 can realize each of the processes described later by being executed on one or more computers 200 having a hardware configuration as illustrated in FIG. 2.
[0038] Still, the request - processing response information 120, the execution status management information 121, the error handling information 122, and the processing time information 123 are stored in the storage device 202.
[0039] FIG. 3 is a block diagram showing an example of the configuration of the base 101 of the first embodiment.
[0040] The base (1) 101 is, for example, an on - premise type base and includes a storage system 310.
[0041] The storage system 310 is an example of a service - providing system and may be composed of a server or the like. The storage system 310 provides volumes.
[0042] The storage system 310 holds service management information 311 that stores data regarding the performance of the storage system 310 and the like.
[0043] Also, the storage system 310 provides an API (1) 312 for the management system 100 to access.
[0044] Also, the storage system 310 has an inter - device connection part 313 for data exchange with another base.
[0045] Data exchange between bases is carried out via the network 102 and is used, for example, for remote backup between volumes of the storage system.
[0046] The base (2) 101 is, for example, a cloud - type base and includes SDS (Software Defined Storage) 320.
[0047] The SDS 320 is an example of a service - providing system and may be composed of various services provided by the cloud base and a server or the like. The SDS 320 provides volumes.
[0048] SDS320 is composed of a server 330 and one or more storage volumes 331.
[0049] SDS320 holds service management information 321 that stores data related to the performance of SDS320 and the like.
[0050] Also, SDS320 provides an API (2) 322 for the management system 300 to access, and also provides an API (3) 323 for the management system to access the server 330.
[0051] Also, SDS320 has a device connection part 324 for data exchange with another infrastructure.
[0052] The infrastructure 101 and the system on the infrastructure 101 shown in FIG. 3 are examples, and the present invention is not limited thereto.
[0053] FIG. 4 is a diagram showing an example of the data structure of the request - process correspondence information 120 of Example 1.
[0054] The request - process correspondence information 120 stores entries including a request 401, an execution type 402, a target 403, a process 404, and a process classification 405.
[0055] For one request 401, there are entries in a form where a plurality of processes 404 correspond. The request 401 stores request information for the infrastructure 101 required in the management system 100. Here, the request information stores information that can uniquely identify a function using the service - providing system of the infrastructure 101 provided by the management system 100.
[0056] The execution type 402 stores whether the process 404 is of the type "normal" during normal execution or "rollback" during rollback.
[0057] The target 403 stores the type of service - providing system in the infrastructure 101 that is the execution target of the process 404, such as a storage system (Storage), SDS, etc.
[0058] The process 404 stores information that can identify the processes executed in the storage system, such as the API name provided by the API (1) 312 of the storage system 310.
[0059] The process classification 405 stores information on whether the process 404 is classified into any of the processes of Create, Update, or Delete.
[0060] In the request 401 of FIG. 4, as an example, "Allocate Volume" that requests volume provision of the storage system 310 of the infrastructure 101, "Change Assign Volume" that requests change of volume capacity and access information, and "Config Cloud Backup" that performs remote backup setting between infrastructures using the device connection parts 313, 324 of the infrastructure 101 are shown.
[0061] The request 401 is not limited to the above examples, and any request that uses the infrastructure 101 provided by the management system 100, such as a request to cancel volume provision of the storage system 310, a request to obtain a snapshot of the volume of the storage system 310, etc., is acceptable.
[0062] FIG. 5 is a diagram showing an example of the data structure of the execution - status management information 121 of Example 1.
[0063] The execution - status management information 121 stores entries including ID 501, request 502, status 503, error type 504, target system 505, process number 506, corresponding process number 507, process 508, process status 509, and process count 510.
[0064] Entries are in a form where a plurality of processes 508 correspond to one request 502.
[0065] ID501 stores an ID for uniquely identifying a request for the infrastructure 101 requested in the management system 100.
[0066] The request 502 stores request information requested in the management system 100.
[0067] The status 503 stores the execution status of the request. For example, when the request fails, "Failed" is stored, when it succeeds, "Success" is stored, and when it is in progress, "Processing" is stored.
[0068] The error type 504 stores the details of the error that occurred when the request ended in error, for example, when the status 503 is "Failed".
[0069] The target system 505 stores an identifier indicating the service - providing system within the infrastructure 101 that is the target for executing the request, and indicates the service that is the target for executing the process 508.
[0070] The process number 506 stores the execution order assigned to the process 508 executed in the request 502. For example, in the example of FIG. 5, for the request "Allocate Volume", P1 (Process number 1) is assigned to the "Server creation" process, which is the normal process executed first, and P2 (Process number 2) is assigned to the "Host WWN setting", which is the normal process executed second.
[0071] Also, R1 (Rollback number 1) is assigned to the "Volume Server setting", which is the rollback process that can be executed first, and R2 (Rollback number 2) is assigned to the "Path setting", which is the rollback process that can be executed second.
[0072] The corresponding process number 507 is set only for the rollback process, and the process number of the normal process that performs the rollback in the rollback process is stored. For example, in the entry where the process number 506 is R5, the corresponding process number P1 is stored, indicating that the process "Server deletion" with the process number 506 being R5 is the rollback process of the process "Server creation" with the process number 506 being P1.
[0073] Process 508 stores the processes executed in the service - providing system. Here, in process 508, the process indicated by process 404 of the request - process correspondence information 120 and the identifier indicating the ID of the object in the target system 505 where the process is performed are stored. For example, when the process is "Server creation", the value indicating the ID of the server is stored, and when the process is "Host WWN setting", the value indicating the ID of the Host WWN is stored.
[0074] Also, when the process is an operation on multiple resources, multiple values indicating the IDs are stored. For example, when the process is "Volume Server setting", the ID of the server and the ID of the volume are stored.
[0075] Also, in the case of a process where the process classification 405 is Update, the content to be changed is also stored. For example, when the process is "Volume setting change", the ID of the volume and the setting change content of changing the capacity "Capacity" of the volume from 100GB to 1TB are stored.
[0076] The process status 509 stores information indicating the execution status of process 508. When the execution of the process is successful, the value "Success" is stored, when the execution of the process fails, the value "Failed" is stored, and when the request is not executed, the value "Not Executed" is stored.
[0077] The processing count 510 stores information indicating the number of times the process 508 is executed. Here, if the process 508 is successfully executed once, "1" is stored in the processing count 510, and if the process 508 fails once and a retry is performed once, "2" is stored in the processing count 510.
[0078] Also, if an error occurs during the execution of the request and the process fails, subsequent normal processing is not performed, and "-" indicating that it has not been executed is stored in the processing count 510.
[0079] Also, since the rollback process is a process of reverting the process up to where the normal process was successful, for the rollback process in which the process number 606 of the failed normal process is stored in the countermeasure process number 507, "-" indicating that it has not been executed is also stored in the processing count 510.
[0080] FIG. 6 is a diagram showing an example of the data structure of the error countermeasure information 122 of the first embodiment.
[0081] The error countermeasure information 122 stores entries including an error code 601, an error classification 602, and a countermeasure start condition 603.
[0082] The error code 601 stores a code indicating the content of an error that can occur in the service providing system of the base 101. For example, as a code indicating the content of an error in the storage system 310 of the base 101, "OPERATION_TIMEOUT" indicating that the process has timed out, "TOO_MANY_REQUEST" indicating that a large number of processing requests have been received and cannot be processed, etc. are stored.
[0083] The error classification 602 stores a value indicating the classification based on the content of the error indicated by the error code 601. For example, "temporary" indicating that it is a temporary error, "configuration inconsistency" indicating that the content specified in the request is inconsistent with the state of the service providing system, "internal" indicating that some problem has occurred inside the service providing system, etc. are stored.
[0084] The coping start condition 603 stores information on what conditions need to be met for automatic coping to be performed when an error with error code 601 occurs. If these conditions are not met, automatic coping is not performed because an error is likely to occur again if automatic coping is carried out.
[0085] The coping start condition 603 stores, for example, "CPU < 100%" indicating that the CPU utilization rate of the base 101 is not 100%, "Number of trials < 3" indicating that the number of trials of the process is less than 3, "Resource Unlocked" indicating that the resources related to the error are not locked in the base 101, "Waiting time 3 min" indicating that automatic coping is performed 3 minutes after the error occurs, "Refresh" indicating that automatic coping is performed after a refresh operation, "-" indicating the non-existence of a condition, and so on.
[0086] FIG. 6 shows an example of the conditions for performing a retry coping as an example of the coping start condition 603.
[0087] In addition, in this embodiment, the conditions for performing rollback coping are "CPU < 100%", "Number of trials = 1", and "Resource Unlocked", which are common for all error codes.
[0088] Here, the conditions for performing rollback coping may also be managed by the error coping information 122. Specifically, it is not limited to this, and multiple copings (retry and rollback) for the same error code and the coping start conditions in each coping may be held.
[0089] FIG. 7 is a diagram showing an example of the data structure of the processing time information 123 of the first embodiment.
[0090] The processing time information 123 stores entries including a process 701, a target type 702, a unit processing time 703, a unit value 704, and a calculation formula 705.
[0091] The process 701 stores information that can identify a process executed in the service providing system of the infrastructure 101, such as the API name provided by the API (1) 312 of the storage system 310.
[0092] The target type 702 stores the type of the service providing system in the infrastructure 101 that is the execution target of the process 701, such as a storage system (Storage) or SDS. Here, it is managed as the target type 702, but the entry of the processing time information 123 may be managed in units of specific instances in the actual environment, that is, not "Storage" but "Storage1", "Storage2", etc.
[0093] The unit processing time 703 stores information on the time required when performing a process on the unit stored in the unit value 704.
[0094] The unit value 704 stores unit values related to capacities such as 1 GB and 1 MB.
[0095] The calculation formula 705 stores a formula for calculating the processing time of the process 701.
[0096] Here, the unit processing time 703, the unit value 704, and the calculation formula 705 may be based on information provided by the vendor providing the service providing system of the target type 702 in a manual, white paper, etc., or based on the actual performance value when actually performing the process 701 on the target type 702, and are not limited to this.
[0097] Also, here only an example of the processing time for a single target type 702 is described, but the processing time for operations on multiple target types 702, specifically, the processing time for data copy operations in operations such as backup from a storage system to SDS, etc., may also be described.
[0098] In addition, as an example of the calculation formula, an example of a formula with only capacity, unit processing time, and quantity as variables has been shown, but other items may also be used as variables. For example, information on system components such as the CPU, memory, and ports of the storage system, the usage status of their configuration information in the actual environment, NW-related information such as the NW bandwidth and the number of switch connections between the storage system and the SDS, etc. may be used as variables, and it is not limited to this.
[0099] FIG. 8 is a flowchart for explaining an example of the request control process executed by the management system 100 of the first embodiment.
[0100] The request control unit 110 receives a request to the base 101 (step S101). The request is input from the input device 203 by a user who uses the management system 100, or is transmitted from a client program (not shown) that accesses the management system 100 via the communication I / F 205.
[0101] Subsequently, an entry indicating the request received in step S101 is extracted from the request - process correspondence information 120, an object to be processed for the request is selected based on the configuration of the service - providing system in the base 101, and information including the information of the selected object and the ID 501 numbered for the request here is stored in the execution status management information 121 (step S102).
[0102] Here, for example, in the example of "Allocate Volume" shown in FIG. 4, the information of each object of Storage of the target 403, Server creation of the process 404, Host WWN setting, Volume creation, Path setting, and Volume Server setting is determined based on the configuration of the service - providing system in the base 101. For example, it is determined to create a Server with the identification information "5".
[0103] Here, although not shown because the configuration of a service providing system such as a storage system is generally held as configuration information, it is information on the internal configuration of devices such as servers and storage (CPU, memory, ports, volumes, pools, disks, etc.) existing in the service providing system, and the connection relationship information between devices such as servers and storage.
[0104] Also, the object extraction method can be any method. For example, in the storage system 310, creating a Volume from the storage area with the most remaining capacity, setting the Volume so that it can be accessed via the Host WWN with the fewest Volume usages, selecting the smallest available number among the numbers for objects that can be set by the device, etc. can be cited, but it is not limited to this.
[0105] Next, in step S103, a list of processes with the execution type of "normal" is created from the execution status management information 121 (step S103).
[0106] Next, a loop process that repeats the processes of S105 to S112 is started until all the processes in the list created in step S103 are completed (step S104).
[0107] In step S105, the processes in the list are executed in the order of process number 506. Here, for example, the management system 100 executes "Server creation" with the process number P1 on the base 101.
[0108] Next, based on the execution result of the process in step S105, the status 503 of the request in the execution status management information 121, the error type 504 if an error has occurred, the process status 509, and the process count 510 indicating how many times the process 508 has been executed are stored, and the execution status management information 121 is updated (step S106).
[0109] Next, it is confirmed whether an error has occurred in the process executed in S105 (step S107).
[0110] If no error has occurred in step S107, the process has succeeded. Therefore, the loop process is transferred to the process with the next process number 506, and the loop process is repeated until all processes in the list are completed.
[0111] If an error has occurred in step S107, the process has failed. Therefore, a coping determination request is sent from the request control unit 110 to the coping determination unit 111, and the coping determination process is called (step S108).
[0112] Next, it is determined whether the determination result of the coping determination process in step S108 is a coping of the coping presentation process that does not execute automatic coping and presents the coping procedure to the user (step S109).
[0113] In step S109, if the determination result of the coping determination process is not a coping of the user presentation process, it is determined whether the determination result of the coping determination process in step S108 includes a retry process for re-executing the process in which the error occurred (step S110).
[0114] In step S110, if the determination result of the coping determination process does not include a retry process, the subsequent processes in the list are deleted (step S111), and the process proceeds to step S112.
[0115] In step S110, if the determination result of the coping determination process includes a retry process, the process proceeds to step S112 without executing step S111.
[0116] In step S112, the automatic coping process (for example, retry process, rollback process, coping start condition process) of the determination result of the coping determination process is inserted into the next process in the list.
[0117] In step S109, if the determination result of the coping determination process is a coping of the user presentation process, since it is determined that automatic coping cannot be executed or is not executed in the coping determination process, the process of presenting the coping procedure to be performed by the user to the user is executed (step S113), and this process is terminated.
[0118] When a request input from the input device 203 by the user is received in step S101, in step S113, an error handling display screen for presenting a handling procedure is created and displayed on the output device 204. Details of the error handling display screen will be described later with reference to FIG. 10.
[0119] Also, when a request from the client program is received in step S101, in step S113, a response including a handling procedure is transmitted to the client program.
[0120] FIG. 9 is a flowchart for explaining an example of a countermeasure determination process executed by the management system 100 of the first embodiment.
[0121] The countermeasure determination unit 111 receives a countermeasure determination request from the request control unit 110 (step S201).
[0122] The countermeasure determination unit 111 refers to the error handling information 122 and acquires the error classification 602 and the countermeasure start condition 603 of the occurring error (step S202).
[0123] Subsequently, with reference to the error classification 602 acquired in step S202, it is determined whether the occurring error is a temporary error (step S203).
[0124] In step S203, if the occurring error is a temporary error, it is determined whether the countermeasure start condition 603 other than the process among the countermeasures for executing the retry process is satisfied (step S204).
[0125] Here, when checking whether there is a possibility of a problem due to the performance, capacity, configuration, etc. of the system components of the service providing system, such as "CPU utilization < 100%", the performance information, capacity information, configuration information, etc. of the service providing system are referred to respectively. Since all of these information are general management information in the service providing system, they are not shown in the figure.
[0126] Also, when checking whether a lock is taken on an object targeted by the process by other processes, or whether there is a possibility of a problem due to the process execution state, such as whether multiple processes are executed simultaneously and the process execution of the service providing system is not delayed, refer to the information of the process being executed in the service providing system. The process information is general management information in the service providing system managed under names such as job information and task information, and is not shown in the figure.
[0127] In step S204, when the coping start condition 603 other than the process is satisfied, check whether the coping start condition 603 includes a process (step S205). The process mentioned here indicates, for example, a process of collecting the latest information of a device called "refresh" into the management system, but is not limited thereto.
[0128] In step S205, when the coping start condition 603 includes a process, return the process of the coping start condition 603 and the retry process as a coping determination result to the request control unit 110 (step S206), and end this process.
[0129] In step S205, when the coping start condition does not include a process, return the retry process as a coping determination result to the request control unit 110 (step S207), and end this process.
[0130] When the error occurring in step S203 is not a temporary error, and when the coping start condition 603 other than the process is not satisfied in step S204, it is determined that there is a high possibility that an error will occur again during the execution of the retry process. Therefore, the retry is not performed, and it is determined whether the error occurring is an error during rollback (step S208).
[0131] If the error occurring in step S208 is not an error during rollback, refer to the execution status management information 121, extract the rollback process having the same corresponding process number as the process number of the last normal process for which the processing status 509 is "Success", and create a list of rollback processes after the said rollback process (step S209).
[0132] Next, start a loop process that repeats the process of S211 until the processing of all rollback processes in the list created in step S209 is completed (step S210).
[0133] In step S211, determine whether the conditions for executing the preset rollback countermeasures are satisfied.
[0134] Here, as described above, the conditions for executing the rollback countermeasures are "CPU < 100%", "number of trials = 1", and "Resource Unlocked" in common for all error codes. By checking whether these conditions are satisfied, it is possible to determine whether there is a high possibility of an error occurring during rollback.
[0135] Note that the conditions for executing the rollback countermeasures may use conditions other than the above and are not limited thereto.
[0136] Also, the conditions for executing the rollback countermeasures may be set to different conditions for each base service - providing system. This is because the usage status and specifications of resources such as the CPU are different in each base service - providing system (for example, a storage system, SDS).
[0137] In this case, it may be determined that the storage system can be rolled back, but the SDS cannot be rolled back. Therefore, if the determination in step S211 is "No" for all the processes in the list for any service providing system, the process for the service providing system is rolled back. If the determination in step S211 for the processes in the list is divided into "Yes" and "No", the process for the service providing system may not be rolled back, etc.
[0138] If it is determined that the conditions for executing rollback handling in step S211 are satisfied for all the processes in the list, the time required to return to the state before request execution and the time required to continue executing the subsequent processes from the process where an error occurred in the request are estimated (step 212). Here, the estimation is performed based on the calculation formula 705 of the processing time information 123. Specifically, the information received in the request is set for the capacity and the number in the calculation formula for estimation.
[0139] Subsequently, in step S213, the rollback completion time required to return to the state before request execution and the request completion time required to execute the processes after the process where an error occurred in the request are estimated and compared.
[0140] If the request completion time is longer than the rollback completion time in step S213 (step S213), the rollback process list created in step S209 is returned to the request control unit 110 as the determination result (step S214), and this process ends.
[0141] Here, the estimation of the rollback completion time and the request completion time may be calculated as the time obtained by adding up the individual processing times according to the limitations of the service providing system, etc. Or for processes that can be parallelized, it may be calculated as the time obtained by adding up after subtracting the processing time of the parallelized part.
[0142] For example, when calculating the rollback completion time and the request completion time of a request including processing to the base platform (1) and processing to the base platform (2), since the processing on the base platform (1) side and the processing on the base platform (2) side are often executed independently and in parallel, the time is calculated taking this into account.
[0143] If it is determined in step S208 that the error that has occurred is not an error during rollback, if the conditions for executing rollback countermeasures are not satisfied in step S211, and if the request completion time is shorter than the rollback completion time in step S213, the automatic countermeasure is not executed and the process is suspended, and a user prompt process for presenting the countermeasure procedure to be performed by the user to the user is returned to the request control unit 110 as the countermeasure determination result (step S215), and this process ends.
[0144] Also, in FIG. 9 of this embodiment, retry is selected as the top priority for error handling. If retry seems difficult, rollback is selected. If rollback also seems difficult, user prompting of the countermeasure procedure is selected, so that error handling can be easily performed.
[0145] Here, in advance as policies, "retry priority policy", "rollback priority policy", "user prompt priority policy", "rollback / retry candidate selection policy", "cost priority policy", "operation completion time priority policy", etc. are set, and error handling may be selected according to the set policy.
[0146] For example, when the "rollback priority policy" is set, after step S202 in FIG. 9 is executed, the process proceeds to step S208, and rollback or user prompting is performed.
[0147] Also, when the "user prompt priority policy" is set, after step S202 in FIG. 9 is executed, the process proceeds to step S215, and user prompting is performed.
[0148] Also, when the "rollback / retry candidate selection policy" is set, the user is allowed to select either rollback or retry.
[0149] Also, when the "cost priority policy" is set, processing that is less costly is prioritized. For example, when the process classification 405 is "Create", rollback is prioritized, and when the process classification is "Delete", retry is prioritized, so that the operation moves in the direction of deleting / reducing resources. This is particularly effective in cases where leaving resources such as SDS on the public cloud incurs significant costs.
[0150] Also, when the "operation completion time priority policy" is set, the time required for automatic handling is estimated, and processing that takes less time until operation completion is prioritized.
[0151] According to this process, when it is determined that there is a high possibility of an error occurring when retrying, rollback is executed or a handling procedure is presented to the user, thus preventing the error handling from being prolonged due to the occurrence of an error during retry and optimizing the time until the management operation is completed.
[0152] Also, when it is determined that there is a high possibility of an error occurring when executing rollback, the handling procedure is presented to the user without executing rollback, thus preventing the error from becoming complicated due to the occurrence of an error during rollback.
[0153] Also, when the rollback completion time is longer than the request completion time, the handling procedure is presented to the user without executing rollback, so that error handling can be performed quickly.
[0154] FIG. 10 is a diagram showing an example of an error handling display screen presented by the management system 100 of Example 1.
[0155] The error handling display screen 2000 is an example of an error handling display screen displayed on the output device 204 in step S113 of FIG. 8.
[0156] On the error handling display screen 2000, error information 2001 and handling procedures 2002 are displayed.
[0157] The error information 2001 includes the ID 501 of the request in which the error occurred, the request 502, the error code 601, the date and time when the error occurred, and an error message defined in advance corresponding to the error code.
[0158] In the handling procedures 2002, the handling procedures to be performed by the user to handle the occurred error are presented. Here, as an example, an example of instructing the user to perform deletion of the Host WWN and deletion of the Server is shown.
[0159] Here, an example of presenting only the procedure for performing a manual rollback as the handling procedure 2002 has been shown, but only the procedure for manually proceeding with the process may be presented, or both of these procedures may be presented.
[0160] Also, when presenting both the rollback procedure and the procedure for proceeding with the process, information that serves as a basis for the user to determine which procedure to choose may be presented together. This information is, for example, the information on the rollback completion time and the request completion time obtained in step S212 of FIG. 9, the information on the cost incurred by leaving the resources without deletion (especially in the case of a service management system that uses public resources), the information on the items that are harmful to the execution of the process obtained in step S204 and step S211, and the like. Also, these information may be presented after being prioritized by comprehensively considering them.
[0161] In this way, by referring to the information displayed on the error handling display screen 2000, the user can easily handle the error during the execution of the request. Also, presenting the cost information on the error handling display screen 2000 facilitates cost management.
[0162] Note that the present invention is not limited to the above-described embodiments and includes various modifications. Also, for example, the above-described embodiments are those in which the configuration is described in detail for easy understanding of the present invention, and are not necessarily limited to those having all the configurations described. Also, for a part of the configuration of each embodiment, it is possible to add, delete, or replace other configurations.
[0163] Further, each of the above configurations, functions, processing units, processing means, etc. may be realized in hardware by designing a part or all of them, for example, by an integrated circuit.
[0164] Also, the present invention can also be realized by a program code of software that realizes the functions of the embodiments. In this case, a storage medium recording the program code is provided to a computer, and a processor included in the computer reads the program code stored in the storage medium. In this case, the program code itself read from the storage medium realizes the functions of the above-described embodiments, and the program code itself and the storage medium storing it constitute the present invention. As a storage medium for supplying such a program code, for example, a flexible disk, CD-ROM, DVD-ROM, hard disk, SSD (Solid State Drive), optical disk, magneto-optical disk, CD-R, magnetic tape, non-volatile memory card, ROM, etc. are used.
[0165] Also, the program code for realizing the functions described in this embodiment can be implemented in a wide range of programs or script languages such as assembler, C / C++, perl, Shell, PHP, Python, Java (registered trademark), etc.
[0166] Furthermore, the program code of the software that realizes the functions of the embodiments may be distributed via a network, stored in a storage means such as a hard disk or memory of a computer, or a storage medium such as a CD-RW or CD-R, and the processor included in the computer may read and execute the program code stored in the storage means or the storage medium.
[0167] In the above embodiments, the control lines and information lines show those considered necessary for explanation, and not necessarily all the control lines and information lines are shown on the product. All the components may be interconnected.
Explanation of Reference Numerals
[0168] 100 Management System 101 Infrastructure 102 Network 110 Request Control Unit 111 Response Judgment Unit 120 Request-Processing Response Information 121 Execution Status Management Information 122 Error Handling Information 123 Processing Time Information 200 Computer 201 Processor 202 Memory Device 203 Input Device 204 Output Device 205 Communication I / F 311, 321 Service Management Information 310 Storage System 320 SDS (Software Defined Storage) 330 Server 331 Storage Volume
Claims
1. A management system for managing one or more bases, wherein the management system includes a processor, and the processor, executes a plurality of consecutive processes on the base in response to a request, determines whether an error has occurred during the execution of the plurality of consecutive processes, and when the error occurs, selects a countermeasure based on at least one of the classification of the error, the execution status of the process, and a preset countermeasure start condition, and executes the selected countermeasure.
2. In the management system according to Claim 1, the processor, when the classification of the error is temporary, selects a retry to re-execute the process in which the error occurred, and when the classification of the error is not temporary, selects a rollback to return to the state before the execution of the plurality of consecutive processes or a suspension in the state during the execution of the plurality of consecutive processes.
3. In the management system according to Claim 2, the processor, when the error occurs, calculates and compares the rollback completion time for the rollback with the request completion time for executing the processes after the process in which the error occurred in the plurality of consecutive processes, and selects the countermeasure based on the result of the comparison.
4. In the management system according to Claim 3, the processor, when the result of the comparison shows that the rollback completion time is shorter than the request completion time, selects the rollback, and when the result of the comparison shows that the rollback completion time is longer than the request completion time, selects the suspension.
5. In the management system according to Claim 2, the countermeasure start condition includes the state of the base, and the processor selects the countermeasure when it determines that the countermeasure can be executed based on the countermeasure start condition.
6. In the management system according to Claim 2, the processor, when it selects the suspension, presents the countermeasure procedure to be performed by the user to the user.
7. In the management system according to Claim 2, the processor selects the countermeasure according to a preset priority order.
8. In the management system according to Claim 7, The priority includes at least one of retry priority, rollback priority, user prompt priority, candidate selection, operation completion time priority, and cost priority in a management system.
9. A management method for a management system that manages one or more bases, The management system includes a processor, The processor Executes a plurality of consecutive processes on the base in response to a request, Determines whether an error has occurred during the execution of the plurality of consecutive processes, When the error occurs, selects a countermeasure based on at least one of the classification of the error, the execution status of the process, and a preset countermeasure start condition, A management method for executing the selected countermeasure.
Citation Information
Patent Citations
Stack management device, stack management method, and stack management program
JP2015170344A