Method and system for managing storage system
The management system improves prediction accuracy for storage system operations by using an analytical and statistical model to account for network load and resource utilization, effectively managing storage system operations.
Patent Information
- Application Number
- JP2024111326
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-10
- Publication Date
- 2026-01-23
AI Technical Summary
Existing technologies struggle to accurately predict the processing time for configuration operations in storage systems due to varying network conditions and insufficient performance resources, leading to unpredictable and inefficient management operations.
A management system that utilizes an analytical model and a statistical model to predict processing time by considering network load and resource utilization, combining historical data with real-time adjustments to improve accuracy.
Enhances the accuracy of predicting management operation times in storage systems, addressing the challenges of network variability and resource allocation inefficiencies.
Smart Images

Figure 2026011061000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention generally relates to the management of storage systems, and for example to a technique for predicting the time required for management operations performed on a storage system from a client using an API (Application Programming Interface). [Background technology]
[0002] In recent years, storage systems in Infrastructure as a Service (IaaS) clouds have been used to store daily business data and as a system recovery destination in the event of a disaster. For cloud storage systems, APIs (Application Programming Interfaces) are typically used to manage the resource configuration of cloud storage systems from remote clients via a communication network. Operations that can be performed using such APIs include management operations, which are at least some of the operations other than I / O operations (typically, inputting and outputting user data to and from volumes). Management operations include reference operations that obtain target configuration information and setting operations that change target configurations. While many reference operations are performed synchronously (requested operations are completed in a short time), setting operations are often performed asynchronously (requested configuration changes often require time to be reflected). When setting operations target storage system resources, configuration change operations such as creating volumes, creating paths connecting volumes to hosts, creating snapshots, and creating replications are examples of setting operations. Setting operations can, for example, create a configuration for remote copying (e.g., volume pairs, paths, etc.) in a storage system.
[0003] Some configuration operations executed through APIs are highly demanding, such as processes that involve changing the configuration of many resources or deleting volumes. Furthermore, the processing time required for configuration operations can vary significantly depending on the configuration of the communication network and the congestion of the servers through which the API passes. As such, the processing time required for configuration operations, especially those related to changing the configuration of resources using an API via a network, depends on a variety of factors and is difficult to predict.
[0004] Furthermore, because storage systems prioritize IO (Input / Output) operations (prioritizing processing according to IO requests for volumes), there are few resources allocated to the management system, which can result in problems with insufficient performance for management operations such as configuration operations.
[0005] A technology for predicting processing time using an API is disclosed in Patent Document 1. According to Patent Document 1, a configuration change processing operation executed by an API on the client side is requested, and the processing time until the processing is completed is managed as an actual value, and the processing time is predicted based on the actual value. [Prior art documents] [Patent documents]
[0006] [Patent Document 1] Japanese Patent Application Publication No. 2018-81431 Summary of the Invention [Problem to be solved by the invention]
[0007] According to the technology disclosed in Patent Document 1, the accuracy of prediction improves as the number of performance data increases sufficiently. However, it takes time to accumulate a sufficient number of performance data for use in prediction, and for APIs that are executed only infrequently or during periods with few performance data, the number of performance data is insufficient, resulting in insufficient prediction accuracy.
[0008] Furthermore, even if a sufficient number of actual performance data are accumulated, the prediction accuracy may be low for events that follow the behavior specific to the device rather than statistical properties.
[0009] Furthermore, in services like cloud computing, where an unspecified number of users connect via a general network, it is difficult to predict the length of processing time required for configuration operations due to external factors. For example, the communication load varies depending on the number of devices connected to the network connecting the client and the cloud system, and the amount of communication via the network by applications running on each device, and this communication load affects the length of processing time. [Means for solving the problem]
[0010] The management system of the storage system calculates a predicted value of the required time using an analytical model, which is a model of ideal operations performed in response to an operation request to the API, and calculates a predicted value of the required time using a statistical model, which is a model constructed based on statistics of the history of operations performed in response to an operation request to the API, and determines the predicted value of the required time from these predicted values and the weights of these models. [Effects of the Invention]
[0011] According to the present invention, it is possible to improve the accuracy of prediction of the time required for a management operation that is performed in a storage system in response to an operation request sent from an API for the management operation via a communication network. [Brief explanation of the drawings]
[0012] [Figure 1] 1 shows an example of the overall configuration of a computer system according to a first embodiment. [Figure 2] 1 shows an example of the configuration of a storage system 101 according to a first embodiment. [Figure 3] 1 shows an example of the configuration of a program 300 and a management table group 308 of the management node 103 according to the first embodiment. [Figure 4]10 illustrates an example of a process for predicting a time required for an operation requested by an API according to the first embodiment. [Figure 5] 3 illustrates an example of a model processing unit 302 according to the first embodiment. [Figure 6] 1 shows an example of a model definition table 309 according to the first embodiment. [Figure 7] 1 illustrates an example of a model coefficient table 310 according to the first embodiment. [Figure 8] 1 shows an example of a resource management table 800 according to the first embodiment. [Figure 9A] 1 illustrates a portion of an example of a required time history table 311 according to the first embodiment. [Figure 9B] 11 illustrates the rest of the example of the required time history table 311 according to the first embodiment. [Figure 10] 10 illustrates an example of a required time prediction method in the statistical model unit 304 according to the first embodiment. [Figure 11] 10 illustrates an example of a processing flow of the management node 103 according to the first embodiment. [Figure 12] 11 illustrates an example of a processing flow of the factor determination process (S1108) according to the first embodiment. [Figure 13] 10 illustrates an example of a sequence diagram for detecting a sign of a problem according to a second embodiment. [Figure 14] 11 shows an example of the configuration of a management table group 308 according to the second embodiment. [Figure 15] 14 illustrates an example of a problem sign threshold table 1401 according to the second embodiment. [Figure 16] 11 illustrates an example of a part of a required time history table 311 according to the second embodiment. [Figure 17] 10 shows an example of the configuration of a program 300 and a management table group 308 of a management node 103 according to a third embodiment. [Figure 18] 16 shows an example of a MAX value table 1603 according to the third embodiment. [Figure 19] 13 illustrates an example of a processing flow for adding a resource that is lacking according to the third embodiment. [Figure 20] 13 illustrates an example of a part of a required time history table 311 according to the fourth embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0013] In the following description, an "interface apparatus" may refer to one or more interface devices, which may be at least one of the following: One or more I / O (Input / Output) interface devices. The I / O (Input / Output) interface devices are interface devices to at least one of the I / O devices and a remote display computer. The I / O interface device to the display computer may be a communications interface device. The at least one I / O device may be a user interface device, for example, either an input device such as a keyboard and a pointing device, or an output device such as a display device. One or more communication interface devices. The one or more communication interface devices may be one or more homogeneous communication interface devices (e.g., one or more NICs (Network Interface Cards)) or two or more heterogeneous communication interface devices (e.g., an NIC and an HBA (Host Bus Adapter)).
[0014] In the following description, "memory" refers to one or more memory devices, typically a primary storage device. At least one memory device in the memory may be a volatile memory device or a non-volatile memory device.
[0015] In the following description, a "persistent storage device" refers to one or more persistent storage devices. A persistent storage device is typically a non-volatile storage device (e.g., an auxiliary storage device), and specifically, for example, an HDD (Hard Disk Drive) or an SSD (Solid State Drive).
[0016] In the following description, the term "storage device" may refer to at least one memory, including memory and persistent storage device.
[0017] In the following description, a "processor" refers to one or more processor devices. The at least one processor device is typically a microprocessor device such as a CPU (Central Processing Unit), but may also be another type of processor device such as a GPU (Graphics Processing Unit). The at least one processor device may be a single-core or multi-core. The at least one processor device may also be a processor core. The at least one processor device may also be a processor device in a broader sense, such as a hardware circuit (e.g., an FPGA (Field-Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit)) that performs part or all of the processing.
[0018] In the following description, data (information) that produces an output in response to an input may be described using expressions such as "xxx table." However, this data (information) may be data of any structure, or may be a learning model such as a neural network, genetic algorithm, or random forest that produces an output in response to an input. Therefore, an "xxx table" may be referred to as "xxx data." In the following description, one table may be divided into two or more tables, or all or part of two or more tables may be one table.
[0019] Furthermore, in the following description, functions are sometimes described using the expression "yyy unit." However, the functions may be realized by one or more computer programs executed by a processor, by one or more hardware circuits (e.g., FPGAs or ASICs), or by a combination thereof. When a function is realized by a program executed by a processor, the specified processing is performed using a storage device and / or an interface device, etc., as appropriate, and therefore the function may be at least a part of the processor. Processing described using a function as the subject may also be processing performed by a processor or a device having the processor. A program may be installed from a program source. The program source may be, for example, a computer from which the program is distributed or a computer-readable recording medium (e.g., a non-transitory recording medium). The description of each function is merely an example; multiple functions may be combined into one function, or one function may be divided into multiple functions.
[0020] In addition, in the following description, when describing elements of the same type without distinguishing between them, common reference symbols will be used, and when describing elements of the same type with distinction between them, reference symbols will be used. [Example]
[0021] 1 to 12, an embodiment of a storage management system will be described, which manages APIs executed for a storage system and predicts the time required for a process requested by the API to be completed in the storage system.
[0022] FIG. 1 is a diagram illustrating the overall configuration of a computer system according to a first embodiment.
[0023] The computer system comprises a storage system 101, a host computer 105, and a management client 107. The storage system 101 comprises two or more (or one) storage nodes 102, and a management node 103 that manages the storage nodes 102. In this embodiment, the storage system 101 is provided in a cloud 100. The cloud 100 may mean cloud computing, and is a system that allows use of computer resources such as CPUs, memory, and storage via a network.
[0024] The host computer 105 runs, for example, a service that forms the core of a business system, and various applications 106 running on the host computer 105 can request the storage system 101 to read / write user data (data that is read / written from / to a volume provided by the storage system 101). The host computer 105 may be a physical computer or a virtual computer.
[0025] The management client 107 can use the API 108 provided by the management node 103, and can obtain configuration information of the storage node 102 managed by the management node 103, and can create, change, or update the configuration of the storage node 102. The management client 107 can be a physical computer or a virtual computer. The API 108 can be used in any unit, such as an operation unit for creating a volume or connecting a path to a volume. The API 108 can be used according to instructions from a user of the management client 107, or can be used automatically according to a script.
[0026] The storage node 102, management node 103, host computer 105, and management client 107 are connected via a communication network 104. The communication network 104 may be configured by combining, for example, a LAN (Local Area Network), a WAN (Wide Area Network), etc.
[0027] The storage system 101 in this embodiment is, for example, a Software Defined Storage (SDS), and although an example of a system on the cloud 100 is shown, it may also be an on-premise system that is owned and operated within a user's company.
[0028] FIG. 2 is an example of the configuration of a storage system 101 according to the first embodiment.
[0029] The storage system 101 includes a storage node 102 and a management node 103, and is configured for the purpose of providing a non-volatile storage area to a host computer 105. The storage node 102 includes a storage controller 202 and hardware resources 207 (e.g., computer resources such as a CPU 29a, memory 29b, network I / F (Interface) 29c, and storage 29d). The management node 103 includes hardware resources 201 (e.g., computer resources such as a CPU 21a, memory 21b, network I / F 21c, and storage 21d) and software resources 200 such as programs that use the hardware resources 201 to operate user data and manage storage configurations. In this embodiment, the hardware resources 207 or 201 are primarily virtual computers on the cloud 100, but if the storage system 101 is an on-premise system, the hardware resources 207 or 201 may be physical computers.
[0030] In the management node 103, the software resource 200 performs API server processing 25 that receives processing requests from the management client 107 via the API 108, resource monitoring processing 26 that monitors the status of the hardware resources 207 of the storage node 102 via the storage controller 202, and time prediction processing 27 that predicts the time required from the request of the API 108 to the completion of processing in the storage system 101. The software resource 200 may be a program, and the above-mentioned processes 25 to 27 may be performed by the program being executed by the CPU 21a of the hardware resource 201.
[0031] The main software resource of the storage node 102 may be a program that controls the storage node 102. The storage controller 202 may be realized by the program being executed by a CPU 29a of the hardware resource 207. The storage controller 202 may configure disk drives as a redundant array of inexpensive disks (RAID), provide logical disk areas called volumes to the host computer 105, provide storage-related functions (for example, functions such as volume creation, duplication, or replication), and control the configuration of the storage node 102 and user data.
[0032] The resources of the management node 103, such as the CPU 21a, memory 21b, network I / F 21c, and storage 21d, are virtual computer resources allocated to processes for managing the system 101, such as changing the configuration of the storage system 101. The cloud 100 has a mechanism for changing the performance of virtual computers by selecting from multiple types of virtual computers, each of which specifies the performance of each resource, such as the CPU operating frequency, the number of cores, memory bandwidth and size, the type and size of storage (e.g., hard disk drive (HDD) or solid state drive (SSD)), and network bandwidth. This mechanism can be used to scale up the resources of the virtual computers. The cloud 100 also has a mechanism for scaling out the virtual computers of the management node 103 to increase resources and improve management performance by adding one virtual computer resource to the management node 103 and introducing a mechanism for distributing processing across multiple machines.
[0033] FIG. 3 shows the configuration of the program 300 and the management table group 308 of the management node 103 according to the first embodiment.
[0034] In the management node 103, when the CPU 21a executes the program 300, functions such as an API server unit 301, a model processing unit 302, an analytical model unit 303, a statistical model unit 304, a resource monitoring unit 305, a required time prediction unit 306, and a required time measurement unit 307 are realized. The management table group 308 includes a model definition table 309, a model coefficient table 310, and a required time history table 311. Examples of each process performed by executing the program 300 will be described later with reference to FIG. 4. The management table group 308 may be, for example, a database, and may be operated from the program 300 using SQL (Structured Query Language).
[0035] FIG. 4 illustrates a process for predicting the time required for an operation requested by an API according to the first embodiment.
[0036] First, the management node 103 receives a processing request from the management client 107 using the API 108 with parameters specified.
[0037] The management node 103 instructs the storage controller 202 to perform processing in accordance with the API parameters received by the API server unit 301. The parameters extracted by the model processing unit 302 are passed to the analytical model unit 303 and the statistical model unit 304. On the other hand, the model processing unit 302 obtains necessary parameters from a plurality of parameters specified by the API, according to a model definition table 309 of a management table group 308, to be passed to the analytical model unit 303 and the statistical model unit 304, which will be described later. An example of processing by the model processing unit 302 is shown in FIG. 5.
[0038] The analytical model unit 303 includes a required time calculation unit 401. The required time calculation unit 401 calculates the time required for processing each operation provided by the API 108 according to a given calculation method. The "analytical model" here may be a model that takes as input a quantity (e.g., explanatory variables) that determines the output required time (e.g., a target variable) and the quantitative relationship between them. For example, it may be a mathematical model or a machine learning model such as a neural network model. For example, as with general industrial products, the designer or manufacturer of a storage system can understand the relationship between the setting items of a setting operation (e.g., a change operation) and the expected required time from prior verification and operational results, and an analytical model is constructed as a model representing such relationship. However, because the predicted time does not match the exact expected value due to individual differences and usage conditions, the model coefficients in the analytical model are adjusted to absorb the error between the predicted time and the expected value. The required time calculation unit 401 calculates (obtains) the predicted required time by inputting a quantity related to the required time into the analytical model described above.
[0039] An example of the relationship between the required time and known quantities implemented in the analytical model is described below. For example, the required time Ta in the storage system 101 using the API for volume creation takes longer the more volumes are created and the larger the volume size. Therefore, the required time Ta is proportional to parameters such as the number of volumes and the volume size. Furthermore, the higher the load on the storage node 102, such as CPU utilization, memory usage, and network utilization, the longer the required time Ta. Therefore, the required time Ta is proportional to parameters related to the load on these storage nodes. Furthermore, since the management node 103 analyzes the processing of the API 108 and issues processing instructions to the storage controller 202 of each storage node 102, when the utilization rate of the management node 103 is high, processing takes longer and the required time Ta increases. The predicted time Ta using the analytical model can be expressed as the following equation: Ta = α × number of volumes × volume size + β0 × storage node CPU utilization rate + β1 × storage node memory utilization rate + β2 × storage node network utilization rate + γ0 × management node CPU utilization rate + γ1 × management node memory utilization rate + γ2 × management node network utilization rate (Equation 1)
[0040] The required time calculation unit 401 of the analytical model unit 303 calculates (predicts) the required time Ta based on the parameters extracted by the model processing unit 302, the model coefficient table 310, and parameters according to the status (operation rate / usage rate) of each resource acquired by the resource monitoring unit 305.
[0041] The initial values of the model coefficients α, β0 to β2, and γ represented by the model coefficient table 310 may be values calculated based on the results of performance evaluation in the environment assumed at the design stage of the storage system 101.
[0042] As described above, when using a wide area network environment such as the cloud 100, the network load changes depending on the number of devices communicating on the network, the amount of communication used by applications running on each device, etc. Since the network load differs depending on the actual system environment used, it is difficult to accurately predict the load before operation.
[0043] Therefore, the statistical model unit 304 also predicts the required time. The statistical model unit 304 includes a statistical model analysis unit 402 and a required time calculation unit 403. The statistical model analysis unit 402 extracts the history information of the target API from the required time history table 311, uses a statistical method to determine the correlation between the utilization rate and the required time of each resource, and creates a statistical model based on the correlation. Various well-known statistical methods, such as principal component analysis and cluster analysis, can be applied as the statistical method. Because the accuracy of statistical analysis varies depending on the amount of accumulated history information (typically, the number of accumulated actual values), the timing of model creation and update may be adjustable. More specifically, for example, the statistical model unit 304 may update the statistical model each time the API 108 is executed and history is recorded, or may accumulate history for a certain period, such as every other day, and then measure the accumulated number. If the accumulated number exceeds a set threshold, the statistical model may be updated. Various other model update timings are possible, but are not limited to these in this example. The required time calculation unit 403 acquires the current status of each resource from the resource monitoring unit 305, and calculates (predicts) the required time for the current resource status by using the statistical model (i.e., the correlation between the availability rate and required time for each resource) obtained by the statistical model analysis unit 402. An example of required time prediction using a statistical model will be described later with reference to FIG.
[0044] The required time prediction unit 306 determines the predicted time (predicted value of the required time) using the predicted time based on the analytical model and the predicted time based on the statistical model. More specifically, for example, the predicted time is the average of the predicted time based on the analytical model and the predicted time based on the statistical model. In this case, the predicted time is determined taking both the analytical model and the statistical model into consideration. Another method is to determine the predicted time by weighting each predicted time based on the analytical model and the statistical model. A characteristic of statistical models is that in situations where there is insufficient history accumulated (e.g., the number of histories is below a certain number), the numerical values may vary significantly and accuracy may be low. In such cases, the accuracy can be improved by increasing the weighting of the predicted time based on the analytical model and increasing the proportion of predicted times based on the analytical model. A more specific weighting value may be calculated by calculating the variance of past predicted times of the statistical model, and the larger the variance, the higher the weighting of the predicted time based on the statistical model, and the smaller the variance, the higher the weighting of the predicted time based on the statistical model.
[0045] Furthermore, the required time prediction unit 306 compares the calculated required time (predicted time) with the actual required time (measured time), and corrects at least one model (for example, the analytical model) of the analytical model and the statistical model using a method described later. This is expected to improve the prediction accuracy of the required time from the next time onwards.
[0046] FIG. 5 illustrates an example of processing by the model processing unit 302 according to the first embodiment.
[0047] The model processing unit 302 acquires parameters that affect the processing time in the storage system 101 for processing requested by each API 108. The API server unit 301 analyzes the parameters specified by the API 108 received from the management client 107. The model processing unit 302 acquires the necessary parameters using a model definition table 309. An example of the model definition table 309 is shown in FIG. 6. That is, the model definition table 309 has an entry for each API 108, and the entry has information such as an API name 601 and a model definition 602. The API name 601 is the name of the API. The model definition 602 is a list of parameters that are used in analytical models and / or statistical models out of the parameters specified by the API 108.
[0048] For example, the following is an example of when the createVolume API for creating a volume is executed from the management client 107. The API server unit 301 analyzes the API name and parameters specified by the API and acquires information in JSON format such as {“API_name”:“createVolume”, “vol_size”:“10GB”, “vol_num”:“50”, . . . , “requested_time”:“2023-12-12T10:05:00.00”}. When creating a volume, the size and number of volumes affect the processing time (required time) in the storage system 101. Therefore, in the model definition table 309, {“vol_size”, “vol_num”, “requested_time”} are specified as parameters (parameter items) to be acquired as the createVolume API. These parameters affect the required time in an analytical model; in a statistical model, “requested_time” is also a parameter required to identify the processing requested by the API. The parameters (parameter values) obtained are {“API_name”:“createVolume”, “vol_size”:“10GB”, “vol_num”:“50”, “requested_time”:“2023-12-12T10:05:00.00”}. The values obtained for each parameter are used in the analytical model and / or statistical model.
[0049] FIG. 7 illustrates an example of the model coefficient table 310 according to the first embodiment.
[0050] The model coefficient table 310 shows model coefficients 701 and their values 702 used to reduce the error in the required time predicted using an analytical model assumed for the storage system 101 .
[0051] FIG. 8 shows an example of a resource management table 800 according to the first embodiment.
[0052] The resource management table 800 has an entry for each virtual computer (storage node 102 or management node 103) in the storage system 101. The entry has information such as a computer ID 801 and operating rates / usage rates 802-804.
[0053] The computer ID 801 represents the ID of the storage node 102 or the management node 103. The operating rates / usage rates 802 to 804 include a CPU operating rate 802, a memory usage rate 803, and a network usage rate 804. The CPU operating rate 802 is for each core of the CPU 21a or 29a and represents the core operating rate. The memory usage rate 803 represents the usage rate of the memory 21b or 29b. The network usage rate 804 represents the usage rate of the network I / F 21c or 29c.
[0054] The resource monitoring unit 305 acquires information on the resource operation status (CPU operation rate, memory usage rate, network usage rate) of each virtual computer in the storage system 101. For each virtual computer, the status of each resource, such as the CPU and memory allocated to the virtual computer, is monitored, a history for a certain period is maintained, and the operation rate / usage rate of each resource is measured. The resource monitoring unit 305 is invoked when the analytical model unit 303 or the statistical model unit 304 calculates a predicted time, and acquires the operation rate / usage rate calculated within each virtual computer at that time and returns that value. In this example, the CPU operation rate, memory usage rate, and network usage rate are used, but information on measured values of latency and throughput of the storage system 101 assigned to each virtual computer may be acquired instead of or in addition to at least one of these. For example, in the analytical model, if latency is high and throughput is low, processing time increases. Therefore, by adding a term representing this to the above (Equation 1), a model that takes into account the impact of the performance of the storage system 101 is created. Similarly, by storing measurement information on the latency and throughput of the storage system 101 as history in the required time history table 311, and by constructing a statistical model that includes this information in the prediction of the required time, a model that takes into account the impact of the performance of the storage system 101 is created. This statistical model is one in which the latency and throughput items of the storage system 101 are added to (Equation 2) described below. This makes it possible to improve the accuracy of the predicted values in the analytical model and statistical model by taking into account the impact that the performance of the storage system 101 assigned to each virtual computer has on each API.
[0055] 9A and 9B show an example of the required time history table 311 according to the first embodiment.
[0056] The required time history table 311 has an entry for each API 108. The entry includes information such as a job ID 900, an API name 901, parameters 902, a start time 903, an end time 904, an actual measurement time 905, a predicted time 902, a management unit resource status 906, and a node resource status 907.
[0057] Job ID 900 represents the ID of the job. API name 901 represents the name of API 108. Parameters 902 represent a list of parameters (pairs of parameter items and parameter values) specified by API 108. Start time 903 represents the start time of an operation according to a request received by API 108. End time 904 represents the end time of an operation according to a request received by API 108. Actual measurement time 905 represents the required time (actual value) of an operation according to a request received by API 108, specifically, the time from the time represented by start time 903 to the time represented by end time 904. Predicted time 1702 represents the predicted time determined by the required time prediction unit 306. Management unit resource status 906 represents the availability / usage of resources of the management node 103 (CPU availability, memory usage, and network usage). The node resource status 907 indicates the availability / usage (CPU availability, memory usage, and network usage) of the resources of the storage node 102. The availability / usage values stored for each of the resource statuses 906 and 907 are statistical values (for example, average values) of the availability / usage over the period from the time indicated by the start time 903 to the time indicated by the end time 904. The required time history table 311 is used to predict the required time using a statistical model.
[0058] FIG. 10 illustrates an example of a required time prediction method in the statistical model unit 304 according to the first embodiment.
[0059] First, the statistical model analysis unit 402 acquires history information of the target API 108 from the required time history table 311. An example of the acquired history information is the table illustrated in Fig. 10. That is, there are three histories (entries) for an API name (for example, "API-A1"), and the information held by each history is information representing a set of parameters, management unit resource status (CPU operating rate, memory usage rate, and network usage rate of the management node 103), node resource status (CPU operating rate, memory usage rate, and network usage rate of the storage node 102), and actual measurement time of the required time.
[0060] The statistical model analysis unit 402 analyzes the correlation between the utilization rate and required time of each resource using a statistical method based on the acquired results (table) shown in FIG. 10. Various statistical methods can be applied, but here, an example using multiple regression analysis is shown. The statistical model analysis unit 402 sets the required time as the objective variable and the parameters specified by the API, the resource status of the management node 103 (CPU utilization rate, memory usage rate, and network usage rate), and the resource status of the storage node 102 (CPU utilization rate, memory usage rate, and network usage rate) as explanatory variables. For example, for a certain target API 108, the following (Equation 2) is constructed as an equation for the required time for an operation following a request to that API 108. Required time = A0 + A1 × number of volumes + A2 × volume size + A3 × CPU utilization rate of storage node 102 + A4 × memory utilization rate of storage node 102 + A5 × network utilization rate of storage node 102 + A6 × CPU utilization rate of management node 103 + A7 × memory utilization rate of management node 103 + A8 × network utilization rate of management node 103 (Equation 2)
[0061] Using the information in the required time history table 311, the statistical model analysis unit 402 calculates the partial regression coefficients of the multiple regression analysis of A0 to A8 using the least squares method. Using this formula 2, the statistical model analysis unit 402 predicts the required time for the target API 108 based on the specified parameters and the acquired resource status (operation rate / usage rate). The required time can also be predicted using other statistical methods, so the statistical method is not limited.
[0062] The accuracy of the statistical analysis of required time (prediction accuracy) varies depending on the number of entries stored in the required time history table 311 for each API 108. Therefore, the timing of creating and updating the statistical model may be adjustable. For example, the statistical model analysis unit 402 may update the statistical model when the API 108 is executed and a history (entry) is recorded once in the required time history table 311, or may accumulate history for a certain period, such as every other day, and then measure the accumulated number. The statistical model may be updated when the accumulated number exceeds a set threshold. The trigger for updating the history is not limited in this embodiment. The required time calculation unit 403 acquires the current status of each resource from the resource monitoring unit 305 and calculates the required time based on the correlation between the status (operation rate / usage rate) of each resource in the statistical model obtained by the statistical model analysis unit 402 and the required time.
[0063] FIG. 11 illustrates an example of a processing flow of the management node 103 according to the first embodiment.
[0064] The model processing unit 302 acquires designated parameters corresponding to the model definition 602 corresponding to the requested API 108 as parameters necessary for using the analytical model and the statistical model (S1101).
[0065] The resource monitoring unit 305 acquires the resource status (operation rate / usage rate) of the management node 103 and the storage node 102 (S1102).
[0066] The analytical model unit 303 acquires the model coefficients from the model coefficient table 310, and calculates the predicted time by inputting the acquired model coefficients, the parameters acquired in S1101, and the numerical values (resource status) acquired in S1102 into the analytical model (S1103).
[0067] The statistical model unit 304 inputs the parameters acquired in S1101 and the numerical values (resource status) acquired in S1102 into the statistical model to calculate the predicted time (S1104).
[0068] The required time prediction unit 306 determines the predicted time for the processing (operation) requested by the target API using the predicted time based on the analytical model calculated in S1103 and the predicted time based on the statistical model calculated in S1104 (S1105).
[0069] After the predicted time is determined in S1105, the API server unit 301 returns the predicted time determined in S1105 in response to a required time query from the management client 107. Furthermore, the required time measurement unit 307 actually monitors the completion of processing in the storage node 102 by polling, confirms whether the processing is complete or not, and, when the processing is completed, records the actual measured time of the processing, etc. in the required time history table 311 (S1106). Thereafter, the required time prediction unit 306 compares the predicted time determined in S1105 with the actual measured time (the actual measured value of the required time) and determines whether the actual measured time is longer than the predicted time by a certain threshold or more (S1107). If the determination result in S1107 is false (S1106: No), the processing flow ends. If the determination result in S1107 is true (S1107: Yes), the management node 103 performs a cause determination process in order to improve the accuracy of the required time prediction (S1108), and the process ends.
[0070] FIG. 12 illustrates an example of a processing flow of the factor determination process (S1108) according to the first embodiment.
[0071] For example, if any of the following timings is detected (S1201), the process proceeds. If the timing is not detected, a certain period of time is waited until the timing is detected, and if the timing is not detected within the certain period of time, the current cause determination process may be terminated. The resource monitoring unit 305 refers to the resource management table 800 (collection results) and identifies that the overall resource availability / usage is low (for example, the availability / usage of each resource is below a threshold). The resource monitoring unit 305 refers to the required time history table 311, identifies a time period with few API executions (for example, a time period when the frequency of API execution per unit time period is below a predetermined value), and determines that the current time period overlaps with the identified time period.
[0072] When the above timing is detected, for an API corresponding to S1106: Yes (an API for which the required time predicted by the analytical model deviates significantly from a certain threshold), the model processing unit 302 changes the parameters acquired in S1101 for that API within a certain range, the analytical model unit 303 predicts the required time using the analytical model using the changed parameters, and the required time measurement unit 307 actually measures the required time for the operation using the changed parameters (S1202). Note that in the example shown in FIG. 12, the volume size is changed to N and the number of volumes is changed to M, but both N and M may be values defined so as not to significantly affect other operations of the virtual computer (e.g., storage node) on which the API is executed. The parameter change may typically be a change to the parameter value of one of the parameter items and the parameter value. Whether to increase or decrease the parameter value as a parameter change may depend on the parameter item and the magnitude of the prediction error (the error in the predicted required time), and the extent to which the parameter value is changed may depend on the magnitude of the prediction error.
[0073] The required time prediction unit 306 compares the actual measured time obtained in S1202 with the predicted time based on the analytical model, and determines whether the comparison result (the relationship between the actual measured time and the predicted time) shows a tendency in accordance with the analytical model (S1203). An example of "the comparison result (the relationship between the actual measured time and the predicted time) shows a tendency in accordance with the analytical model" is that the prediction error (the difference between the actual measured time and the predicted time) is within a certain range.
[0074] For example, if the difference between the analytical model's predicted time and the actual measured time when 10 volumes with a volume size of 10 GB are created using an API for creating volumes is greater than a threshold, there is a possibility that the current performance of the storage system 101 has changed from that assumed by the current analytical model. Therefore, measurements are performed with a volume size of 10 GB and 20 volumes, or with a volume size of 20 GB and 10 volumes (S1202). The predicted time based on the analytical model at this time is compared with the actual measured time, and it is determined whether the difference is within a certain threshold (S1203).
[0075] If the determination result of S1203 is true (S1203: Yes), the analytical model technique is still usable, so the required time prediction unit 306 corrects the value of the model coefficient (value 702) used in the analytical model so that the predicted time of the analytical model matches the actual measured time (S1204).
[0076] If the determination result in S1203 is false (S1203: No), the required time prediction unit 306 estimates that the cause of the prediction error is something other than the analytical model, and increases the weighting of the statistical model in the required time prediction unit 306 (S1205). This makes it possible to improve the accuracy of the predicted time determined by the required time prediction unit 306. [Example]
[0077] The second embodiment will be described below, focusing mainly on the differences from the first embodiment, and explanations of the commonalities with the first embodiment will be omitted or simplified.
[0078] Some of the processes (operations) executed by the API 108 on the storage system 101 take a long time because they involve many changes to the resource configuration or high-load processes such as deleting volumes. Furthermore, the processing time may vary depending on the degree of congestion of the communication network 104 or the server (here, the management node 103) through which the API 108 passes. These problems are caused by insufficient performance or internal failures of the management node 103. In the second embodiment, a sign of the above-described problem is detected by determining whether the difference between the predicted value of the required time for processing executed by the API, calculated by the method in the first embodiment, and the measured value of the actual required time, is equal to or exceeds a specified threshold.
[0079] FIG. 13 illustrates an example of a flow of a sequence process for detecting a sign of a problem according to the second embodiment.
[0080] An API is requested from the management client 107 to the storage system 101 (S1301). In the storage system 101, the required time prediction unit 306 of the management node 103 determines the predicted time T for the requested API (S1302). The determined predicted time T is saved in the required time history table 311 together with the job ID of the API processing. At the same time, the management node 103 performs processing on the specified storage node 102 in accordance with the request. The management node 103 obtains the actual measured time A of the processing performed on the storage node 102 by monitoring the processing time actually performed on the storage node 102 by polling, or by receiving a completion notification from the storage node 102 (S1303).
[0081] When the management client 107 wants to obtain the required time for the processing executed by the API requested in S1301, it obtains the job ID managed by the management node 103 as the return value of the API requested in S1301. This allows the management client 107 to identify the job ID of the processing executed by the API using the job ID managed by the management node 103. By specifying this job ID and executing a predicted time acquisition API that requests a predicted time for the required time, the management client 107 obtains the predicted time T for the processing executed by the API corresponding to the job ID from the management node 103 (S1304). By obtaining the predicted time T, the management client 107 can predict the required time, making it easier to plan other tasks.
[0082] Next, a method for detecting a problem sign in the management node 103 will be described. The management node 103 has a problem sign flag 2001 (see FIG. 16) in the required time history table 311. If there is a possibility that a problem has occurred with API execution for each job ID, the management node 103 sets this problem sign flag to "1"; if there is no problem, the management node 103 sets it to "0." The management node 103 also acquires a threshold P for detecting a problem sign from the problem sign threshold table 1401 (S1305), and compares the acquired actual measurement time A with the sum of the predicted time T and the threshold P for detecting a problem sign. If the actual measurement time A exceeds the sum of the predicted time T and the threshold P (A>T+P), it is determined that a problem has occurred, and the management node 103 sets the problem sign flag 2001 for the target job ID to "1" (S1306). If the actual measurement time A does not exceed the sum of the predicted time T and the threshold P (A≦T+P), the management node 103 sets the problem sign flag 2001 of the target job ID to “0”.
[0083] The management client 107 can recognize whether a problem has occurred by executing a problem diagnosis API to check whether a problem has occurred and checking the problem prediction flag for the target job ID recorded in the required time history table 311 (S1307). This allows for early detection of the possibility of a problem occurring, enabling early countermeasures to be taken.
[0084] Another method for enabling the management client 107 to recognize the result of the problem sign flag is to return the return value of the API requested by the management client 107 in 1301, including the predicted time determined by the required time prediction unit 306. In this case, in the case of an asynchronous API, the return value is returned after the required time prediction unit 306 has calculated the predicted time. On the other hand, in the case of a synchronous API, the return value is not returned until the processing is complete, so it cannot be used as a prediction, but if the problem sign flag 2001 is set, it can be checked whether or not measures are being considered to improve the API for the next and subsequent executions of the API.
[0085] Another possible method is to notify the management client 107 of the predicted time determined by the required time prediction unit 306. For example, the management client 107 may have a mechanism for receiving push notification messages, and the management node 103 creates a push notification type message containing the calculated predicted value and the API job ID, and then pushes the created message to the management client 107. This allows the management client 107 to know the predicted value of the time required for processing executed by the requested API. As such, various methods for notifying the predicted time are possible, and the present application does not limit the methods.
[0086] FIG. 14 shows an example of the configuration of the management table group 308 according to the second embodiment.
[0087] In addition to tables 309 to 311, management table group 308 in FIG. 3 also includes problem sign threshold table 1401 in which a threshold value for the difference between the predicted time required for each API and the actual measured time is recorded.
[0088] FIG. 15 illustrates an example of the problem sign threshold table 1401 according to the second embodiment.
[0089] The problem sign threshold table 1401 has an entry for each API. The entry has information such as an API name 1501 and a threshold 1502. The threshold 1502 represents the threshold for the difference between the predicted time and the actual measured time. Since the frequency of performance problems varies depending on the processing content of the API, a threshold can be set for each API. The threshold 1502 may be set to a fixed value based on performance problems that have occurred in the past. For APIs where the predicted time that is determined is likely to vary, the threshold 1502 can also be set so that a problem can be detected if a certain percentage of the predicted time is exceeded.
[0090] Furthermore, in the second embodiment, if the predicted time exceeds a certain threshold and is longer than the actually measured time, the actually measured time may be too short and the processing that should actually be performed may not have been performed, so the required time measurement unit 307 may check whether the processing for the requested API has been completed (whether it is being executed). If the processing has been completed, the required time measurement unit 307 may perform a verification process to check whether the configuration using the requested API has been constructed. [Example]
[0091] The third embodiment will be described below, focusing mainly on the differences from the first and second embodiments, and explanations of the commonalities with the first and second embodiments will be omitted or simplified.
[0092] In the second embodiment, by checking the difference between the required time and the actual measured time for API execution, it is possible to predict the signs of some kind of problem occurring, even if the cause is unknown. One of the factors that causes the required time is a lack of resources for processing configuration changes in the storage system 101. In this case, it is possible to improve the required time by estimating the resource that is causing a performance shortage and expanding the insufficient resource. In the third embodiment, it is expected that the performance problem will be resolved by estimating the resource that is causing a performance bottleneck using the required time history table 311 and expanding the resource.
[0093] FIG. 17 shows an example of the configuration of the program 300 and the management table group 308 of the management node 103 according to the third embodiment.
[0094] In the management node 103, when the CPU 21a executes the program 300, functions such as a resource shortage estimation unit 1601 and a resource addition processing unit 1602 are realized in addition to the functions 301 to 307. Furthermore, the management table group 308 includes a MAX value table 1603 in addition to the tables 309 to 311.
[0095] The resource shortage estimation unit 1601 estimates which resource is the bottleneck when the actual required time is longer than the required time predicted from the history table 311. The resource shortage estimation unit 1601 estimates the bottleneck resource and the resource amount according to the processing flow of Fig. 19. The resource addition processing unit 1602 performs resource addition processing according to the resource amount of shortage estimated by the resource shortage estimation unit 1601. The MAX value table 1603 indicates the MAX value of the resource amount of the resources that can be used by the management client 107 among the resources of the storage system 101.
[0096] FIG. 18 shows an example of the MAX value table 1603 according to the third embodiment.
[0097] The MAX value table 1603 has an entry for each virtual computer, and the entry has information such as a computer ID, a CPU core MAX 1802, a memory usage rate MAX 1803, and a network usage rate MAX 1804.
[0098] The computer ID indicates the ID of the virtual computer. The CPU core MAX 1802 indicates the maximum value of the CPU operating rate of the CPU core. The memory usage rate MAX 1803 indicates the maximum value of the memory usage rate. The network usage rate MAX 1804 indicates the maximum value of the network usage rate.
[0099] The resources of the storage system 101 are used for various purposes, such as user data processing and configuration change processing. The available resource MAX value table 1603 can be used by the management client 107. That is, the MAX value table 1603 indicates the maximum value of the availability / usage rate of resources that can be used mainly for configuration change processing of the storage system 101. In the storage system 101, priority is given to processing requests (typically I / O requests) from the application 106 of the host computer 105, so there are cases where restrictions are placed on the resources used for requests from the management client 107. The MAX value table 1603 is a table of the MAX values of these restrictions. If resources are being used at or above the availability / usage rate listed in the available resource MAX value table 1603, it means that there is a shortage of resources.
[0100] FIG. 19 illustrates an example of a processing flow for adding a resource that is lacking according to the third embodiment.
[0101] The resource shortage estimation unit 1601 obtains the history of the same API (entries with the same API name 901) from the required time history table 311, compares the actual measured time (the actual value of the required time) with the predicted time 1702, and obtains the history in which the actual measured time is longer than the predicted time 1702 (S1901).
[0102] The resource shortage estimation unit 1601 calculates the contribution to the actual measured time (the contribution influenced by the availability / usage rate of each resource) (S1902). Here, the "contribution" refers to the proportion of how much each resource affects the required time. There are various possible methods for calculating the contribution, but the resource shortage estimation unit 1601 calculates the standard partial regression coefficient as the contribution. Here, one example is shown, but the method is not limited to this. For example, the resource shortage estimation unit 1601 performs multiple regression analysis and calculates the standard partial regression coefficient from the obtained partial regression coefficient. Take the volume creation API as an example. The objective variable is the required time, and the explanatory variables are parameters specified by the API (e.g., volume size: 10 GB, number of volumes: 50), the resource status of the management node 103 (CPU availability, memory usage, and network usage), and the resource status of the storage node 102 (CPU availability, memory usage, and network usage). For example, for a certain target API 108, the following (Equation 3) is constructed as an equation for the time required for an operation following a request to that API 108. Required time = B0 + B1 × number of volumes + B2 × volume size + B3 × CPU utilization rate of storage node 102 + B4 × memory utilization rate of storage node 102 + B5 × network utilization rate of storage node 102 + B6 × CPU utilization rate of management node 103 + B7 × memory utilization rate of management node 103 + B8 × network utilization rate of management node 103 (Equation 3)
[0103] The resource shortage estimation unit 1601 obtains the partial regression coefficients of the multiple regression analysis of B0 to B8 by the least squares method using the history acquired in S1901. That is, the resource shortage estimation unit 1601 calculates which explanatory variable has a greater impact on the required time, which is the objective variable, in cases where the predicted time is longer than the actual measured time. With respect to the obtained partial regression coefficients of the multiple regression analysis of B0 to B8, the resource shortage estimation unit 1601 calculates each standard partial regression coefficient by the formula: Standard partial regression coefficient = (Partial regression coefficient) × (Standard deviation of explanatory variable) ÷ (Standard deviation of objective variable). This standard partial regression coefficient represents the degree of contribution of each explanatory variable to the required time, which is the objective variable.
[0104] Next, the resource shortage estimation unit 1601 corrects the operating rate / usage rate of each resource based on the contribution calculated in S1902 (S1903). There are various methods for correction, but if it is assumed that more time is required, the impact of the operating rate / usage rate of a resource with a large contribution rate will be even greater. The resource shortage estimation unit 1601 sets the corrected value of the operating rate / usage rate of each resource as the virtual resource operating rate / usage rate by multiplying and normalizing the value by the contribution rate.
[0105] Next, the resource shortage estimation unit 1601 obtains the MAX value of the resources in the storage system 101 that can be used by the management client 107 from the MAX value table 1603 (S1904). The resource shortage estimation unit 1601 calculates the difference between the MAX value of the resources that can be used by the management client 107 and the operating rate / usage rate of each resource after correction in S1903 (S1905). For each resource, the calculated difference is an estimate of the surplus resource that would be expected if the resource were further used.
[0106] The resource with the smallest surplus calculated in S1905 is likely to become a resource bottleneck. The resource addition processing unit 1602 detects the resource with the smallest surplus calculated in S1905, and adds a resource in an amount corresponding to the surplus to the detected resource (S1906).
[0107] For example, if the following values are used: {Storage node CPU utilization: 20, Storage node memory utilization: 10, Storage node network utilization: 10, Management node CPU utilization: 40, Management node memory utilization: 20, Management node network utilization: 10}, and their respective contributions are {10, 10, 10, 40, 20, 10}, the adjustment value is estimated by increasing the utilization / utilization rate by the contribution rate. In other words, if the contribution rate is 10, it is considered a 10% increase, so it is multiplied by 1.1. Multiplying the resource utilization / utilization rate by {1.1, 1.1, 1.1, 1.4, 1.2, 1.1}, which is the contribution rate, results in an adjustment value of {22, 11, 11, 56, 24, 11}. Subtracting this from the maximum available resource value in Figure 18, {30, 20, 20, 60, 30, 20}, gives {8, 9, 9, 4, 6, 9}. This is the surplus of resources. The CPU utilization rate of the management node, which has the smallest surplus of "4," is likely to become a bottleneck due to a lack of resources.
[0108] In this way, it is possible to detect resources that may be causing a bottleneck in performance, and by enhancing the detected bottleneck resource, performance can be improved. Widely known methods of scaling up or scaling out can be used to enhance resources. In this example, the management node 103 and the storage node 102 are virtual computers that have resources allocated on the cloud 100, and an example of a method of scaling up or scaling out a virtual computer on the cloud 100 will be described. Many cloud services provide multiple virtual computer types with specified resource performance, such as the CPU operating frequency, number of cores, memory bandwidth and size, storage type and size (e.g., HDD (Hard Disk Drive) or SSD (Solid State Drive)), and network bandwidth. A mechanism allows the performance of the virtual computer to be changed by selecting one of these types. If a computer's resources are insufficient, scaling up is possible by changing to a computer type with a higher-performance CPU or memory. In S1906, CPU resources allocated to the management node 103 can be added. For example, if it is detected that the CPU utilization rate is insufficient, a computer type provided by the cloud service is selected that has the same specifications as the current virtual computer in terms of resources such as memory, storage, and network, but with more CPU cores than the current resources. In cloud services, using high-performance resources is generally more expensive, so in S1906, the minimum amount of resources is added, thereby reducing costs. If a computer type with a higher operating frequency is cheaper than increasing the number of cores, the computer type with the higher operating frequency can be selected before increasing the number of cores. This method makes it possible to achieve cost-effective scale-up by adding only the resources that are insufficient in terms of performance.
[0109] In this example, there was a shortage of a specific resource, but if S1906 detects that there are shortages of a wide range of resources, such as CPU, memory, or network, the scale-out method is used. For example, a virtual machine of the same type and with the same specifications as the current virtual machine is added to increase processing power.
[0110] By using this method, it is possible to analyze the history of cases where a process takes longer than expected, detect resource bottlenecks, and add resources to eliminate the bottlenecks. [Example]
[0111] The fourth embodiment will be described below, focusing mainly on the differences from the first to third embodiments, and explanations of the commonalities with the first to third embodiments will be omitted or simplified.
[0112] FIG. 20 illustrates an example of a part of the required time history table 311 according to the fourth embodiment.
[0113] Each entry in the required time history table 311 includes information such as a problem symptom flag 2001, an analytical model weight 2002, and a statistical model weight 2003 in addition to the information described with reference to FIG.
[0114] The problem sign flag 2001 is a flag indicating whether or not there is a problem sign, with "1" meaning that there is a problem sign and "0" meaning that there is no problem sign.
[0115] The analytical model weight 2002 represents the weight of the analytical model, and the statistical model weight 2003 represents the weight of the statistical model. In this embodiment, the sum of the two is 100.
[0116] By using an analytical model and a statistical model in combination, the accuracy of the predicted time is expected to be improved. The predicted time calculation using the analytical model and the predicted time calculation using the statistical model are performed in parallel, and the predicted time calculated using the analytical model may reflect the weight of the analytical model, and the predicted time calculated using the statistical model may reflect the weight of the statistical model. The predicted time (e.g., the average value of the two predicted times) may be determined based on the predicted time reflecting the weight of the analytical model and the predicted time reflecting the weight of the statistical model. Because the weight of the ideal operation (analytical model) and the weight of the actual operation (statistical model) are obtained, a reduction in the calculation cost or time of cause analysis is expected.
[0117] For example, suppose the difference between the predicted time and the actual measured time is greater than a predetermined value. In this case, if the weight of the analytical model is greater than the weight of the statistical model, since sufficient history has not yet been accumulated and the variance is large, the analysis can proceed based on the prediction that the cause is likely a sudden network delay or that some kind of configuration information change is taking time. For example, since the phenomenon is likely to occur during a time period when the communication network 104 is congested, the analysis can begin by investigating the network load during that time period to find the cause of the difference between the predicted time and the actual measured time being greater than a predetermined value. On the other hand, if the weight of the statistical model is greater than the weight of the analytical model, the weight of the predicted time based on the statistical model is greater than the predicted time based on the analytical model, and the statistical model has higher prediction accuracy than the analytical model, suggesting that network delays are constantly occurring. If the actual measured time is longer than the predicted time, it is possible that network delays are constantly occurring, but that sudden network delays are highly likely.
[0118] Furthermore, in a process of estimating a resource that is causing a performance bottleneck using the required time history table 311, when the actual measured time is longer than the predicted time, the model weights are used to estimate which resource is causing the bottleneck. This is expected to improve the accuracy of bottleneck estimation. For example, if the weight of the analytical model is higher than the weight of the statistical model, and the actual measured time is longer than the predicted time, there is a high possibility that network resources, which were an uncertain factor at the time of design, are insufficient. Therefore, if the estimated numerical values for the CPU and network are the same, it is possible to determine that network resources are insufficient and prioritize adding network resources. On the other hand, if the weight of the statistical model is higher than the weight of the statistical model, it is possible that the on-site network environment has been learned to a certain extent. In this case, when estimating the insufficient resources as described above, if the CPU and network resources are the same, CPU resources can be prioritized for addition. In this way, the weights of the analytical model and the statistical model can be used to prioritize the resources to be added.
[0119] Although several embodiments have been described above, these are merely examples for explaining the present invention, and the scope of the present invention is not limited to these embodiments. The present invention can be implemented in various other forms.
[0120] The above description can be summarized as follows, for example: The following summary may include supplementary explanations to the above description and explanations of modifications of the above embodiment.
[0121] A management system (e.g., management node 103) includes an interface device (e.g., network I / F 21c), a storage device (e.g., memory 21b and storage 21d), and a processor (e.g., CPU 21a) connected to the interface device and the storage device. The interface device communicates with a client (e.g., management client 107) that requests processing from one of a plurality of APIs (e.g., API 108) for management operations on a storage system (e.g., storage system 101) via a communication network. The storage device stores management data (e.g., management table group 308). The processor is connected to the interface device and the storage device.
[0122] The management data includes required time history data (e.g., required time history table 311). The required time history data is data that accumulates a history including actual measurement time, which is an actual measurement value of the time required for a management operation requested by an API, each time the management operation is performed.
[0123] The processor receives an operation request associated with a specified parameter for a target API, which is an API for which parameters are specified by a client among multiple APIs. The processor acquires resource loads (e.g., resource status including resource availability / usage) of at least the hardware resources (e.g., hardware resources 201 and / or 207) of the management system and / or storage system involved in the operation performed in response to the received operation request. For example, the management data may include data defining the hardware resources (e.g., nodes) involved in the operation requested by each API, and the "hardware resources involved in the operation performed in response to the received operation request" may be identified from this data. The processor inputs at least some of the specified parameters and the acquired resource load into an analytical model, which is a model of an ideal operation performed in response to an operation request for the target API, to calculate an analytically predicted time, which is a value predicted by the analytical model for the time required for the operation performed in response to the received operation request. The processor may input at least some of the parameters and the acquired resource load into a statistical model, which is a model constructed based on statistics of the history of operations performed in response to operation requests for the target API, among the required time history data, to calculate a statistical predicted time, which is a predicted value by the statistical model for the required time of the operation performed in response to the received operation request.The processor may determine the predicted time as a predicted value of the required time of the operation performed in response to the received operation request, based on the analytical predicted time, the statistical predicted time, and each weight of the analytical model and the statistical model.
[0124] This improves the accuracy of prediction of the required time for a management operation performed in the storage system in response to an operation request sent from the management operation API via a communication network. Note that if the first condition is met (e.g., if the weight of the analytical model is zero), the analytical predicted time need not be calculated. Also, if the second condition is met (e.g., if the weight of the statistical model is zero, or if a predetermined number or more of histories for the target API have not been accumulated in the required time history data), the statistical predicted time need not be calculated. Also, an analytical model and a statistical model may exist for each API, and the predicted time may be calculated using the analytical model and statistical model corresponding to the requested API. The analytical model and statistical model for each API may be stored in a storage device. Also, for example, the analytical model may be any of the above-mentioned (Equation 1) to (Equation 3).
[0125] The processor may perform a prediction accuracy determination (e.g., S1107) to determine whether a difference between a predicted time determined for an operation performed in response to a received operation request and an actual measured time required for the operation is equal to or greater than a threshold. If the result of the prediction accuracy determination is true, the processor may perform a factor determination process (e.g., S1108) for the target API. In the factor determination process, the processor may perform a factor determination (e.g., S1202 and S1203) to determine whether an analytical model is responsible for the difference between the predicted time and the actual measured time being equal to or greater than the threshold. Depending on the result of the factor determination, the processor may perform at least one of correcting the analytical model or the statistical model, and changing the weight of at least one of the analytical model and the statistical model. This is expected to further improve prediction accuracy.
[0126] The factor determination may be performed one or more times to determine whether a change in accordance with the analytical model exists by performing a process including changing a parameter value to be input into the analytical model from the parameter value of a specified parameter and calculating the difference between the predicted time calculated by inputting the changed parameter value into the analytical model and the actual measured time of the operation according to the changed parameter value. For example, the processor may perform S1202 one or more times to determine whether the change in the difference is a change in accordance with the analytical model, and this determination may be the factor determination. This is expected to improve the accuracy of determining whether the analytical model is a factor, and therefore, it is expected that the prediction accuracy will be improved by correcting the model and / or changing the weights according to the determination result.
[0127] If the result of the factor determination is true, the processor may correct the model coefficients used in the analytical model as a correction of the analytical model. This is expected to further improve prediction accuracy. Note that if the result of the factor determination is true, instead of or in addition to correcting the values of the model coefficients, the processor may perform at least one of the following for the target API: a correction other than changing the values of the model coefficients of the analytical model; a change in the weight of the analytical model; and a change in the weight of the statistical model. The weight change may be a change to relatively increase the weight of the analytical model, for example, by increasing the weight of the analytical model and / or decreasing the weight of the statistical model (the weight of the analytical model does not necessarily have to be higher than the weight of the statistical model).
[0128] If the result of the factor determination is false, the processor may relatively increase the weight of the statistical model, which is expected to further improve the prediction accuracy. Note that this weight change may be, for example, by decreasing the weight of the analytical model and / or increasing the weight of the statistical model (the weight of the statistical model does not necessarily have to be higher than the weight of the analytical model).
[0129] If the target API is an API whose determined predicted time tends to vary by more than a certain degree (for example, if the variation in the determined predicted time 1702 is greater than a certain degree), the threshold may be a value according to the product of the determined predicted time and a predetermined ratio (for example, "10% of the predicted required time"). This is expected to improve the accuracy of the prediction accuracy determination.
[0130] The processor may notify the client of the presence of a problem sign when the actual time required for an operation performed in response to a received operation request is equal to or greater than the sum of the predicted time determined for the operation and a predetermined threshold, thereby enabling the client to take measures in advance to prevent the problem from occurring.
[0131] The processor may receive a request to a predetermined API other than the target API after the request to the target API, set a return value included in a response to the request indicating that a problem is suspected, and return the response to the client. In this way, the client can be notified of the presence of a problem suspected by the response to the request from the API, making it easier to realize a notification of the presence of a problem suspected compared to realizing so-called push-type notifications.
[0132] If a prediction error occurs in which the actual time required for an operation performed in response to a received operation request is equal to or greater than the sum of the predicted time determined for that operation and a predetermined threshold, the processor may calculate the contribution of each of multiple types of resources included in the hardware resources of the management system and / or storage system to the prediction error, and calculate the required resource amount (e.g., the product of a value based on the contribution and the actual measured resource load) from the calculated contribution and the actual measured load of that resource. The processor may estimate the type of resource with the smallest difference between the maximum resource amount (e.g., MAX values 1802 to 1804) and the required resource amount as a bottleneck resource, and add the estimated bottleneck resource. This reduces the possibility of a prediction error occurring in the future.
[0133] The multiple types of resources may include a processor and a network. If the processor and the network are estimated as bottleneck resources, the following may be performed. In the former case, measures are taken to address the possibility that there is a shortage of network resources, which was an uncertain factor at the time of design. In the latter case, measures are taken to address the possibility that there is a shortage of processor resources because there is a possibility that the network environment has been learned to a certain extent. In this way, the possibility of a prediction error occurring in the future can be reduced. If the weight of the analytical model is higher than the weight of the statistical model, the processor adds network resources. If the weight of the statistical model is higher than the weight of the analytical model, the processor adds more preprocessor resources.
[0134] The management data may include model definition data representing definition parameters, which are predetermined parameters for each API as parameters that affect the time required for an operation performed in response to an operation request for that API. At least some of the parameters may be parameters that correspond to the definition parameters corresponding to the target API. In this way, the processor can identify parameters required as model input for each API.
[0135] The storage system may be a system including one or more virtual computers defined on a cloud (e.g., cloud 100) that are one or more storage nodes, and the storage system may include a management system (e.g., management node 103) that is a virtual computer separate from the one or more storage nodes. [Explanation of symbols]
[0136] 100: Cloud, 101: Storage system, 102: Storage node, 103: Device management unit
Claims
1. an interface device that communicates with a client that requests processing via a communication network using one of a plurality of APIs (Application Programming Interfaces) for management operations on the storage system; a storage device that stores management data; a processor connected to the interface device and the storage device; Equipped with The management data includes required time history data, The required time history data is data that is accumulated as a history including an actual measurement time, which is an actual measurement value of the required time for a management operation requested by an API, each time the management operation is performed, and The processor: receiving an operation request associated with a specified parameter for a target API, which is an API for which a parameter is specified in the client, among the plurality of APIs; acquire a resource load of at least a hardware resource involved in an operation performed in response to the received operation request among the hardware resources of the management system and / or the storage system; calculating an analytically predicted time, which is a predicted value by the analytical model for a required time for the operation to be performed in response to the received operation request, by inputting at least some of the specified parameters and the acquired resource load into an analytical model, which is a model of an ideal operation to be performed in response to the operation request for the target API; a statistical model constructed based on statistics of a history of operations performed in response to an operation request for the target API among the required time history data, and inputting at least some of the parameters and the acquired resource load into the statistical model; and calculating a statistical predicted time, which is a predicted value by the statistical model for the required time of the operation performed in response to the received operation request; determining a predicted time as a predicted value of a time required for an operation to be performed in response to the received operation request, based on the analytical predicted time, the statistical predicted time, and the weights of the analytical model and the statistical model; Management system.
2. The processor: performing a prediction accuracy determination, which is a determination as to whether or not a difference between the predicted time determined for the operation to be performed in response to the received operation request and an actual measured time required for the operation is equal to or greater than a threshold value; If the result of the prediction accuracy determination is true, a factor determination process is performed on the target API; In the factor determination process, the processor: performing a factor determination to determine whether the analytical model is responsible for the difference between the predicted time and the actual measured time being equal to or greater than the threshold; performing at least one of correcting the analytical model or the statistical model and changing a weight of at least one of the analytical model and the statistical model according to a result of the factor determination; The management system according to claim 1 .
3. The factor determination is to determine whether or not there is a change according to the analytical model by performing a process including changing a parameter value to be input to the analytical model from the parameter value of the specified parameter, and calculating a difference between a predicted time calculated by inputting the changed parameter value into the analytical model and an actual measured time of an operation according to the changed parameter value, at least once. The management system according to claim 2 .
4. If the result of the factor determination is true, the processor corrects the analytical model by correcting values of model coefficients used in the analytical model. The management system according to claim 2 .
5. If the result of the factor determination is false, the processor relatively increases the weight of the statistical model. The management system according to claim 2 .
6. When the target API is an API for which the determined prediction time tends to vary by more than a certain degree, the threshold value is a value according to the product of the determined prediction time and a predetermined ratio. The management system according to claim 2 .
7. the processor notifies the client that there is a problem sign when an actual measured time required for an operation performed in response to the received operation request is equal to or greater than the sum of the predicted time determined for the operation and a predetermined threshold value; The management system according to claim 1 .
8. the processor accepts a request made to a predetermined API other than the target API after the request made to the target API, sets the presence of a problem symptom as a return value included in a response to the request, and returns the response to the client; The management system according to claim 7.
9. When a prediction error occurs in which the actual measured time required for the operation performed in response to the received operation request is equal to or greater than the sum of the predicted time determined for the operation and a predetermined threshold, the processor: calculating a contribution of each of a plurality of types of resources included in the hardware resources of the management system and / or the storage system to the prediction error; calculating a required resource amount for each of the plurality of types of resources from the calculated contribution rate and the load of the resource as an actual measurement; The type of resource with the smallest difference between the maximum resource amount and the required resource amount is estimated as the bottleneck resource, Add resources to estimated bottlenecks, The management system according to claim 1 .
10. the plurality of types of resources include a processor and a network; When the processor and the network are estimated as the bottleneck resources, If the weight of the analytical model is higher than the weight of the statistical model, the processor adds resources to the network; If the weight of the statistical model is higher than the weight of the analytical model, the processor adds additional processor resources. The management system according to claim 9.
11. the management data includes model definition data representing definition parameters, which are predetermined parameters for each API as parameters that affect the required time for an operation performed in response to an operation request for the API; The at least some parameters are parameters corresponding to definition parameters corresponding to the target API. The management system according to claim 1 .
12. the storage system is a system including one or more virtual computers that are defined on a cloud and are one or more storage nodes, a virtual computer separate from the one or more storage nodes in the storage system is the management system; The management system according to claim 1 .
13. The storage system management system receiving an operation request associated with a specified parameter for a target API, which is an API for which a parameter is specified in a client, among a plurality of APIs (Application Programming Interfaces) for management operations on the storage system; acquire a resource load of at least a hardware resource involved in an operation performed in response to the received operation request among the hardware resources of the management system and / or the storage system; calculating an analytically predicted time, which is a predicted value by the analytical model for a required time for the operation to be performed in response to the received operation request, by inputting at least some of the specified parameters and the acquired resource load into an analytical model, which is a model of an ideal operation to be performed in response to the operation request for the target API; a statistical model, which is a model constructed based on statistics of a history of operations performed in response to an operation request for the target API, among required time history data that accumulates a history including actual measured times that are actual measured values of the time required for the management operation requested by the API each time the management operation requested by the API is performed, by inputting at least some of the parameters and the acquired resource load into the statistical model, and calculating a statistical predicted time, which is a predicted value by the statistical model for the time required for the operation performed in response to the received operation request; determining a predicted time as a predicted value of a time required for an operation to be performed in response to the received operation request, based on the analytical predicted time, the statistical predicted time, and the weights of the analytical model and the statistical model; A management method for doing this.
Citation Information
Patent Citations
Program components management apparatus and management method
JP2018081431A