Non-transitory machine-readable storage medium, computing system, and method
By calculating factor values and correlation coefficients to optimize the workload distribution, the distribution problem of the increased workload part on the request buckets of different characteristics is solved, and the performance prediction and resource utilization efficiency of the storage system are improved.
Patent Information
- Application Number
- CN202010522809.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-17
- Filing Date
- 2020-06-10
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2040-06-10
AI Technical Summary
The prior art is difficult to effectively determine and optimize the appropriate distribution of the increased workload portion on different characteristic request buckets, resulting in a degradation of storage system performance.
By calculating the factor value, the distribution of the operation amount of the increased workload part on the request buckets is determined on different characteristics. The correlation coefficient is calculated using the correlation function and normalized to generate a modified workload bucket to predict the performance of the storage system.
Improves the performance prediction accuracy and resource utilization efficiency of the storage system, and avoids performance degradation due to workload changes.
Smart Images

Figure CN112099732B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to data storage. Background Art
[0002] The configuration of a storage device can be used to store data. A requester can issue a request to access data, which may result in a workload being executed at a storage system that controls access to the storage device. Changes in the workload can affect the performance at the storage system. Summary of the Invention
[0003] According to one aspect of the present disclosure, a non-transitory machine-readable storage medium is provided. The non-transitory machine-readable storage medium includes instructions that, when executed, cause a computing system to perform the following operations: receiving an indication of an increased workload portion to be added to a workload of a storage system, the workload including buckets representing operations of different characteristics; calculating a factor value indicating the distribution of the amount of operations of the increased workload portion to the buckets representing the operations of the different characteristics based on the amount of the operations of the different characteristics in the workload; and distributing the amount of operations of the increased workload portion to the buckets representing the operations of the different characteristics according to the factor value.
[0004] According to another aspect of the present disclosure, a computing system is provided. The computing system includes: a processor; and a non-transitory storage medium storing instructions that can be executed on the processor to perform the following operations: receiving an indication of an increased workload portion to be added to a workload of a storage system, the workload being represented by a representation of buckets of input / output (I / O) operations of different characteristics; calculating a factor value indicating the distribution of the amount of I / O operations of the increased workload portion to the buckets representing the I / O operations of the different characteristics based on the amount of the I / O operations of the different characteristics in the workload; distributing the amount of the I / O operations of the increased workload portion to the buckets representing the I / O operations of the different characteristics according to the factor value; and generating a modified representation of the buckets of the I / O operations of the different characteristics in a modified workload, the modified workload including the workload and the increased workload portion.
[0005] According to another aspect of the present disclosure, a method executed by a system including a hardware processor is provided. The method includes: receiving an indication of an increased workload portion to be added to a workload of a storage system, the workload including buckets of operations representing different input / output (I / O) sizes; calculating, based on the amounts of operations of the different I / O sizes in the workload, a factor value indicating a distribution of the amount of operations of the increased workload portion to the buckets of operations representing different I / O sizes; distributing, according to the factor value, the amount of operations of the increased workload portion to the buckets of operations representing the different I / O sizes to form modified buckets of operations representing the different I / O sizes; and predicting a performance of the storage system based on the modified buckets of operations representing the different I / O sizes. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] Some implementations of the present disclosure are described with reference to the following drawings.
[0007] Figure 1 A histogram of buckets representing a workload according to some examples.
[0008] Figure 2 is a block diagram of a configuration including a host system, a storage system, and a workload distribution engine according to some examples.
[0009] Figure 3 is a flowchart of a workload distribution process according to some examples.
[0010] Figure 4 illustrates an example of calculating a rolling correlation coefficient according to some examples.
[0011] Figure 5 illustrates an example of a calculated normalized correlation coefficient according to some examples.
[0012] Figure 6 illustrates an example of a delta number of I / O for different buckets calculated based on the normalized correlation coefficient according to some examples.
[0013] Figure 7 illustrates an example of the number of buckets of I / O for different buckets of an updated workload that has been increased from an original workload according to some examples.
[0014] Figure 8 is a block diagram of a storage medium storing machine-readable instructions according to some examples.
[0015] Figure 9 is a block diagram of a computing system according to some examples.
[0016] Throughout the drawings, like reference numerals represent similar but not necessarily identical elements. These drawings are not necessarily to scale, and some portions may be exaggerated in size to more clearly illustrate the examples shown. Additionally, the drawings provide examples and / or implementations consistent with the specification; however, the specification is not limited to the examples and / or implementations provided in the drawings. Detailed Description
[0017] In the present disclosure, unless the context clearly indicates otherwise, the use of the terms "a," "an," or "the" is also intended to include the plural forms. Additionally, the terms "comprising," "including," "containing," "equipped with," "having," or "with" used in the present disclosure specify the presence of the recited elements, but do not preclude the presence or addition of other elements.
[0018] A "storage system" may refer to a platform including hardware and machine-readable instructions to implement the storage of data. The storage system may be implemented by using a combination of processing resources, memory resources, and communication resources.
[0019] The processing resources may include a processor, multiple processors, or a portion of a processor (e.g., one or more cores). The memory resources may include a memory device or multiple memory devices, such as dynamic random access memory (DRAM) and static random access memory (SRAM), etc. An example of the memory resources is a cache memory for temporarily storing data held in persistent storage. The communication resources may include a network interface controller (or a portion thereof) or a port for communicating over a network. In a further example, the storage system may include other resources, which include physical resources and / or virtual resources.
[0020] The storage system may include or be able to access a storage device (or multiple storage devices). A "storage device" may include persistent storage implemented by using one or more non-volatile storage devices, such as one or more disk-based storage devices (e.g., one or more hard disk drives (HDDs), etc.) and one or more solid-state storage devices (e.g., one or more solid-state drives (SSDs), etc.).
[0021] A requester may issue a request to access data stored in a storage device. Such a request results in a workload executed by the storage system. A "workload" may refer to a collection of activities in response to one or more requests from one or more requesters.
[0022] In some cases, it may be difficult to determine how an increase in workload can affect the performance of a storage system or various resources of the storage system. Heavy use of the resources of the storage system (e.g., processing resources, memory resources, and / or communication resources) can cause the overall performance of the storage system to degrade, which can cause a requester to experience slow data access speeds (increased latency).
[0023] A workload can be represented by a set of parameters that describe the corresponding characteristics of the workload. A "set of parameters" can refer to a set that contains a single parameter or a set that contains multiple parameters. Examples of parameters that describe workload characteristics can include the input / output (I / O) size of the corresponding requests associated with the workload, the count of the number of requests of a corresponding I / O size (I / O size), and the type of the corresponding requests.
[0024] Although example parameters for representing a workload are listed, it should be noted that in other examples, additional or alternative parameters can be used to represent the characteristics of the workload, such as the characteristics of the requests associated with the workload. A request is "associated with" a workload if the request submitted by a requester causes the workload (or a portion of the workload) to be executed.
[0025] A request of a specific I / O size refers to a request that accesses (reads or writes) data of a specific I / O size (e.g., data having a size of 4 kilobits (kb), 8 kb, 16 kb, 32 kb, etc.). The count of the number of requests of a specific I / O size refers to the amount of requests of a specific I / O size. The type of requests can include write requests (the first type of requests) and read requests (the second type of requests), etc.
[0026] A given workload can be associated with multiple requests of different characteristics (e.g., I / O size, request type, etc.). For example, a given workload can be associated with multiple requests of different I / O sizes.
[0027] Generally, a given workload can be represented by a set of "buckets", where each bucket represents the amount of requests of a corresponding different characteristic. For example, the first bucket can include a representation of the amount of requests of a first I / O size, the second bucket can include a representation of the amount of requests of a second I / O size different from the first I / O size, and so on. The different buckets representing a given workload can represent the corresponding amount of requests for corresponding intervals (such as time intervals). The different buckets representing a given workload can be represented in the form of a histogram (or any other representation form), which represents the amount of the buckets at the corresponding intervals (such as time intervals).
[0028] Figure 1An example histogram 100 showing buckets representing request volumes for different I / O sizes is presented. In the example histogram 100, the horizontal axis represents time and the vertical axis represents the number (volume) of requests. The histogram 100 includes a plurality of bars 102-1, 102-2, …, 102-N (N≥2). Each bar 102-i (i = 1 to N) is divided into a plurality of bar segments, and each bar segment represents a bucket representing the request volume for a corresponding I / O size (or other characteristic) at a specific time interval. In Figure 1 the example, different patterns of each bar segment represent corresponding different I / O sizes, such as 128 kb, 256 kb, and 512 kb. For example, bar 102-1 at time interval T1 is divided into bar segments 102-11, 102-12, and 102-13, which represent the corresponding buckets representing the request volumes for the corresponding different I / O sizes at time interval T1. Bar segment 102-11 represents the bucket representing the 128 kb request volume (i.e., the request volume for which data of 128 kb size is read or written) at time interval T1, bar segment 102-12 represents the bucket representing the 256 kb request volume at time interval T1, and bar segment 102-13 represents the bucket representing the 512 kb request volume at time interval T1. The length of each bar segment indicates the number of requests for the corresponding I / O size represented by the bar segment.
[0029] Although three I / O sizes are depicted in Figure 1 , it should be noted that different examples can represent requests for additional or alternative I / O sizes.
[0030] Similarly, bar 102-2 at time interval T2 is divided into corresponding bar segments 102-21, 102-22, and 102-23, which represent the corresponding buckets representing requests for different I / O sizes at time interval T2, and bar 102-N at time interval TN is divided into corresponding bar segments 102-N1, 102-N2, and 102-N3, which represent the corresponding buckets representing requests for different I / O sizes at time interval TN.
[0031] The requests represented by bars 102-1 to 102-N can be part of a given workload (or multiple workloads).
[0032] When simulating an increase in a given workload (e.g., an X% increase in the workload, which results in adding an increased workload portion on top of the given workload), it may be difficult to determine the appropriate distribution of the increased workload portion across buckets representing requests with different characteristics (e.g., different I / O sizes) to provide a useful simulation of such an increase in the workload. The "increased workload portion" includes the number of requests added to the given workload. For example, it may be difficult to determine how to allocate the requests of the increased workload portion to a first bucket for requests of a first I / O size and a second bucket for requests of a second I / O size, etc. In other words, if the increased workload portion includes Y requests (Y≥1), it may be difficult to accurately determine which portions of the Y requests are to be distributed across N buckets (N≥1).
[0033] According to some implementations of the present disclosure, techniques or mechanisms are provided to distribute the amount of I / O operations of the increased workload portion across buckets representing I / O operations with different characteristics (e.g., different I / O sizes). As used herein, a bucket of I / O operations "representing" a characteristic refers to a bucket of the amount of I / O operations representing the characteristic. I / O operations may refer to operations related to data access (reads and / or writes), and here, the I / O operations may be performed in response to requests. The distribution of the amount of I / O operations among the buckets is based on the calculation of factor values, which are based on the amounts of I / O operations with different characteristics in the workload (e.g., different sizes or different types of I / O operations). The calculated factor values indicate the distribution of the amount of I / O operations of the increased workload portion to the buckets of I / O operations with different characteristics. For example, a factor value may indicate that the amount of a first subset of the I / O operations of the increased workload portion is to be distributed to a first bucket of I / O operations of a first I / O size, and the amount of a second subset of the I / O operations of the increased workload portion is to be distributed to a second bucket of I / O operations of the first I / O size, and so on.
[0034] Figure 2 is a block diagram of an exemplary arrangement including a plurality of host systems 202-1 to 202-N, where N≥1. Although a plurality of host systems are depicted in Figure 2 note that in other examples, there may be only one host system.
[0035] Host system 202-1 includes application 204-1. Host system 202-N includes application 204-N. Although the applications are represented as examples of requesters that can issue requests to access data according to the Figure 2 example, note that in other examples, different types of requesters may issue requests to access data.
[0036] Data accessed by a request from a requester is stored by a storage device 208 of a storage system 210 (or storage systems 210). The storage device 208 may be implemented using disk-based storage devices (such as HDDs), solid-state memory devices (such as SSDs), other types of persistent storage devices, or combinations thereof. The storage device 208 may be configured as an array (or arrays) of storage devices.
[0037] Although Figure 2 the storage device 208 is shown as being part of the storage system 210, in other examples, the storage device 208 may be external to the storage system 210 but accessible by the storage system 210.
[0038] The storage system 210 includes a processor 212 and a cache memory 214. The processor 212 may execute a data access program (in the form of machine-readable instructions) that manages access to data stored in the storage device 208 in response to requests received from a requester based on the requester's access to a storage volume (or storage volumes). The cache memory 214 may be used to temporarily store data, such as write data to be written to the storage device 208.
[0039] The storage system 210 also includes various ports 216. A "port" may refer to a communication interface through which a host system 202-i (i = 1 to N) can access the storage system 210 via a network 218. Examples of the network 218 may include any or some combination of the following: a storage area network (SAN), a local area network (LAN), and a public network such as the Internet.
[0040] Each port 216 has a corresponding bandwidth available for communicating data in response to an access request from a requester. A port may refer to a physical port or a logical port.
[0041] Although not shown, the host systems 202-1 and 202-N may also each include ports for communicating via the network 218.
[0042] The ports 216, the processor 212, and the cache memory 214 are examples of resources of the storage system 210 that can be used to perform tasks associated with a workload executed in response to requests from corresponding requesters.
[0043] The ports 216 are examples of communication resources. The processor 212 is an example of a processing resource. The cache memory 214 is an example of a memory resource.
[0044] Figure 2 Also shown is a storage system workload distribution engine 220 that can perform storage system workload distribution processing (such as Figure 3 the processing 300 shown).
[0045] As used herein, an "engine" may refer to a hardware processing circuit, which may include any or some combination of a microprocessor, cores of a multi-core microprocessor, a microcontroller, a programmable integrated circuit, a programmable gate array, a digital signal processor, or other hardware processing circuits. Optionally, an "engine" may refer to a combination of a hardware processing circuit and machine-readable instructions (software and / or firmware) executable on the hardware processing circuit.
[0046] The storage system workload distribution engine 220 may be part of the storage system 210 or, optionally, may be separate from and coupled to the network 218.
[0047] Further reference is made below in connection with Figure 2 described Figure 3 . The workload distribution process 300 includes (at 302) receiving an indication of an increased workload portion of the workload to be added to the storage system. The workload includes buckets representing operations of different I / O sizes. Examples of buckets representing operations are buckets representing requests of corresponding I / O sizes.
[0048] (The indication received at 302), e.g., Figure 2 233 in
[0049] In a further example, there may be multiple user consoles 230.
[0050] As an example, the UI 232 may present an option that allows a user at the user console 230 to specify a selected percentage increase of a given workload, such as a 10% increase or some other percentage. The user may specify a percentage increase of a given workload to determine the impact on the resources of the storage system 210 due to the increased workload.
[0051] In other examples, the indication of the increased workload may be provided by a different entity, such as a machine or a program.
[0052] The workload distribution process 300 calculates factor values (at 304) based on the amounts of operations of different I / O sizes in the workload, which indicate the amounts of operations of the increased workload portion to be distributed to the buckets representing operations of different I / O sizes. It may be based on a correlation function (e.g., Figure 2a 240) or another type of function to calculate the factor value. In an example using a correlation function, the calculated factor value is based on the correlation coefficient calculated by the correlation function 240. A correlation function may refer to a function that calculates the statistical correlation between variables based on the spatial or temporal distance between the variables. In some examples of the present disclosure, the variables associated by the correlation function 240 may include the amount of operations in the buckets of the operations and the total amount of operations (discussed in further detail below).
[0053] The workload distribution process 300 further distributes, according to the factor value (at 306), the amount of operations of the increased workload portion into the buckets representing operations of different I / O sizes to form modified buckets representing operations of different I / O sizes. The distribution of the amount of operations of the increased workload portion includes distributing a first amount of operations of the increased workload portion into a first bucket representing operations of a first I / O size, distributing a second amount of operations of the increased workload portion into a second bucket representing operations of a second I / O size, and so on.
[0054] The workload distribution process 300 predicts (at 308) the performance of the storage system 210 based on the modified buckets representing operations of different I / O sizes. As discussed further below, the prediction may be based on using a prediction model (e.g., Figure 2 a 242) that receives the modified buckets representing operations of different I / O sizes as input.
[0055] As described above, in some examples, the factor value used to distribute the amount of operations of the additional workload portion into the corresponding buckets is based on the correlation coefficient calculated for each of the buckets representing I / O operations at respective time intervals (such as Figure 1 the time intervals T1 to TN in). For example, in Figure 1 an example, at each time interval Ti (i = 1 to N), three buckets (represented by the respective three segments of bar 102-i) representing I / O operations according to different characteristics may be associated with respective variables X1, X2, and X3. Here, X1 may represent the amount of I / O operations (e.g., the request volume) of a first I / O size (e.g., 128 kb), X2 may represent the amount of I / O operations (e.g., the request volume) of a second I / O size (e.g., 256 kb), and X3 may represent the amount of I / O operations (e.g., the request volume) of a third size (e.g., 512 kb). Given the variables X1, X2, and X3 for each time interval Ti, the following correlation functions can be calculated:
[0056] CC1 = corr(X1, Xt), (Equation 1)
[0057] CC2 = corr(X2, Xt), (Equation 2)
[0058] CC3 = corr(X3, Xt), (Equation 3) Here, Xt represents the total amount of I / O operations (e.g., requests) at each time interval Ti; in fact, at time interval Ti, Xt = X1 + X2 + X3. In the above Equations 1 - 3, corr() is Figure 2 the correlation function 240 of
[0059] For each time interval T1,..., TN, the correlation coefficients CC1, CC2, and CC3 are calculated. In other words, at each time interval Ti, in an example where there are three buckets in each time interval, three correlation coefficients CC1, CC2, and CC3 are calculated.
[0060] As further explained below, the storage system workload distribution engine 220 can calculate a factor value for distributing an additional workload portion based on the correlation coefficients.
[0061] Figure 4 Shows different examples of four buckets A, B, C, and D of I / O operations with different I / O characteristics in each time interval. Figure 4 Shows two tables, Table 1 and Table 2. Each row of Table 1 and Table 2 represents a corresponding time interval (rows 1 - 10 in Table 1 and Table 2 represent 10 different time intervals).
[0062] In Table 1, the bucket A column includes entries containing values representing the amount of I / O operations with the first I / O characteristic in the workload, the bucket B column includes entries containing values representing the amount of I / O operations with the second I / O characteristic in the workload, the bucket C column includes entries containing values representing the amount of I / O operations with the third I / O characteristic in the workload, and the bucket D column includes entries containing values representing the amount of I / O operations with the fourth I / O characteristic in the workload.
[0063] In the following discussion, the amount of I / O operations in the entries of the bucket columns can be referred to as the "number of buckets of I / O". For example, the number of buckets of I / O for time interval 6 in the bucket A column of Table 1 is 1565.
[0064] In a further example, the correlation function 240 applied to calculate the corresponding correlation coefficients is a rolling coefficient function calculated based on values in a rolling window including multiple time intervals. In Figure 4 the example of
[0065] In the Bucket A column of Table 1, window 402 includes rows 1 - 4 (corresponding to time intervals 1 - 4) of the number of four corresponding buckets of IOs that contain the first I / O characteristic corresponding to Bucket A. In the Bucket A column, row 1 includes the value 575, which indicates that there are 575 I / O operations with the first I / O characteristic in time interval 1. Similarly, in rows 2 - 4 of the Bucket A column, the values 490, 3185, and 600 indicate that there are 490 I / O operations with the first I / O characteristic in time interval 2, 3185 I / O operations with the first I / O characteristic in time interval 3, and 600 I / O operations with the first I / O characteristic in time interval 4.
[0066] The number of four buckets of IOs in rows 1 - 4 of the Bucket A column is part of a window of value (402). Figure 4 Another window 410 showing the values in the Bucket B column, where window 410 includes the values of rows 2 - 5 indicating the number of corresponding buckets of IOs of the second I / O characteristic.
[0067] The Total_IOs column of Table 1 represents the total number of I / O operations for each time interval (or more simply, "Total IOs"). Thus, in row 1, the value in the Total_IOs column is 18208, which is the sum of the values in row 1 of the Bucket A column, Bucket B column, Bucket C column, and Bucket D column.
[0068] The window 404 of the values in the Total - IOs column includes the total number of IOs in the corresponding time intervals 1, 2, 3, and 4. Another window 412 of the values in the Total - IOs column includes the total number of IOs in the corresponding time intervals 2, 3, 4, and 5.
[0069] In Table 2, the entries in the Rolling_corr(A, Total) column include the corresponding rolling coefficient values calculated based on the corresponding different windows in the Bucket A column of Table 1, the entries in the Rolling_corr(B, Total) column include the corresponding rolling coefficient values calculated based on the corresponding different windows in the Bucket B column of Table 1, the entries in the Rolling_corr(C, Total) column include the corresponding rolling coefficient values calculated based on the corresponding different windows in the Bucket C column of Table 1, and the entries in the Rolling_corr(D, Total) column include the corresponding rolling coefficient values calculated based on the corresponding different windows in the Bucket D column of Table 1.
[0070] The rolling correlation coefficient values in the Rolling_corr(A, Total) column (e.g., 0.34 in entry 406 at time interval 4) are calculated based on the following:
[0071]
[0072] Here, Rolling_CCA represents the rolling coefficient value, w represents the window size (in the example of Figure 4 , the window size is 4 time intervals), XA j represents the number of buckets of IOs in row j of the window in column A of bucket in Table 1, and Xt j represents the total number of IOs in row j of the window in column Total_IOs of Table 1. In the example of Figure 4 , each window includes 4 rows, and row j of the window refers to one of the 4 rows in the window. Thus, for example, in window 410 in column B of bucket Figure 4 , row 1 of window 410 corresponds to time interval 2, row 2 of window 410 corresponds to time interval 3, and so on.
[0073] Based on the number of buckets of IOs in window 402 in column A of bucket in Table 1, the rolling correlation coefficient value in entry 406 of Table 2 is calculated. In addition, based on the number of buckets of IOs in another window including the number of buckets of IOs in time intervals 2 - 5 in column A of bucket in Table 1, the rolling correlation coefficient value in entry 414 of Table 2 is calculated, and based on the number of buckets of IOs in another window including the number of buckets of IOs in time intervals 3 - 6 in column A of bucket in Table 1, the rolling correlation coefficient value in entry 416 of Table 2 is calculated, and so on.
[0074] For bucket B, according to the following formula, based on the number of buckets of IOs in window 410 in column B of bucket in Table 1 and the total number of IOs in window 412 in column Total_IO of Table 1, the rolling correlation coefficient value Rolling_CCB in entry 408 of column Rolling_corr(B, Total) is calculated:
[0075]
[0076] Here, XB j represents the number of IOs in row j of the window in column B of bucket.
[0077] The rolling correlation coefficient values for buckets C and D are calculated in a similar manner.
[0078] In Table 2, since the rolling correlation coefficient value is based on a window of four values in the example, rows 1 - 3 of Table 2 are blank.
[0079] A positive rolling correlation coefficient value indicates that the sum of the number of buckets of IOs in the corresponding bucket window (e.g., 402) and the sum of the total number of IOs in the corresponding total IO window (e.g., 404) are positively correlated. A negative rolling correlation coefficient value indicates that the sum of the number of buckets of IOs in the corresponding bucket window and the sum of the total number of IOs in the corresponding total IO window are negatively correlated.
[0080] In some examples, normalization of the rolling correlation coefficient values is performed. The normalization includes setting any negative rolling correlation coefficient values to zero (or some other predetermined value), and scaling the positive rolling correlation coefficient values to generate a normalized correlation coefficient value Normalized_CC according to the following formula:
[0081]
[0082] In Equation 6, if the rolling correlation coefficient value is negative, max(0, CC) outputs 0, and if it is positive, it outputs the rolling correlation coefficient value CC.
[0083] In addition, Equation 6 scales the positive rolling correlation coefficient value CC based on the ratio of the sum of the number of buckets of IOs in each window to the sum of the total number of IOs in each window.
[0084] If the number of buckets of IOs in the first bucket and the number of buckets of IOs in the second bucket are both roughly in the same direction of movement (increase or decrease in value) as the total amount of IOs in the Total_IOs column, the correlation coefficient values calculated for the first and second buckets may have similar magnitudes, even if one of the two buckets has a smaller number of buckets of IOs that contributes less to the workload. For example, in Table 2, compared to the correlation coefficient value of bucket A, the correlation coefficient value of bucket C has a relatively large value, even though bucket A generally has a larger number of buckets of IOs than bucket C. Therefore, the relatively large correlation coefficient value of bucket C does not accurately represent the contribution of the I / O operations of bucket C relative to the contribution of the I / O operations of bucket A.
[0085] The scaling performed according to Equation 6 takes into account the relative number of buckets of IOs in different buckets, such that due to the smaller number of buckets of IOs in bucket C, the scaling reduces the correlation coefficient value calculated for bucket C.
[0086] After normalization using Equation 6, the generated normalized correlation coefficient values are as shown in Figure 5 Table 3. The normalized correlation coefficient value of bucket C in column 502 of Table 3 is much lower than the corresponding normalized correlation coefficient value of bucket A in column 504 of Table 3. Although the normalized correlation coefficient value of bucket C in column 502 has a value of 0.00, it should be noted that this is due to rounding to 2 decimal places.
[0087] Once the normalized correlation coefficient values are derived, for example, according to Equation 6, the number of IOs added to each bucket at each corresponding time interval can be derived according to Equation 7 below:
[0088]
[0089] Here, Delta_IO_A represents the increment of IO of the increased workload portion (represented by Increased_Workload) to be added to bucket A, Normalized_CCA represents the normalized correlation coefficient of bucket A normalized according to Equation 6, and
[0090]
[0091] represents the sum of the normalized correlation coefficient values Normalized_CCk (k = A, B, C, D) of buckets A, B, C, and D within the corresponding time intervals.
[0092] In Figure 6 Table 4, the increments of IO calculated for each bucket in each time interval are provided. In Table 4, the DeltaIOs A column represents the increment of IO added to bucket A in the corresponding time interval, the Delta IOs B column represents the increment of IO added to bucket B in the corresponding time interval, the Delta IOs C column represents the increment of IO added to bucket C in the corresponding time interval, and the Total Delta IOs column contains the value representing the total number of Delta IOs added in each time interval to buckets A, B, C, and D. For example, in row 4 of Table 4, for time interval 4, the total number 1670 of delta IOs is based on the sum of the increments of IO of buckets A, B, C, and D in the Delta IOs A column, Delta IOs B column, Delta IOs C column, and Delta IOs D column respectively for time interval 4.
[0093] Figure 7 Table 5 in Figure 4 shows the updated quantities of IO of the updated workload (the original workload increased by the increased workload portion) in the corresponding buckets A, B, C, and D, which is based on adding the increment of IO of each bucket at each time interval to the corresponding original bucket numbers of the IO of the workload in
[0094] Table 1. In Table 5, the Updated A column contains the updated quantity of IO of the updated workload in bucket A, the Updated B column contains the updated quantity of IO of the updated workload in bucket B, the Updated C column contains the updated quantity of IO of the updated workload in bucket C, the Updated D column includes the updated quantity of IO of the updated workload in bucket D, and the Updated Total IOs column includes the total number of updated IOs.
[0095] Once the updated number of IOs in the respective buckets of operations with different I / O characteristics in the updated workload is calculated, the updated number of IOs (in histogram form or other representation) can be provided as an input to a prediction model 242, such as a regression model that predicts latency, saturation, or other operational metrics of the storage system 210. The operational metrics can be used to indicate the performance and / or usage of the storage system 210, or one or more resources of the storage system 210. The latency of the storage system 210 or a resource of the storage system 210 can refer to the amount of time required for the storage system 210 or the resource to execute a workload. The saturation of the storage system 210 or a resource can refer to the amount of the storage system 210 or the resource consumed by the workload. In other examples, other operational metrics can be used.
[0096] The prediction model 242 can be trained by using training data based on any of various machine learning techniques. Optionally, or in addition, when the prediction model performs calculations, the prediction model 242 can self-update and provide feedback (by a person, machine, or program) regarding the accuracy of the calculations performed by the prediction model.
[0097] An administrator (e.g., using Figure 2 the user console 230) can select a past period of time as the analysis period during which the user may wish to evaluate the impact of workload changes on the performance of the storage system 210.
[0098] Representations of the updated number of IOs in the respective buckets of operations with different I / O characteristics in the updated workload are input into the prediction model 242 to estimate new predicted values of the operational metrics. For example, the percentage change in latency relative to the original latency can be calculated. When there are expected changes in the workload during a particular period of time, this change in latency can be provided as an insight into the expected change in the behavior of the storage system 210 in terms of latency. This can help determine whether the expected latency will violate a threshold that may be unacceptable to the enterprise, or whether the expected latency is within an acceptable threshold range due to changes in the workload.
[0099] Similarly, representations of the updated number of IOs of the updated workload can be fed into the prediction model 242 to generate an estimated saturation of the storage system 210 or a resource to determine whether the estimated saturation will violate a threshold. In other examples, the prediction model 242 can be used to estimate the values of other operational metrics.
[0100] The storage system workload distribution engine 220 can send a report 228 for viewing in the UI 232 of the user console 230 ( Figure 2)。In some examples, the report 228 may include estimated operation metrics generated by the prediction model 242. In further examples, the report 228 may also include information on the increment of I / O (such as those in Figure 6 ) to be added to the corresponding buckets of the updated workload and / or the updated quantity of I / O in the corresponding buckets of the updated workload (such as those in Figure 7 ).
[0101] Based on the report 228, the enterprise can take actions to address the expected increase in the workload. If one or more operation metrics generated by the prediction model 242 indicate that the storage system 210 or resources for executing the updated workload are expected to operate within the target specifications (e.g., the estimated latency or saturation is below the specified threshold), the enterprise can simply allow the storage system 210 to continue operating.
[0102] However, if one or more operation metrics generated by the prediction model 242 indicate that the storage system 210 or resources for executing the updated workload are expected to operate outside the target specifications (e.g., the estimated latency or saturation exceeds the specified threshold), the enterprise can take actions to update the storage system 210 (e.g., by adding resources or updating resources). As another example, the enterprise can choose not to deploy the updated workload on the storage system 210 and instead can deploy the updated workload on another storage system.
[0103] Optionally, the enterprise can also take other actions, such as any one or some combination of the following actions: the use of resources currently generated by the workload (such as by reducing the rate of submitting data requests to the storage system for the set of storage volumes corresponding to the workload); configuring quality of service (QoS) settings for the workload type, where the QoS settings can affect the priority of the workload in resource usage; changing the resource distribution for the workload (such as by changing the amount of resources allocated for data access to data for processing the workload of a given workload type); migrating the data of the workload from a first set of storage devices to a different second set of storage devices; and so on.
[0104] One or more actions taken by the enterprise in response to one or more operation metrics generated by the prediction model 242 can be performed by a person or an automated system. The automated system can include the Figure 2 resource management engine 250 in. The resource management engine 250 can automatically perform any of the above actions in response to one or more operation metrics generated by the prediction model 242.
[0105] Rather than calculating a predicted change in an operational metric of storage system 210, a predicted change in an operational metric (e.g., latency or saturation) can be calculated for a specific storage volume of storage system 210 or a subset of storage volumes of storage system 210.
[0106] A "storage volume" can refer to a collection of data stored in a storage device or multiple storage devices (or portions thereof) that can be managed as a unit. A storage volume is a logical entity for storing data. A storage volume can be presented to a host system capable of reading and writing data to the storage volume (e.g., Figure 2 202-1 to 202-N in). For example, a storage volume can be exported by a storage system to a host system for use by the host system. More generally, a storage volume can be available to a host system such that the host system (such as an application in the host system) can access the data in the storage volume.
[0107] Figure 8 is a block diagram of a non-transitory machine-readable or computer-readable storage medium 800 storing machine-readable instructions that cause a system (a computer or multiple computers) to perform a corresponding task when executed.
[0108] The machine-readable instructions include an instruction 802 for receiving an indication of an increased workload portion to be added to a workload of a storage system, the workload including buckets of operations representing different characteristics (e.g., different I / O sizes, different I / O types, etc.). In some examples, the indication of the increased workload portion is based on an indication that the workload is to be increased by a specified percentage or the number of operations by which the workload is to be increased, etc.
[0109] The machine-readable instructions further include a factor value calculation instruction 804 for calculating factor values based on the amounts of operations of different characteristics in the workload, the factor values indicating the distribution of the amounts of operations of the increased workload portion to the buckets of operations representing different characteristics.
[0110] In some examples, the factor values are based on correlation coefficients. For example, a first factor value is calculated based on (using a correlation function) correlating the amount of operations of a first characteristic in the workload with the total amount of operations in the workload, a second factor value is calculated based on (using a correlation function) correlating the amount of operations of a second characteristic in the workload with the total amount of operations in the workload, and so on.
[0111] The calculation of the factor values further includes normalizing correlation values generated by correlating the amounts of operations of each characteristic within each time interval with the total amount of operations in the workload within each time interval of the workload. The normalization of the correlation values can be based on the ratio of the amount of operations of the corresponding characteristic to the total amount of operations.
[0112] The machine-readable instructions further include an operation distribution instruction 806 that distributes the amount of operations of the increased workload portion to buckets of operations of different characteristics according to factor values.
[0113] Each corresponding bucket representing an operation includes the amount of operations representing the corresponding characteristic, and distributing the amount of operations of the increased workload portion to the buckets representing operations causes an increase in the amount of operations in the buckets of operations. Distributing the amount of operations of the increased workload portion to the buckets representing operations, a first factor value of the factor value causes a first increase in the amount of operations in the first bucket representing operations, a second factor value of the factor value causes a second increase in the amount of operations in the second bucket representing operations, and so on.
[0114] Figure 9 is a block diagram of a computing system 900 (implemented as a computer or multiple computers) including a hardware processor 902 (or multiple hardware processors). The hardware processor may include a microprocessor, a core of a multi-core microprocessor, a microcontroller, a programmable integrated circuit, a programmable gate array, a digital signal processor, or other hardware processing circuits.
[0115] The computing system 900 further includes a storage medium 904 that stores machine-readable instructions executable on the hardware processor 902 to perform various tasks. Machine-readable instructions executable on the hardware processor may refer to instructions executable on a single hardware processor or instructions executable on multiple hardware processors.
[0116] The machine-readable instructions include an instruction 906 for receiving an indication of an increased workload portion of the workload to be added to the storage system, the workload being represented by a representation of buckets of I / O operations representing different characteristics.
[0117] The machine-readable instructions further include a factor value calculation instruction 908 that calculates factor values based on the amounts of I / O operations of different characteristics in the workload, the factor values indicating the distribution of the amount of I / O operations of the increased workload portion to the buckets of I / O operations representing different characteristics.
[0118] The machine-readable instructions further include an I / O operation distribution instruction 910 that distributes the amount of I / O operations of the increased workload portion to the buckets of I / O operations representing different characteristics according to the factor values.
[0119] The machine-readable instructions further include a modified representation generation instruction 912 that generates a modified representation of the buckets of I / O operations representing different characteristics in the modified workload including the workload and the increased workload portion.
[0120] The storage medium (e.g., Figure 8 in 800 orFigure 9 The 904) in may include any or some combination of the following: semiconductor memory devices, such as dynamic or static random access memory (DRAM or SRAM), erasable and programmable read only memory (EPROM), electrically erasable and programmable read only memory (EEPROM), and flash memory; disks, such as fixed, floppy, and removable disks; another magnetic medium including magnetic tape; optical media, such as compact disc (CD) or digital video disc (DVD); or another type of storage device. Note that the instructions discussed above may be provided on one computer-readable or machine-readable storage medium, or alternatively, may be provided on multiple computer-readable or machine-readable storage media distributed in a large system that may have multiple nodes. Such a computer-readable or machine-readable storage medium or media is considered to be part of an article (or article of manufacture). An article or article of manufacture may refer to any single manufactured component or multiple components. The storage medium or media may be located in the machine running the machine-readable instructions or at a remote site where the machine-readable instructions can be downloaded via a network for execution.
[0121] In the foregoing description, numerous details are set forth to provide an understanding of the subject matter disclosed herein. However, implementations may be practiced without these details. Other implementations may include modifications and variations to the above details. The appended claims are intended to cover these modifications and variations.
Claims
1. A non-transitory machine-readable storage medium, the non-transitory machine-readable storage medium including instructions that, when executed, cause a computing system to perform the following operations: Receive an indication of an increased workload portion to be added to an existing workload of a storage system, the existing workload including buckets of operations representative of different input / output (I / O) sizes; Based on the amounts of operations of the different I / O sizes in the existing workload, calculate a factor value indicative of how to distribute the amount of operations of the increased workload portion to the buckets of operations representative of the different I / O sizes; And Distribute the amount of operations of the increased workload portion to the buckets of operations representative of the different I / O sizes according to the factor value, wherein the amounts of operations of the different I / O sizes in the existing workload are within a window of a time interval, and the total amount of operations in the existing workload is within the window of the time interval, and wherein the window is a rolling window, and the calculation of the factor value is based on the amounts of operations of the different I / O sizes in the rolling window of the time interval.
2. The non-transitory machine-readable storage medium according to claim 1, wherein, A first bucket among the buckets includes an amount of operations of a first I / O size, and a second bucket among the buckets includes an amount of operations of a second I / O size.
3. The non-transitory machine-readable storage medium according to claim 1, wherein, Each respective bucket among the buckets includes an amount of operations of a respective I / O size among the different I / O sizes, and wherein distributing the amount of operations of the increased workload portion to the buckets representative of the operations results in an increase in the amount of operations in the buckets.
4. The non-transitory machine-readable storage medium according to claim 3, wherein, Distributing the amount of operations of the increased workload portion to the buckets representative of the operations results in a first increase in the amount of operations in a first bucket representative of the operations according to a first factor value among the factor values, and results in a second increase in the amount of operations in a second bucket representative of the operations according to a second factor value among the factor values.
5. The non-transitory machine-readable storage medium according to claim 1, wherein, The instructions, when executed, cause the computing system to perform the following operations: Calculate a first factor value among the factor values based on correlating the amount of operations of a first I / O size among the different I / O sizes in the existing workload with the total amount of operations in the existing workload.
6. The non-transitory machine-readable storage medium according to claim 5, wherein, The calculation of the first factor value further includes normalizing a correlation value generated by correlating the amount of operations of the first I / O size with the total amount of operations.
7. The non-transitory machine-readable storage medium according to claim 6, wherein, The normalization of the correlation value is based on a ratio of the amount of operations of the first I / O size to the total amount of operations.
8. The non-transitory machine-readable storage medium according to claim 1, wherein, Distributing the amount of operations of the increased workload portion to the buckets representative of the operations generates modified buckets representative of operations including the operations of the existing workload and the operations of the increased workload portion, and wherein the instructions, when executed, cause the computing system to perform the following operations: Use a prediction model based on the modified buckets representative of the operations to predict the performance of the storage system.
9. The non-transitory machine-readable storage medium according to claim 8, wherein, The predicted performance includes a predicted latency of the storage system or a predicted latency of resources of the storage system, or a predicted saturation of the storage system or a predicted saturation of resources of the storage system.
10. A computing system, comprising: A processor; And A non-transitory storage medium storing instructions executable on the processor to perform the following operations: Receiving an indication of an increased workload portion to be added to an existing workload of a storage system, the existing workload being represented by a representation of buckets of I / O operations representing different input / output (I / O) sizes; Based on the amounts of I / O operations of the different I / O sizes in the existing workload, calculating a factor value indicating how to distribute the amount of I / O operations of the increased workload portion among the buckets representing the I / O operations of the different I / O sizes; Distributing the amount of I / O operations of the increased workload portion among the buckets representing the I / O operations of the different I / O sizes according to the factor value; And Generating a modified representation of buckets of I / O operations of different I / O sizes in a modified workload, the modified workload including the existing workload and the increased workload portion, Wherein the amounts of operations of the different I / O sizes in the existing workload are within a window of a time interval, and the total amount of I / O operations in the existing workload is within the window of the time interval, and Wherein the window is a rolling window, and the calculation of the factor value is based on the amounts of operations of the different I / O sizes within the rolling window of the time interval.
11. The computing system according to claim 10, wherein, The indication of the increased workload portion is based on an indication that the existing workload is to be increased by a specified percentage.
12. The computing system according to claim 10, wherein, The representation of the buckets of I / O operations in the existing workload includes a histogram of buckets of I / O operations in a plurality of time intervals, wherein in each respective time interval of the plurality of time intervals of the histogram, a plurality of buckets represent the buckets of I / O operations in the respective time interval.
13. The computing system according to claim 10, wherein, Each respective bucket of the buckets of I / O operations in the existing workload includes a respective amount of I / O operations of a respective one of the different I / O sizes.
14. The computing system according to claim 10, wherein, The instructions are executable on the processor to calculate a respective factor value among the factor values based on correlating the amount of I / O operations of a respective one of the different I / O sizes in the existing workload with the total amount of I / O operations in the existing workload.
15. The computing system according to claim 14, wherein, The calculation of the respective factor value further includes normalizing a correlation value generated by correlating the amount of operations of the respective I / O size with the total amount of operations, wherein the normalization of the correlation value is based on the ratio of the amount of I / O operations of the respective I / O size to the total amount of I / O operations.
16. A method performed by a system including a hardware processor, the method comprising: Receiving an indication of an increased workload portion to be added to an existing workload of a storage system, the existing workload including buckets of operations representing different input / output (I / O) sizes; Based on the amounts of operations of the different I / O sizes in the existing workload, calculating a factor value indicating how to distribute the amount of operations of the increased workload portion among the buckets representing operations of different I / O sizes; Distribute the amount of operations of the increased workload portion into the buckets representing operations of the different I / O sizes according to the factor value to form modified buckets representing operations of the different I / O sizes; and Predict the performance of the storage system based on the modified buckets representing operations of the different I / O sizes, wherein the amount of operations of the different I / O sizes in the existing workload is within a window of a time interval, and the total amount of I / O operations in the existing workload is within the window of the time interval, and wherein the window is a rolling window, and the calculation of the factor value is based on the amount of operations of the different I / O sizes within the rolling window of the time interval.
17. The method according to claim 16, wherein The calculation of the factor value includes calculating each corresponding factor value in the factor value based on correlating the amount of operations of a corresponding I / O size among the different I / O sizes in the existing workload with the total amount of I / O operations in the existing workload.
Citation Information
Patent Citations
I / O request processing method and device
CN107688546A
Provisioning advisor
US20160349992A1