Input-output scheduling of virtualized computing instances

By using IO workload classification and multi-queue priority scheduling algorithms in a container-based hyperconverged infrastructure environment, the priority of IO requests is dynamically adjusted, and the performance degradation caused by large IO virtualization overhead and resource contention is solved, and efficient scheduling and resource optimization of IO requests are achieved.

CN120353534APending Publication Date: 2025-07-22DELL PROD LP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410088807.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-22
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In container-based hyperconverged infrastructure environments, IO virtualization is expensive and resource contention leads to performance degradation and unpredictable behavior, and it is difficult for the prior art to effectively optimize the scheduling of IO requests.

Method used

Through IO workload classification and priority determination logic, virtualized computing instances are grouped and multiple IO queues are generated. Using the multi-queue priority IO scheduling algorithm of the NVMe driver, the priority of IO requests is dynamically adjusted to optimize resource allocation and ensure the responsiveness of high-priority requests.

Benefits of technology

Improve IO performance, ensure fairness and responsiveness of IO operations, optimize resource allocation, and improve the overall performance of container-based HCI environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353534A_ABST
    Figure CN120353534A_ABST
Patent Text Reader

Abstract

An apparatus includes at least one processing device configured to identify an input-output (IO) workload classification for a virtualized computing instance issuing an IO request to a shared storage system, and determine a packet of the virtualized computing instance based on the IO workload classification. The at least one processing device is further configured to generate a plurality of IO queues associated with different IO priority levels for a given workload group, classifying IO requests received from virtualized computing instances of the given workload group into different queues of the plurality of IO queues based on information characterizing: (i) a time to serve IO requests given available resources of the shared storage system; and (ii) latency of IO requests received from virtualized compute instances of the given workload group, and processing IO requests based on different priority levels of the plurality of IO queues.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Information processing systems increasingly utilize reconfigurable virtual resources to meet changing user needs in an efficient, flexible, and cost-effective manner. For example, cloud computing and storage systems implemented using virtual resources such as virtual machines have been widely adopted. Other virtual resources now widely used in information processing systems include Linux containers. Such containers can be used to provide at least a portion of the virtualized infrastructure of an information processing system. Applications running on containers, virtual machines, or other virtual resources can include one or more processes that execute application functionality and issue input-output (IO) requests to be delivered to a storage system, including a storage system shared by multiple containers, virtual machines, or other virtual resources. The storage controller of the storage system services such IO requests. Summary of the Invention

[0002] Exemplary embodiments of the present disclosure provide techniques for IO scheduling of virtualized computing instances that issue IO requests to a shared storage device.

[0003] In one embodiment, a device includes at least one processing device including a processor coupled to a memory. The at least one processing device is configured to identify an IO workload classification for each of a plurality of virtualized computing instances that issue IO requests to a shared storage system. The at least one processing device is further configured to determine two or more virtualized computing instance workload groups based at least in part on the identified IO workload classifications of the plurality of virtualized computing instances, each of the two or more virtualized computing instance workload groups including a different subset of the plurality of virtualized computing instances. The at least one processing device is further configured to generate two or more IO queues associated with different IO priority levels for at least one given virtualized computing instance workload group of the two or more virtualized computing instance workload groups. The at least one processing device is further configured to classify IO requests received from the subset of the plurality of virtualized computing instances in the given virtualized computing instance workload group into the two or more IO queues, wherein a given IO request received from a given virtualized computing instance of the subset of the plurality of virtualized computing instances in the given virtualized computing instance workload group is placed in a given IO queue of the two or more IO queues based at least in part on information characterizing servicing of the IO request by the shared storage system. The at least one processing device is further configured to process the IO requests received from the subset of the plurality of virtualized computing instances in the given virtualized computing instance workload group based at least in part on the different priority levels associated with the two or more IO queues.

[0004] These and other illustrative embodiments include, but are not limited to, methods, devices, networks, systems, and processor-readable storage media. BRIEF DESCRIPTION OF THE DRAWINGS

[0005] Figure 1 is a block diagram of an information processing system configured for IO scheduling of virtualized computing instances that issue IO requests to a shared storage device in an illustrative embodiment.

[0006] Figure 2 is a flowchart of an exemplary process for IO scheduling of virtualized computing instances that issue IO requests to a shared storage device in an illustrative embodiment.

[0007] Figure 3 illustrates a system configured for multi-priority IO request scheduling in a container-based hyper-converged infrastructure environment in an illustrative embodiment.

[0008] Figure 4Illustrates a process for separating an I / O request stream into I / O request queues with different priorities in an illustrative embodiment.

[0009] Figure 5 Illustrates a table showing sample characteristics of an I / O data set for characterizing an I / O workload in an illustrative embodiment.

[0010] Figure 6 and Figure 7 Illustrates an example of a processing platform that can be used to implement at least a portion of an information processing system in an illustrative embodiment. Detailed Description

[0011] Illustrative embodiments will be described herein with reference to exemplary information processing systems and associated computers, servers, storage devices, and other processing devices. However, it should be understood that the embodiments are not limited to use with the specific illustrative system and device configurations shown. Thus, as used herein, the term "information processing system" is intended to be broadly construed to encompass, for example, processing systems including cloud computing and storage systems, as well as other types of processing systems including various combinations of physical and virtual processing resources. Thus, an information processing system can include, for example, at least one data center or other type of cloud-based system that includes one or more clouds hosting tenants accessing cloud resources.

[0012] Figure 1 Illustrates an information processing system 100 configured according to an illustrative embodiment. The information processing system 100 is assumed to be built on at least one processing platform and provides functionality for I / O scheduling of virtualized computing instances that issue I / O requests for a shared storage system service. The information processing system 100 includes a collection of client devices 102-1, 102-2, ..., 102-M (collectively client devices 102) coupled to a network 104. An information technology (IT) infrastructure environment 105 is also coupled to the network 104 and includes a collection of virtualized computing instances 106-1, 106-2, …, 106-N (collectively virtualized computing instances 106) that run corresponding collections of one or more applications 108-1, 108-2, …, 108-N (collectively applications 108). The client devices 102 are assumed to utilize the applications 108 running on the virtualized computing instances 106, which will generate various I / O requests that need to be serviced by the shared storage device 118 in the IT infrastructure environment. The IT infrastructure environment 105 includes an I / O scheduling system 110 that facilitates the servicing of such I / O requests for the shared storage device 118.

[0013] The I / O scheduling system 110 implements a set of I / O workload classification logic 112, I / O priority determination logic 114, and multi-priority I / O queues 116. A virtualized computing instance 106, which is part of running an application 108, is assumed to issue I / O requests to be served on a shared storage device 118 of an IT infrastructure environment 105. The I / O scheduling system 110 is configured to use the I / O workload classification logic 112 to organize the virtualized computing instance 106 or the application 108 running thereon into different I / O workload groups. Within each of the I / O workload groups, the I / O priority determination logic 114 assigns the I / O requests to different priority queues represented as the set of multi-priority I / O queues 116. The shared storage device 118 utilizes multi-priority I / O scheduling logic 120 to serve requests from the multi-priority I / O queues 116. The multi-priority I / O scheduling logic 120 can be implemented by a storage controller or a storage driver of the shared storage device 118. For example, the shared storage device can utilize a storage access protocol such as Non-Volatile Memory Express (NVMe), and the multi-priority I / O scheduling logic 120 can be implemented using the NVMe driver of the shared storage device 118.

[0014] The virtualized computing instance 106 is assumed to be implemented using one or more IT assets of the IT infrastructure environment, such as physical computing resources running virtualization infrastructure. The virtualized computing instance 106 is illustratively a software container or other type of virtual computing resource, such as a virtual machine (VM). In some embodiments, the IT infrastructure environment 105 includes a hyper-converged infrastructure (HCI) environment. In the case where the virtualized computing instance 106 is implemented as a software container, this can be a container-based HCI environment.

[0015] Although the I / O scheduling system 110 is shown as being external to the shared storage device 118 in Figure 1 , it should be understood that in some embodiments, the I / O scheduling system 110 can be implemented inside the shared storage device 118 (e.g., within one or more storage controllers or drivers of the shared storage device, which can also run or implement the multi-priority I / O scheduling logic 120).

[0016] The client device 102 can include, for example, physical computing devices such as IoT devices, mobile phones, laptops, tablets, desktop computers, or other types of devices used by enterprise members, in any combination. Such devices are examples of what are more generally referred to herein as "processing devices". Some of these processing devices are also generally referred to herein as "computers". The client device 102 also or alternatively includes virtual computing resources such as VMs, containers, etc.

[0017] In some embodiments, client device 102 includes various computers associated with a particular company, organization, or other enterprise. Thus, client device 102 can be considered an example of an asset of an enterprise system. Additionally, at least portions of information processing system 100 may also be referred to herein as collectively constituting one or more "enterprises." Many other operational scenarios involving a wide variety of different types and arrangements of processing nodes are possible, as will be understood by those skilled in the art.

[0018] Network 104 is assumed to include a portion of a global computer network such as the Internet, but other types of networks can be part of network 104, including wide area networks (WANs), local area networks (LANs), satellite networks, telephone or wireline networks, cellular networks, wireless networks (such as WiFi or WiMAX networks), or portions or combinations of these and other types of networks.

[0019] Shared storage device 118 can be implemented using one or more storage systems. The term "storage system" as used herein is intended to be construed broadly. A given storage system as used broadly herein can include, for example, content addressable storage devices, flash-based storage devices, network attached storage devices (NAS), storage area networks (SAN), direct attached storage devices (DAS), and distributed DAS, as well as combinations of these and other storage types (including software-defined storage devices). Other specific types of storage products that can be used to implement a storage system in an illustrative embodiment include all-flash and hybrid flash storage arrays, software-defined storage products, cloud storage products, object-based storage products, and scale-out NAS clusters. Combinations of multiple storage products among these and other storage products can also be used to implement a given storage system in an illustrative embodiment.

[0020] In some embodiments, shared storage device 118 is part of a software-defined IT infrastructure, such as hyper-converged infrastructure (HCI) including virtualized computing instances 106, software-defined storage devices providing shared storage device 118, and virtualized networking (e.g., software-defined networking) linking virtualized computing instances 106 to shared storage device 118.

[0021] Although not explicitly shown in Figure 1 one or more input-output devices such as a keyboard, a display, or other types of input-output devices can be used to support one or more user interfaces to IO scheduling system 110 and to support communication between IO scheduling system 110 and other related systems and devices not explicitly shown.

[0022] In some embodiments, the client device 102 is assumed to be associated with a user of an enterprise, organization, or other entity that also operates the IT infrastructure environment 105. In other embodiments, the client device 102 may be associated with a user of one or more enterprises, organizations, or other entities that are different from the enterprise, organization, or other entity that operates the IT infrastructure environment 105.

[0023] Figure 1 In the embodiments of Figure 1 , the IO scheduling system 110 and the shared storage device 118 are assumed to be implemented using at least one processing device. Each such processing device typically includes at least one processor and associated memory, and implements one or more functional modules or logic to control certain features of the IO scheduling system 110 and the shared storage device 118. In Figure 1 the embodiments of Figure 1 , the IO scheduling system 110 implements IO workload classification logic 112, IO priority determination logic 114, and a multi-priority IO queue 116, while the shared storage device 118 implements multi-priority IO scheduling logic 120. As described above, in some embodiments, the IO scheduling system 110 may be implemented inside the shared storage device 118 such that the same set of one or more processing devices (e.g., jointly providing or implementing the storage controller or storage driver of the shared storage device 118) may implement the IO workload classification logic 112, the IO priority determination logic 114, as well as the multi-priority IO queue 116 and the multi-priority IO scheduling logic 120. In other embodiments, the multi-priority IO scheduling logic 120 may be implemented outside the shared storage device 118, such as within the IO scheduling system 110. Various other combinations are possible. Thus, it should be understood that Figure 1 the specific arrangement of the client device 102, the IT infrastructure environment 105, the virtualized computing instance 106, the IO scheduling system 110, and the shared storage device 118 shown in the embodiments of Figure 1 is presented only by way of example, and alternative arrangements may be used in other embodiments.

[0024] At least part of the IO workload classification logic 112, the IO priority determination logic 114, and the multi-priority IO queue 116 and the multi-priority IO scheduling logic 120 may be implemented at least in part in the form of software stored in memory and executed by a processor.

[0025] As will be described in more detail below, the IO scheduling system 110 and other parts of the information processing system 100 may be part of a cloud infrastructure.

[0026] The IO scheduling system 110 and Figure 1The other components of the information processing system 100 in the embodiments are assumed to be implemented using at least one processing platform including one or more processing devices, each processing device having a processor coupled to a memory. Such processing devices may illustratively include a particular arrangement of computing, storage, and network resources.

[0027] Although the client device 102, the IT infrastructure environment 105, the virtualized computing instance 106, the IO scheduling system 110, and the shared storage device 118 or their components (e.g., the application 108, the IO workload classification logic 112, the IO priority determination logic 114, and the multi-priority IO queue 116 and the multi-priority IO scheduling logic 120) may be implemented on correspondingly different processing platforms, many other arrangements are possible. For example, in some embodiments, at least part of the IO scheduling system 110 and the shared storage device 118 are implemented on the same processing platform. Additionally, a given client device (e.g., 102-1) may be at least partially implemented within at least one processing platform that implements at least part of the IT infrastructure environment 105.

[0028] As used herein, the term "processing platform" is intended to be interpreted broadly so as to include, for example but not limited to, multiple sets of processing devices and associated storage systems configured to communicate via one or more networks. For example, a distributed implementation of the information processing system 100 is possible, where certain components of the system reside in one data center at a first geographic location, while other components of the system reside in one or more other data centers at one or more other geographic locations that may be remote from the first geographic location. Thus, in some embodiments of the information processing system 100, the client device 102 and the IT infrastructure environment 105 or parts or components thereof may reside in different data centers. Many other distributed implementations are possible.

[0029] The following will be combined with Figure 6 and Figure 7 to more specifically describe additional examples of the processing platforms for implementing the IO scheduling system 110 and the other components of the information processing system 100 in the illustrative embodiments.

[0030] It should be understood that these and other features of the illustrative embodiments are presented only by way of example and should not be construed as restrictive in any way.

[0031] It should be understood that Figure 1The specific set of elements for the I / O scheduling of virtualized computing instances for issuing I / O requests to a shared storage device shown is presented only by way of illustrative example, and in other embodiments, additional or alternative elements may be used. Thus, another embodiment may include additional or alternative systems, devices, and other network entities, as well as different arrangements of modules and other components.

[0032] It should be understood that these and other features of the illustrative embodiments are presented only by way of example and should not be construed as limiting in any way.

[0033] Reference will now be made Figure 2 to the flowchart of to more particularly describe an exemplary process for the I / O scheduling of virtualized computing instances for issuing I / O requests to a shared storage device. It should be understood that this particular process is merely an example, and additional or alternative processes for the I / O scheduling of virtualized computing instances for issuing I / O requests to a shared storage device may be used in other embodiments.

[0034] In this embodiment, the process includes steps 200 to 208. These steps are assumed to be performed by an I / O scheduling system 110 and / or a shared storage device 118 using I / O workload classification logic 112, I / O priority determination logic 114, and a multi-priority I / O queue 116 and multi-priority I / O scheduling logic 120. The process begins at step 200: identifying an I / O workload classification for each of the virtualized computing instances 106 that issue I / O requests to the shared storage device 118. The IT infrastructure environment may include an HCI environment. The virtualized computing instances 106 may include software containers, and the HCI environment may include a container-based HCI environment. The virtualized computing instances 106 and the shared storage device 118 may operate on a common physical infrastructure in a container-based HCI environment.

[0035] In step 202, two or more virtualized computing instance workload groups are determined based at least in part on the identified I / O workload classifications of the virtualized computing instances 106 in the IT infrastructure environment 105. Each of the virtualized computing instance workload groups includes a different subset of the virtualized computing instances 106.

[0036] In step 204, two or more IO queues associated with different IO priority levels are generated for at least one given virtualized computing instance workload group among two or more virtualized computing instance workload groups. In step 206, IO requests received from a subset of virtualized computing instances 106 in the given virtualized computing instance workload group are classified into two or more IO queues. A given IO request received from a given virtualized computing instance in the subset of virtualized computing instances 106 in the given virtualized computing instance workload group is placed in a given IO queue among the two or more IO queues, at least in part based on information characterizing the servicing of the IO request by the shared storage device 118. The information characterizing the servicing of the IO request by the shared storage device 118 may include (i) the time to service the given IO request given the available resources of the given shared storage device 118, and (ii) the waiting time of IO requests received from the subset of virtualized computing instances 106 in the given virtualized computing instance workload group.

[0037] The two or more IO queues generated for a given virtualized computing instance workload group may include a first IO queue associated with a first priority level and at least a second IO queue associated with a second priority level different from the first priority level.

[0038] Step 204 may include placing IO requests having a responsibility ratio greater than a threshold in a first IO queue associated with the first priority level among the two or more IO queues, and placing IO requests having a responsibility ratio less than or equal to the threshold in a second IO queue associated with the second priority level among the two or more IO queues. The responsibility ratio of a given IO request received from a given virtualized computing instance is determined at least in part based on the total waiting time associated with the IO request received from the given virtualized computing instance and the amount of time taken to service the given IO request. The threshold may include a value range determined at least in part based on analyzing the IO request stream and the available resources of the physical infrastructure on which the virtualized computing instances 106 and the shared storage device 118 are running.

[0039] Information characterizing the latency of I / O requests received from a given virtualized computing instance can include one or more of the following: the average responsible time for each of the I / O requests received from the given virtualized computing instance during a specified time period; the average wait time for each of the I / O requests received from the given virtualized computing instance during a specified time period; the number of I / O requests received from the given virtualized computing instance per second during a specified time period; the rate of random write I / O requests received from the given virtualized computing instance during a specified time period; and the amount of data written by the given virtualized computing instance to the shared storage device 118 during a specified time period.

[0040] In step 208, I / O requests received from a subset of virtualized computing instances 106 of a given virtualized computing instance workload group are processed at least in part based on different priority levels associated with two or more I / O queues. Step 208 can utilize a multi-priority I / O scheduling algorithm. The multi-priority I / O scheduling algorithm can be implemented using the NVMe driver of the shared storage device 118.

[0041] Exemplary embodiments provide technical solutions for optimizing the scheduling of I / O requests in a virtualized computing environment, including but not limited to a container-based hyper-converged infrastructure (HCI) environment. HCI is an infrastructure deployment model that combines storage, computing, and network resources into a single cluster. A container-based HCI environment provides flexibility and agility in deploying and managing workloads. However, the I / O characteristics of a container-based HCI environment have not been fully understood and pose various technical challenges.

[0042] A container-based HCI environment provides several unique characteristics in terms of I / O virtualization. For example, faster and more efficient I / O operations can be achieved using lightweight containerization techniques. Containers share the underlying host operating system kernel, which reduces the need for redundant I / O operations and can improve overall performance. Additionally, containers can be quickly deployed and scaled, which can help optimize I / O performance by quickly allocating and deallocating resources as needed. Container-based I / O has several unique characteristics, but it also faces some technical challenges related to efficiency.

[0043] One of the key technical challenges in container-based IO is the overhead of IO virtualization. In a VM-based virtualized computing environment, IO virtualization is implemented using a hypervisor, which can incur significant overhead. In contrast, container-based IO uses lightweight virtualization technologies such as Linux Containers (LXC) or Docker, which have lower overhead than VMs. However, especially when multiple containers access shared storage resources, the overhead of IO virtualization is still significant. Another challenge in container-based IO is potential resource contention. Containers running on the same host can compete for shared resources such as CPU, memory, and network bandwidth. Especially when multiple containers access shared storage resources simultaneously, this competition can lead to performance degradation and unpredictable behavior.

[0044] The technical solution described in this paper provides a scheduling method for classifying and prioritizing different types of IO requests (e.g., originating from containers in a container-based HCI environment), thereby improving IO performance. In some embodiments, a ratio is calculated based on the IO data to ensure fairness and responsiveness of IO operations, while also optimizing or improving resource allocation and ensuring that sufficient resources are available for high-priority IO requests. Thus, the technical solution provides a novel method for prioritizing concurrent IO (e.g., in a container-based HCI environment). In some embodiments, the technical solution utilizes the multi-queue priority IO scheduling functionality of NVMe or other storage drivers to achieve optimal or improved performance.

[0045] IO performance metrics are collected from an IT infrastructure environment such as one or more container-based HCI environments. Such IO performance metrics are used as input to an algorithm that calculates and differentiates different priorities (e.g., high, medium, and low-priority IO requests) of IO requests based on the relevant IO performance metrics of the IO requests. To further optimize or improve IO performance, the metrics calculated using the algorithm are integrated with multi-queue priority IO scheduling features (e.g., features provided in an NVMe driver). Such features allow for the creation of multiple IO queues with different priorities, ensuring that higher-priority IO requests are processed first and then lower-priority IO requests. By leveraging such features in combination with the IO scheduling method described in this paper, the technical solution is able to consistently prioritize higher-priority IO requests, thereby achieving optimal or improved performance in container-based HCI and other IT infrastructure environments.

[0046] Accordingly, the technical solution provides a novel scheduling algorithm that calculates metrics for prioritizing I / O performance in container-based HCI and other IT infrastructure environments. In some embodiments, such metrics are integrated with the multi-queue priority I / O scheduling features of NVMe to achieve optimal or improved performance, which ensures that higher-priority I / O requests are served first and then lower-priority I / O requests. Advantageously, the technical solutions described herein have the potential to enhance the efficiency and reliability of container-based HCI or other IT infrastructure environments, enabling organizations to fully leverage the advantages of such environments.

[0047] Figure 3 Shown is a system 300 configured to implement a technical solution for scheduling I / O requests received from a collection of containers 301-1, 301-2, 301-3, ..., 301-C (collectively, containers 301) in a container-based HCI environment to improve overall performance and meet specific service requirements. It should be noted that Figure 3 Containers 301 can be examples of virtualized computing instances 106 running one or more applications 108 that issue I / O requests to a shared storage device 118. Each of the containers 301 typically runs only one application, also referred to as the workload of the container. At the start of each of the containers 301, a container workload I / O classification engine 303 is utilized to determine the I / O workload type and priority of that container. The containers 301 are then grouped into a collection of container workload groups 305-1, 305-2, ..., 305-G (collectively, container workload groups 305) based on such classification.

[0048] Then, according to an algorithm referred to herein as Highest Response Ratio Next (HRRN), the workloads within each of the container workload groups 305 are sorted by priority. This results in multiple I / O queues for each of the container workload groups 305. For example, container workload group 305-1 has a high-priority queue (HPQ) 350-1-1 and a low-priority queue (LPQ) 350-1-2. HPQ 350-1-1 and LPQ 350-1-2 are collectively referred to as the multi-priority queue 350-1 associated with container workload group 305-1. Similarly, container workload group 305-2 has HPQ 350-2-1 and LPQ 350-2-2, which are collectively referred to as the multi-priority queue 350-2 associated with container workload group 305-2. Container workload group 305-G has HPQ 350-G-1 and LPQ 350-G-2, which are collectively referred to as the multi-priority queue 350-G associated with container workload group 305-G. Although in Figure 3In the example, there are two IO queues (e.g., HPQ and LPQ) associated with each of the container workload groups 305, but this is not a requirement. One or more of the container workload groups 305 can be associated with three or more IO queues (e.g., HPQ, Medium Priority Queue (MPQ), and LPQ). Additionally, different container workload groups within the container workload group 305 can be associated with different numbers of IO queues. For example, some container workload groups can be associated with only two IO queues (e.g., HPQ and LPQ), while other container workload groups can be associated with three IO queues (e.g., HPQ, MPQ, and LPQ) or more than three IO queues. Various other combinations are possible. The storage controller 307 associated with the shared storage system utilized by the container 301 implements multi-queue priority IO scheduling logic 309 to service the IOs in the multi-priority queues 350-1, 350-2, ..., 350-G (collectively referred to as the multi-priority queues 350). In some embodiments, the storage controller 307 includes an NVMe driver for the shared storage system.

[0049] IO prioritization within each of the container workload groups 305 can be performed using the HRRN algorithm, which calculates the responsible ratio (RR) for each IO to dynamically adjust the priority of that IO. The output of the HRRN algorithm is to separate the IOs into the multi-priority queues 350. The multi-queue priority IO scheduling logic 309 of the storage controller 307 services the IOs from the multi-priority queues 350, ensuring fairness and responsiveness of the IO requests from the container 301, optimizing or improving resource allocation, and guaranteeing sufficient resources for high-priority IO operations.

[0050] The HRRN algorithm (which can also be referred to as container-based dynamic HRRN) dynamically optimizes the IO scheduling from the container 301. The HRRN algorithm considers the application or workload type, IO pattern, and responsible time information to optimize the IO scheduling to help quickly improve the performance of container-based HCI environments and other IT infrastructure environments.

[0051] In each of the containers 301, the IO priority can be determined by a threshold T according to Figure 4Filtered by the schema 400. IO requests with a high RR (e.g., RR greater than T) will be placed in a higher priority queue (e.g., one of HPQ 350-1-1, 350-2-1,..., 350-G-1). Otherwise, the IO request will remain in a lower priority queue (e.g., one of LPQ 350-1-2, 350-2-2,..., 350-G-2). The threshold T is not a fixed value and can be dynamically modified within a certain range (e.g., within the range [T - Δt, T + Δt]), where Δt depends on the IO request characteristics and the system resource allocation within a short time period. The threshold T is a key value used to determine whether a specific IO request should be moved to a higher or lower priority queue. Considering the fairness of each IO request stream, the threshold T may vary within the range [T - Δt, T + Δt]. The method for determining T and Δt will be discussed in further detail below.

[0052] It should be noted that when using more than two priority queues (e.g., such as HPQ, MPQ, and LPQ), a similar separation can be applied by using two thresholds (e.g., threshold T1 for separating HPQ and MPQ and threshold T2 for separating MPQ and LPQ). This can be extended as needed according to the number of queues utilized.

[0053] Let the IO stream (also known as the IO request stream) and the system resources be represented as a matrix, where C j represents the IO stream, and X j represents the system resources. The optimized result max z is the expected value determined according to the following terms:

[0054]

[0055]

[0056] This can be transformed into a matrix:

[0057]

[0058] Then, according to the requirement of the minimum value of X (X = z = Δt) as needed, the required value is a constrained value, such as:

[0059] x1 + x2 + x3 +... + x n ≤ m

[0060] Therefore, the problem in the original format can be described as:

[0061]

[0062] The optimized result maxz is the desired value, where Δt = maxz. In the above equation, k and h are constant values that provide a basic feasible solution for the matrix of the linear programming problem, while assuming x1 = x2 = x3 = … = x n = 0.

[0063] Now, the responsibility ratio RR will be described. In some embodiments, the system will preprocess different types of IOs to generate a data set. By executing various applications and collecting IO traces, different IO workloads may be generated. Different IO workloads may include different operations, such as reading, writing, and waiting operations on 10,000 files through multithreading. Container IO metrics can be collected to identify the characteristics of each IO and reform them into a data set of the IO workload. In this way, a thread can be defined to handle IO requests, and the thread will include various IO operations and related system resources. For each IO data set I i , the corresponding column T i that captures the IO characteristics may include Figure 5 those shown in table 500, where column T i includes characteristics such as the following: the average responsibility time of container IO (AVG-RT), that is, the average responsibility time of each IO in the container; the average wait time of container IO (AVG-WT), that is, the average wait time of each IO in the container; the container IO request (CR), that is, the number of IO requests per second for each running container; the container random write rate (CRWR), that is, the rate of random write IO requests for each running container; and the container write bytes (CWB), that is, the number of bytes written by each running container. Thus, I i = {AVG-RT, AVG-WT, CR, CRWR, CWB}. It should be noted that Figure 5 the specific values of the different variables shown are presented only by way of example. For example, the value of AVG-RT depends on the responsibility time for the system to process a single IO request. For a hard disk drive (HDD), the average single-IO responsibility time may be approximately 5 milliseconds (ms). The value of AVG-WT is the average wait time when the system processes an IO request, and for an HDD, depending on the load of the system, it may be approximately 70 ms or longer. For other types of drives (e.g., such as solid state drives (SSDs) or other flash-based memories), these values may be different.

[0064] Therefore, the total wait time for each I i is determined according to the following items:

[0065]

[0066] Tr represents the requested service time, which indicates how long it takes to service a single IO request in the IO dataset I i The RR can be calculated based on the following items:

[0067]

[0068] RR is used to determine the IO priority, where a higher value of RR indicates a higher relative IO priority.

[0069] It should be understood that the specific advantages described above and elsewhere in this document are associated with specific illustrative embodiments and need not be present in other embodiments. Moreover, the specific types of information processing system features and functionality shown in the figures and described above are merely exemplary, and numerous other arrangements may be used in other embodiments.

[0070] Now reference will be made to Figure 6 and Figure 7 to describe in more detail an illustrative embodiment of a processing platform for implementing the functionality of an IO scheduler for a virtualized computing instance that issues IOs to a shared storage device. Although described in the context of system 100, in other embodiments, these platforms may also be used to implement at least portions of other information processing systems.

[0071] Figure 6 An exemplary processing platform including a cloud infrastructure 600 is shown. The cloud infrastructure 600 includes a combination of physical and virtual processing resources that can be used to implement at least a portion of the information processing system 100 in Figure 1 The cloud infrastructure 600 includes a plurality of virtual machines (VMs) and / or container sets 602-1, 602-2,... 602-L implemented using a virtualization infrastructure 604. The virtualization infrastructure 604 runs on a physical infrastructure 605 and illustratively includes one or more hypervisors and / or operating system-level virtualization infrastructures. The operating system-level virtualization infrastructure illustratively includes kernel control groups of a Linux operating system or other types of operating systems.

[0072] The cloud infrastructure 600 also includes a set of applications 610-1, 610-2,... 610-L that run on respective ones of the VM / container sets 602-1, 602-2,... 602-L under the control of the virtualization infrastructure 604. The VM / container sets 602 may include respective VMs, respective sets of one or more containers, or respective sets of one or more containers running within a VM.

[0073] In Figure 6In some embodiments of the implementation, the VM / container set 602 includes corresponding VMs implemented using a virtualization infrastructure 604 that includes at least one hypervisor. A hypervisor platform can be used to implement a hypervisor within the virtualization infrastructure 604, where the hypervisor platform has an associated virtual infrastructure management system. The underlying physical machine can include one or more distributed processing platforms, and the distributed processing platforms include one or more storage systems.

[0074] In Figure 6 In other embodiments of the implementation, the VM / container set 602 includes individual containers implemented using a virtualization infrastructure 604 that provides operating system-level virtualization functionality (such as support for Docker containers running on bare metal hosts or Docker containers running on VMs). Containers are illustratively implemented using the corresponding kernel control groups of the operating system.

[0075] As is apparent from the above, one or more of the processing modules or other components of the system 100 can each run on a computer, server, storage device, or other processing platform element. A given such element can be regarded as an example of what is more generally referred to herein as a "processing device". Figure 6 The illustrated cloud infrastructure 600 can represent at least a portion of a processing platform. Another example of such a processing platform is Figure 7 the processing platform 700 shown in

[0076] In this implementation, the processing platform 700 includes a portion of the system 100 and includes a plurality of processing devices represented as 702-1, 702-2, 702-3,..., 702-K that communicate with each other through a network 704.

[0077] The network 704 can include any type of network, such as including a global computer network (such as the Internet), WAN, LAN, satellite network, telephone or wired network, cellular network, wireless network (such as a WiFi or WiMAX network), or portions or combinations of these and other types of networks.

[0078] The processing device 702-1 in the processing platform 700 includes a processor 710 coupled to a memory 712.

[0079] The processor 710 can include a microprocessor, microcontroller, application specific integrated circuit (ASIC), field programmable gate array (FPGA), central processing unit (CPU), graphics processing unit (GPU), tensor processing unit (TPU), video processing unit (VPU), or other types of processing circuits, as well as portions or combinations of such circuit elements.

[0080] The memory 712 can include random access memory (RAM), read only memory (ROM), flash memory, or other types of memory in any combination. The memory 712 and other memories disclosed herein should be regarded as illustrative examples of what is more generally referred to as a "processor-readable storage medium" that stores executable program code for one or more software programs.

[0081] Articles of manufacture that include such processor-readable storage media are considered illustrative embodiments. A given such article of manufacture can include, for example, a storage array, a storage disk, or an integrated circuit that includes RAM, ROM, flash memory, or other electronic memory, or any of a variety of other types of computer program products. As used herein, the term "article of manufacture" should be understood to exclude transient propagated signals. Many other types of computer program products that include a processor-readable storage medium can be used.

[0082] The processing device 702-1 also includes a network interface circuit 714 for docking the processing device with the network 704 and other system components and can include a conventional transceiver.

[0083] The other processing devices 702 of the processing platform 700 are assumed to be configured in a manner similar to that shown for the processing device 702-1 in the figure.

[0084] Moreover, the specific processing platform 700 shown in the figure is presented only by way of example, and the system 100 can include additional or alternative processing platforms, and can include numerous different processing platforms in any combination, where each such platform includes one or more computers, servers, storage devices, or other processing devices.

[0085] For example, other processing platforms for implementing illustrative embodiments can include converged infrastructure.

[0086] Accordingly, it should be understood that in other embodiments, different arrangements of additional or alternative elements can be used. At least a subset of these elements can be implemented together on a common processing platform, or each such element can be implemented on a separate processing platform.

[0087] As previously indicated, the components of the information processing system disclosed herein can be implemented at least in part in the form of one or more software programs stored in memory and executed by a processor of a processing device. For example, at least a portion of the functionality for IO scheduling of virtualized computing instances for issuing IO requests to a shared storage device as disclosed herein is illustratively implemented in the form of software running on one or more processing devices.

[0088] It should be emphasized again that the above-described implementation examples are presented for illustrative purposes only. Many variations and other alternative implementation examples can be used. For example, the disclosed technology can be applied to various other types of information processing systems, virtualization infrastructures, etc. In addition, the specific configurations of the system and device elements illustratively shown in the drawings and the associated processing operations can be changed in other implementation examples. Furthermore, the various assumptions made above during the description of the illustrative implementation examples should also be considered exemplary and not as requirements or limitations of the present disclosure. Numerous other alternative implementation examples within the scope of the appended claims will be apparent to those skilled in the art.

Claims

1. A device, comprising: at least one processing device, the at least one processing device including a processor coupled to a memory; the at least one processing device is configured to: identify an input-output workload classification for each of a plurality of virtualized computing instances that issue input-output requests to a shared storage system; determine two or more virtualized computing instance workload groups at least in part based on the identified input-output workload classifications of the plurality of virtualized computing instances, each of the two or more virtualized computing instance workload groups including a different subset of the plurality of virtualized computing instances; generate two or more input-output queues associated with different input-output priority levels for at least one given virtualized computing instance workload group of the two or more virtualized computing instance workload groups; classify input-output requests received from the subset of the plurality of virtualized computing instances in the given virtualized computing instance workload group into the two or more input-output queues, wherein a given input-output request received from a given virtualized computing instance in the subset of the plurality of virtualized computing instances in the given virtualized computing instance workload group is placed in a given input-output queue of the two or more input-output queues at least in part based on information characterizing the servicing of input-output requests by the shared storage system; and process the input-output requests received from the subset of the plurality of virtualized computing instances in the given virtualized computing instance workload group at least in part based on the different priority levels associated with the two or more input-output queues.

2. The device according to claim 1, wherein the plurality of virtualized computing instances and the shared storage device are part of a hyperconverged infrastructure environment.

3. The device according to claim 2, wherein the plurality of virtualized computing instances include software containers, and wherein the hyperconverged infrastructure environment includes a container-based hyperconverged infrastructure environment.

4. The device according to claim 1, wherein processing the input-output requests received from the subset of the plurality of virtualized computing instances in the given virtualized computing instance workload group includes utilizing a multi-priority input-output scheduling algorithm.

5. The device according to claim 4, wherein the multi-priority input-output scheduling algorithm is implemented using a fast non-volatile memory driver of the shared storage system.

6. The device according to claim 1, wherein the plurality of virtualized computing instances and the shared storage system operate on a common physical infrastructure in an information technology infrastructure environment.

7. The device according to claim 1, wherein classifying the input-output requests received from the subset of the plurality of virtualized computing instances in the given virtualized computing instance workload group into the two or more input-output queues includes placing the input-output requests having a responsibility ratio greater than a threshold in a first input-output queue associated with a first priority level among the two or more input-output queues, and placing the input-output requests having a responsibility ratio less than or equal to the threshold in a second input-output queue associated with a second priority level among the two or more input-output queues.

8. The device according to claim 7, wherein the responsibility ratio of a given input-output request among the input-output requests received from the given virtualized computing instance is determined based at least in part on the total waiting time associated with the input-output request received from the given virtualized computing instance and the amount of time taken to process the given input-output request.

9. The device according to claim 7, wherein the threshold includes a value range determined based at least in part on analyzing the flow of the input-output requests and the available resources of the physical infrastructure on which the plurality of virtualized computing instances and the shared storage system are running.

10. The device according to claim 1, wherein the information characterizing the service of the input-output requests by the shared storage system includes (i) the time to service the given input-output request given the available resources of the shared storage system, and (ii) the waiting time of the input-output requests received from the subset of the plurality of virtualized computing instances in the given virtualized computing instance workload group.

11. The device according to claim 10, wherein the information characterizing the waiting time of the input-output requests received from the given virtualized computing instance includes the average responsibility time of each of the input-output requests received from the given virtualized computing instance within a specified time period.

12. The device according to claim 10, wherein the information characterizing the waiting time of the input-output requests received from the given virtualized computing instance includes the average waiting time of each of the input-output requests received from the given virtualized computing instance within a specified time period.

13. The device according to claim 10, wherein the information characterizing the waiting time of the input-output requests received from the given virtualized computing instance includes at least one of the following: the number of input-output requests received per second from the given virtualized computing instance within a specified time; and the amount of data written by the given virtualized computing instance to the shared storage system within a specified time period.

14. The apparatus according to claim 10, wherein the information characterizing the latency of the input-output requests received from the given virtualized computing instance includes the rate of random write input-output requests received from the given virtualized computing instance within a specified period of time.

15. A computer program product, the computer program product comprising a non-transitory processor-readable storage medium having program code stored therein with one or more software programs, wherein the program code, when executed by at least one processing device, causes the at least one processing device to: Identify an input-output workload classification for each of a plurality of virtualized computing instances that issue input-output requests to a shared storage system; Determine two or more virtualized computing instance workload groups based at least in part on the identified input-output workload classifications of the plurality of virtualized computing instances, each of the two or more virtualized computing instance workload groups including a different subset of the plurality of virtualized computing instances; Generate two or more input-output queues associated with different input-output priority levels for at least one given virtualized computing instance workload group of the two or more virtualized computing instance workload groups; Classify input-output requests received from the subset of the plurality of virtualized computing instances in the given virtualized computing instance workload group into the two or more input-output queues, wherein a given input-output request among the input-output requests received from a given virtualized computing instance in the subset of the plurality of virtualized computing instances in the given virtualized computing instance workload group is placed in a given input-output queue of the two or more input-output queues based at least in part on information characterizing the servicing of the input-output request by the shared storage system; and Process the input-output requests received from the subset of the plurality of virtualized computing instances in the given virtualized computing instance workload group based at least in part on the different priority levels associated with the two or more input-output queues.

16. The computer program product according to claim 15, wherein the plurality of virtualized computing instances include software containers, and wherein the plurality of virtualized computing instances and the shared storage device are part of a container-based hyper-converged infrastructure environment.

17. The computer program product according to claim 15, wherein the information characterizing the servicing of the input-output request by the shared storage system includes (i) the time to service the given input-output request given the available resources of the shared storage system, and (ii) the latency of the input-output requests received from the subset of the plurality of virtualized computing instances in the given virtualized computing instance workload group.

18. A method, comprising: Identify an input-output workload classification for each of a plurality of virtualized computing instances that issue input-output requests to a shared storage system; Determine two or more virtualized computing instance workload groups at least in part based on the identified input-output workload classifications of the plurality of virtualized computing instances, each of the two or more virtualized computing instance workload groups including a different subset of the plurality of virtualized computing instances; Generate two or more input-output queues associated with different input-output priority levels for at least one given virtualized computing instance workload group of the two or more virtualized computing instance workload groups; Classify input-output requests received from the subset of the plurality of virtualized computing instances in the given virtualized computing instance workload group into the two or more input-output queues, wherein a given input-output request received from a given virtualized computing instance in the subset of the plurality of virtualized computing instances in the given virtualized computing instance workload group is placed in a given input-output queue of the two or more input-output queues at least in part based on information characterizing the servicing of the input-output request by the shared storage system; And Process the input-output requests received from the subset of the plurality of virtualized computing instances in the given virtualized computing instance workload group at least in part based on the different priority levels associated with the two or more input-output queues; Wherein the method is performed by at least one processing device including a processor coupled to a memory.

19. The method of claim 18, wherein the plurality of virtualized computing instances includes software containers, and wherein the plurality of virtualized computing instances and the shared storage device are part of a container-based hyperconverged infrastructure environment.

20. The method of claim 18, wherein the information characterizing the servicing of the input-output request by the shared storage system includes (i) the time to service the given input-output request given the available resources of the shared storage system, and (ii) the latency of the input-output requests received from the subset of the plurality of virtualized computing instances in the given virtualized computing instance workload group.