Storage-level load balancing

The method addresses I/O queue duplication in NVMe storage systems by rebalancing I/O traffic across CPU cores, enhancing performance and load distribution, thereby reducing latency and improving throughput.

JP7702198B2Active Publication Date: 2025-07-03INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023513203
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-08-26
Filing Date
2021-08-12
Publication Date
2025-07-03
Estimated Expiration
2041-08-12

AI Technical Summary

Technical Problem

Existing NVMe storage systems face performance degradation due to I/O queue duplication, leading to unbalanced CPU core loads and increased latency, especially during peak workloads.

Method used

A computer-implemented method for storage-level load balancing that monitors CPU core utilization and rebalances I/O queues by detecting overloaded states and recommending the transfer of I/O traffic to underutilized CPU cores, using NVMe-oF architecture to manage I/O queues across multiple hosts.

Benefits of technology

Reduces queue duplication bottlenecks, improves performance, increases IOPS, and achieves better load distribution among CPU cores, minimizing latency and resource underutilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007702198000001
    Figure 0007702198000001
  • Figure 0007702198000002
    Figure 0007702198000002
  • Figure 0007702198000003
    Figure 0007702198000003
Patent Text Reader

Abstract

In a storage-level load balancing approach, a load level of a storage system is monitored, the load level being a utilization rate of multiple CPU cores in the storage system. An overload condition is detected based on the utilization rate of one or more CPU cores exceeding a threshold, the overload condition being caused by overlapping of one or more I / O queues from multiple host computers accessing a single CPU core. In response to detecting the overload condition, a new I / O queue on a second CPU core is selected, the second CPU core having a utilization rate less than a second threshold. A recommendation is sent to the first host computer, the recommendation being to move I / O traffic from the first CPU core to the new I / O queue on the second CPU core to rebalance the load level of the storage system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to the field of computer storage, and more particularly to storage-level load balancing.

Background Art

[0002] Non-Volatile Memory Express (NVMe (trademark)) is an optimal and high-performance extensible host controller interface designed to meet the needs of computer storage systems that utilize solid-state storage based on the Peripheral Component Interconnect Express (registered trademark) (PCIe (registered trademark)) interface. Designed from the ground up for non-volatile memory technology, NVMe is designed to efficiently access storage devices with non-volatile memory, from current NAND flash technology to future high-performance persistent memory technology.

[0003] The NVMe protocol, like high-performance processor architectures, utilizes parallel and low-latency data paths to the underlying media. This enables significantly higher performance and lower latency compared to conventional storage interfaces such as serial-attached SCSI (SAS) and serial ATA (SATA) protocols. NVMe can support up to 65,535 input / output (I / O) queues, with each queue having 65,535 entries. Legacy SAS and SATA interfaces only support a single queue with 254 entries per SAS queue and 32 entries per SATA queue. NVMe host software can create queues up to the maximum allowed by the NVMe controller, depending on the system configuration and the expected workload. NVMe supports scatter / gather I / O, which minimizes CPU overhead during data transfer, and also provides a function to change priorities based on workload requirements.

[0004] NVMe over Fabrics (NVMe-oF) is a network protocol used for communication between a host and a storage system over a network (also known as a fabric). NVMe-oF defines a common architecture to support various storage networking fabrics for the NVMe block storage protocol over a storage networking fabric. This includes realizing a front-end interface to the storage system, scaling out to a large number of NVMe devices, and extending the distance to which NVMe devices and NVMe subsystems can be accessed.

Summary of the Invention

[0005] One aspect of the present invention includes a computer-implemented method for load distribution at the storage level. In a first embodiment, the load level of a storage system is monitored, and the load level is the utilization rate of a plurality of CPU cores in the storage system. An overloaded state is detected based on the utilization rate of one or more CPU cores exceeding a threshold, and the overloaded state is caused by the duplication of one or more I / O queues from a plurality of host computers accessing a single CPU core in the storage system. In response to detecting the overloaded state, a new I / O queue is selected on a second CPU core in the storage system, and the second CPU core has a utilization rate lower than a second threshold. A recommendation is sent to the host computer, and the recommendation is to move I / O traffic from the first CPU core to the new I / O queue on the second CPU core to balance the load level of the storage system.

[0006] Another aspect of the present invention includes a computer-implemented method for storage-level load balancing. In a second embodiment, in response to receiving a command from a host computer to establish an I / O queue pair, processor resources and memory resources are allocated in a storage system, and the storage system implements a Non-Volatile Memory Express over Fabrics (NVMe-oF) architecture. An overload condition is detected on a CPU core in the storage system, and the overload condition is an overlap of multiple host computers using the same I / O queue pair. In response to detecting the overload condition, a recommendation is sent to the host computer, and the recommendation is to move I / O traffic from a first CPU core to a new I / O queue on a second CPU core to balance the load level of the storage system.

Brief Description of the Drawings

[0007]

Figure 1

Figure 2

Figure 3a

Figure 3b

Figure 3c

Figure 4

Figure 5

DETAILED DESCRIPTION OF THE INVENTION

[0008] With the explosive increase in data volume and data usage in modern data processing systems, new methods are needed to improve the throughput of data transfer between hosts and storage and to reduce latency in modern systems. In a typical system, multiple transport channels and protocols coexist in one storage system, which may include NVMe Remote Direct Memory Access (NVMe-RDMA), NVMe over Fiber Channel (NVMe-FC), Fiber Channel-to-Small Computer System Interface (FC-SCSI), Fiber Channel over Ethernet (FCoE), Internet Small Computer Systems Interface (iSCSI), and the like.

[0009] NVMe is a storage protocol designed to speed up data transfer between servers, storage devices, and flash controllers that generally use the PCIe bus as a transfer mechanism. The NVMe specification provides a register interface and command set that enable high-performance I / O. NVMe is a replacement for the traditional Small Computer System Interface (SCSI) standard (and other standards such as SAS and SATA) in data transfer between hosts and storage systems. One of the major advantages of NVMe-based PCIe flash over SAS- or SATA-based SSDs is that it reduces the access latency in the host software stack, resulting in higher Input / Output Operations Per Second (IOPS) and lower CPU utilization.

[0010] NVMe supports parallel I / O processing by multi-core processors, and by speeding up I / O dispatching, it reduces I / O latency. Multiple CPU cores process I / O requests simultaneously, so CPU resources are optimally utilized and system performance is improved. Additionally, NVMe is designed to reduce the number of CPU instructions used per I / O. Also, NVMe supports 64,000 commands in a single message queue and can utilize up to 65,535 I / O queues.

[0011] NVMe over Fabrics (NVMe-oF) is an extension of local PCIe NVMe that enables the high performance and low latency benefits provided by NVMe over a network fabric rather than a local connection. Servers and storage devices can be connected via an Ethernet network or Fibre Channel (FC), both of which support NVMe commands over the fabric and extend the benefits of the NVMe protocol to interconnected system components. The design goal of NVMe-oF is said to be to add no more than 10 microseconds of latency to the communication between an NVMe host computer and a network-connected NVMe storage device, in addition to the latency associated with accessing a PCIe NVMe storage device.

[0012] NVMe-oF supports multiple I / O queues for normal I / O operations from the host to the storage system. NVMe supports a maximum of 65,535 queues, with each queue having a maximum of 65,535 entries. It is the role of the host driver to create queues after the connection is established. When the host connects to the target system, a special-purpose queue called the Admin Queue is created. As the name implies, the Admin Queue is used to transfer control commands from the initiator to the target device. Once the Admin Queue is created, it is used by the host to create I / O queues based on system requirements. The host can establish multiple I / O queues for a single controller with the same NQN (NVMe Qualified Name, used to identify a remote NVMe storage target) and map multiple namespaces (or volumes) to it. Once the I / O queues are established, I / O commands are sent to the I / O Submission Queue (SQ), and I / O responses are collected from the Completion Queue (CQ). These I / O queues can be added or removed using control commands sent via the Admin Queue for that session.

[0013] Upon receiving an I / O queue creation command, the target device performs an initial system check of the maximum supported queue and other related fields, creates an I / O queue, and assigns the I / O queue to a CPU core on the storage controller. Next, the target device sends a response to the queue creation request via the Admin Completion queue. Each I / O queue is assigned to a different CPU core on the storage controller. This enables parallel processing and improves the system throughput. The core assignment logic is implemented in the target storage controller, and the mapping of the I / O queue to the CPU core is executed based on a predefined policy in the storage controller.

[0014] In the current technology, performance degradation due to queue duplication has become a problem. NVMe can support approximately 65,535 queues that can be assigned to different CPU cores to achieve parallelism. When the host issues a command to establish an I / O queue pair with the storage system, the storage system allocates processor resources and memory resources to the I / O queue pair. For example, consider the case where two or more hosts have established connections to a common NVMe target. The I / O queues created by multiple hosts are likely to start duplicating on individual CPU cores. That is, the host "A" primary I / O queue on Core1 may overlap with the host "B" primary I / O queue on Core1. In such a scenario, the I / O workloads transmitted on the NVMe queue from the I / O queue pairs from both hosts are provided by a single core in the storage controller. This reduces the parallelism on the storage controller side and affects the I / O performance of the host application. In the current state-of-the-art technology, since there is no means to tie the allocation of CPU cores to the expected workload, the I / O load can become highly unbalanced among the available CPU cores in the storage controller node. Since each CPU core is shared by multiple I / O queues, there is no means to detect the workload imbalance due to queue duplication from one or more hosts, nor is there a means to notify the server of the workload imbalance. Also, when multiple hosts are connected to the storage target via NVMe queues, since the I / O workloads of the hosts are different, some CPU cores may become overloaded and some CPU cores may become underloaded. Also, there is no mechanism for the storage system to predict how much load will occur in each queue during I / O queue generation. In the host multipathing driver, the host will use a specific I / O queue as the primary queue. When multiple hosts connect their primary queues to the same CPU core, that CPU core will become overloaded, and the applications accessing the data will experience increased I / O latency and will not be able to gain the benefits of parallelization.

[0015] As a result of I / O queue duplication, the load among CPU cores becomes unbalanced, which may reduce IOPS. When the host is executing a small-scale I / O-intensive workload, the severity of this overhead due to queue duplication worsens, and at peak workloads, it may cause a slowdown in the application along with unexpected I / O latency issues. Also, the imbalance of CPU cores in the entire storage controller system loads some CPU cores while other CPU cores are idle, reducing parallel processing and increasing overall latency and delay, which also causes performance problems in the storage controller.

[0016] In various embodiments, the present invention solves this problem by detecting duplicate I / O queues in CPU core allocation within an NVMe storage controller and rebalancing the allocation of I / O queues to CPU cores. In one embodiment, the queue distribution program monitors the queues, workloads, and CPU core availability established for all available CPU cores. When encountering a situation of queue duplication, the queue distribution program determines the CPU workload and load imbalance. The queue distribution program identifies the I / O queues connected to the CPU cores and analyzes the I / O queues of the IOPS workload with a high bandwidth utilization rate. Since the IOPS workload is susceptible to the influence of the CPU, the queue distribution program collects this information and maps the CPU consumption for each I / O queue connected to the overloaded CPU core. In one embodiment, the queue distribution program traverses all the I / O queues created from the same host and analyzes their workloads as well.

[0017] In one embodiment, based on the collected workload information, the queue distribution program determines which I / O queue workloads can be increased to obtain better performance. The queue distribution program achieves this by performing symmetric workload distribution of the workloads of the I / O queues on the storage system.

[0018] In one embodiment, when the queue distribution program makes a transfer decision for a new I / O workload, the information is sent as a signal to the management control unit of the NVMe controller and as an asynchronous notification of the queue duplication status to the host. This Advanced Error Reporting (AER) message includes the I / O queue ID (IOQ_ID) for which the storage system plans to move traffic to balance the CPU workload.

[0019] When the signal is sent to the host, the host's NVMe driver will decide whether to continue with the current I / O transmission policy or adopt the proposal from the queue distribution program to prioritize a specific IOQ. If the host decides to adopt the proposal from the queue distribution program, the path policy of the IOQ is adjusted by the host-side NVMe driver. In some examples, if the host can tolerate a performance degradation, if the host can tolerate a decrease in the total amount of IOPS, or if the host does not want to change the IOQ policy for other reasons, the proposal is rejected and a signal notifying the rejection is sent to the queue distribution program. In one embodiment, when the queue distribution program receives the rejection signal, the queue distribution program sends an AER message to another host to shift its I / O workload from the overloaded CPU core. In this way, both the queue distribution program and the host become parties to the decision, and the distribution of the workload is well achieved by sending a signal to the second host.

[0020] Advantages of the present invention include a reduction in queue duplication bottlenecks, an improvement in performance, an increase in IOPS, an avoidance of IOQ recreation, and an improvement in load distribution among CPU cores.

[0021] In the present invention, since the host IOQ priority is changed, the queue duplication bottleneck is reduced, and the imbalance of CPU cores is reduced or eliminated.

[0022] In the case of queue duplication where the host is performing I / O simultaneously, as the cores process each queue one by one, the performance deteriorates. Therefore, the present invention brings better performance. However, when two queues belong to different hosts, the present invention can rebalance the I / O queues to avoid overall performance degradation.

[0023] Since the present invention avoids queue duplication situations, it leads to an increase in IOPS. Consequently, the host I / O turnaround time is shortened, and the overall IOPS is improved.

[0024] In the present invention, without detaching the IOQ from the storage system or the host, only an on-the-fly target change instruction is given to the host NVMe driver, thus avoiding the recreation of the IOQ. Therefore, the storage-level workload is dispersed, and a performance gain is transparently created.

[0025] The present invention achieves a greater balance in the load across all CPU cores in the storage system. Consequently, the storage system is in a more balanced state, resulting in improved load distribution among the CPU cores.

[0026] FIG. 1 is a functional block diagram showing a distributed data processing environment generally designated 100 suitable for the operation of a queue distribution program 112 according to at least one embodiment of the present invention. As used herein, the term "distributed" describes a computer system that includes a plurality of physically different devices that operate together as a single computer system. FIG. 1 provides only an example of one implementation and is not meant to impose any limitations with respect to the environments in which different embodiments may be implemented. Many modifications to the illustrated environment can be made by those skilled in the art without departing from the scope of the present invention as set forth by the claims.

[0027] In various embodiments, the distributed data processing environment 100 includes a plurality of host computers. In the embodiment depicted in FIG. 1, the distributed data processing environment 100 includes hosts 130, 132, and 134, all of which are connected to network 120. Network 120 can be, for example, a telecommunications network, a local area network (LAN), a wide area network (WAN) such as the Internet, or a combination of these three, and can include wired, wireless, or fiber optic connections. Network 120 can include one or more wired or wireless or both networks that can receive and transmit data, voice, or video signals, or combinations thereof, including multimedia signals that include voice, data, and video information. Generally, network 120 can be any combination of connections and protocols that will support communication between hosts 130, 132, 134, and other computing devices (not shown) within the distributed data processing environment 100.

[0028] In various embodiments, host 130, host 132, and host 134 can each be a stand-alone computing device, a management server, a web server, a mobile computing device, or any other electronic device or computing system capable of receiving, transmitting, and processing data. In one embodiment, host 130, host 132, and host 134 can each be a personal computer, a desktop computer, a laptop computer, a netbook computer, a tablet computer, a smartphone, or any programmable electronic device capable of communicating with other computing devices (not shown) within distributed data processing environment 100 via network 120. In another embodiment, host 130, host 132, and host 134 can each represent a server computing system that utilizes multiple computers as a server system, such as in a cloud computing environment. In yet another embodiment, host 130, host 132, and host 134 can each represent a computing system that utilizes clustered computers and components (e.g., database server computers, application server computers, etc.) that function as a single pool of seamless resources when accessed within distributed data processing environment 100.

[0029] In various embodiments, the distributed data processing environment 100 also includes a storage system 110 connected to hosts 130, 132, and 134 via a fabric 140. The fabric 140 can be, for example, an Ethernet fabric, a Fibre Channel fabric, Fibre Channel over Ethernet (FCoE), or an InfiniBand® fabric. In another embodiment, the fabric 140 can include any of the RDMA technologies including InfiniBand, RDMA over Converged Ethernet (RoCE), and iWARP. In other embodiments, the fabric 140 can be any fabric capable of interfacing hosts to a storage system as known to those skilled in the art.

[0030] In various embodiments, the storage system 110 can be a stand-alone computing device, a management server, a web server, a mobile computing device, or any other electronic device or computing system capable of receiving, transmitting, and processing data. In some embodiments, the storage system 110 can be connected to a network 120 via the fabric 140.

[0031] In one embodiment, the storage system 110 includes a queue dispersion program 112. In one embodiment, the queue dispersion program 112 is a program, application, or sub-program of a larger program for intelligently selecting a transport channel between protocols based on drive type.

[0032] In one embodiment, the storage system 110 includes an information repository 114. In one embodiment, the information repository 114 may be managed by a queue distribution program 112. In an alternative embodiment, the information repository 114 may be managed by the operating system of the storage system 110, either alone or in conjunction with the queue distribution program 112. The information repository 114 is a data repository capable of storing, collecting, comparing, or combining information, or a combination thereof. In some embodiments, the information repository 114 is located external to the storage system 110 and is accessed through a communication network such as a fabric 140. In some embodiments, the information repository 114 is stored on the storage system 110. In some embodiments, the information repository 114 may exist on another computing device (not shown), provided that the information repository 114 is accessible by the storage system 110. The information repository 114 can include, for example, transport channel and protocol data, protocol class data, drive type and drive tier data, link connection data, transport channel tables, raw data transferred between a host initiator and a target storage system, other data received by the queue distribution program 112 from one or more sources, and data created by the queue distribution program 112.

[0033] As is known in the art, the information repository 114 may be implemented using any volatile or non-volatile storage medium for storing information. For example, the information repository 114 may be implemented in a tape library, an optical library, one or more independent hard disk drives, multiple hard disk drives in a redundant array of independent disks (RAID), a SATA drive, a solid state drive (SSD), or random access memory (RAM). Similarly, the information repository 114 may be implemented in any suitable storage architecture known in the art, such as a relational database, an object-oriented database, or one or more tables.

[0034] Figure 2 is an example of the mapping of I / O queues to CPU cores in a basic NVMe storage system according to an embodiment of the present invention. In one embodiment, the storage system 200 is an example of one possible configuration of the queue mapping of the storage system 110 of FIG. 1. In one embodiment, the processor in the storage system 200 has a controller management core 210. In one embodiment, when a host is connected to a target system, a special-purpose queue called an Admin Queue is created during association. The Admin Queue is used to transfer control commands from an initiator to a target device. In one embodiment, the Admin Queue in the controller management core 210 consists of an Admin Submission Queue for sending I / O requests to the I / O queue and an Admin Completion Queue for receiving completion messages from the I / O queue.

[0035] In a typical storage system, there is one or more CPUs, and each CPU has multiple CPU cores. In the example shown in FIG. 2, the processor in storage system 200 has n cores, depicted here as Core_0 212, Core_1 214 to Core_n-1 216. In some embodiments of the present invention, each CPU core has an I / O Submission Queue for sending requests to an I / O queue and an I / O Completion Queue for receiving completion messages from the I / O queue. The exemplary storage system of FIG. 2 also includes a controller 220, which is a controller of the storage system according to some embodiments of the present invention.

[0036] Note that in the example depicted in FIG. 2, only a single I / O queue assigned to each CPU core is shown. In a typical embodiment of the present invention, multiple I / O queues are assigned to each CPU core. This more typical embodiment is illustrated below in FIGS. 3a - 3c.

[0037] FIG. 3a is an explanatory diagram of a typical storage configuration generally referred to as 300, and shows an example of the problem statement from the above. In this example, host A 310 and host B 312 are examples of hosts (130-134) in the distributed data processing environment 100 of FIG. 1. Fabric 320 is a fabric that enables communication between any number of host devices and a storage system. The various communication fabrics that can constitute the fabric 320 in various embodiments of the present invention have been listed above. Storage system 330 is an example of a possible embodiment of storage system 110 in FIG. 1. Storage system 330 includes a disk subsystem 336, a virtualization and I / O management stack 337, an NVMe queue manager 338, and a CPU 335. CPU 335 includes cores 331, 332, 333, and 334, and two I / O queues are connected to each core. In other embodiments, any number of I / O queues may be connected to cores 331-334 up to the maximum number of queues supported as described above. Connections 321 and 322 are examples of connections between the CPU cores and the NVMe-oF I / O queues in the host.

[0038] In this example, both host A and host B are connected to the storage system, and I / O queues are established to all four CPU cores by the host. In this example, the queues of A1 and B1 have a higher I / O workload than the other queues, and thus are overloaded. As a result, the entire system becomes unbalanced and resources are underutilized.

[0039] Figure 3b is an example of the system of Figure 3a, but incorporates one embodiment of the present invention. In this example, the storage system 330 sends an AER message to host B to notify that core 331 is overloaded and core 332 is underutilized. By moving traffic from core 331 to core 332, the system is balanced and performance is improved. In this example, both in-band signaling 341 and out-of-band signaling 342 are shown. In one embodiment, the out-of-band signaling 342 uses an out-of-band API instance 339 to communicate with the host. In one embodiment, either in-band signaling or out-of-band signaling is used. In another embodiment, both in-band signaling and out-of-band signaling are used.

[0040] Figure 3c illustrates an example of the system of Figure 3a, but incorporates one embodiment of the present invention. In this example, the storage system 330 moves the traffic that was previously on core 331 of queue B1 (via connection 322) in Figure 3b to core 332 and queue B2 (via connection 323), which was previously underutilized, thereby rebalancing the utilization of the CPU cores and I / O queues, increasing throughput, and decreasing latency.

[0041] Figure 4 is a flowchart diagram of a workflow 400 depicting the operational steps of a queue distribution program 112 to improve workload management in an IOQ subsystem. In one embodiment, the queue distribution program 112 continuously monitors the CPU core utilization for all CPU cores in the NVMe system using a daemon that collects information regarding the CPU cores, and checks the CPU core consumption for all available CPU cores. In one embodiment, the queue distribution program 112 determines whether one or more CPU cores are detected as being overloaded and whether another set of one or more CPU cores are detected as being underutilized. In one embodiment, if the queue distribution program 112 determines that one or more CPU cores are detected as being overloaded and another set of one or more CPU cores are detected as being underutilized, the queue distribution program 112 will use the daemon to send a signal containing an imbalance message to the NVMe controller. In one embodiment, upon receiving a CPU_IMBALANCE message from the monitoring daemon, the queue distribution program 112 will collect I / O statistics by traversing all I / O queues connected to the overloaded CPU cores and accessing the data access map maintained by the storage controller. In one embodiment, the queue distribution program 112 analyzes all I / O queues for a host that is part of the overloaded CPU cores and captures other IOQ information that would be considered for overload distribution. In one embodiment, the queue distribution program 112 selects a new IOQ recommended for I / O distribution. In one embodiment, the queue distribution program 112 uses an IOQ manager to map the new IOQ_ID to the NVMe Qualified Name. In one embodiment, the queue distribution program 112 generates an AER message with the new IOQ_ID proposed for the specified NQN to recommend shifting the workload to this IOQ. In one embodiment, the queue distribution program 112 receives a response from the host having the new IOQ selected in step 412.In one embodiment, the queue distribution program 112 determines whether the host has accepted the recommendation. In one embodiment, the queue distribution program changes the host IOQ priority setting and sends more workloads onto the queue with the specified IOQ_ID.

[0042] In an alternative embodiment, the steps of the workflow 400 may be performed by any other program while collaborating with the queue distribution program 112. It should be understood that the embodiments of the present invention are provided to improve workload management at least in the IOQ subsystem. However, FIG. 4 provides only an illustration of one implementation and does not imply any limitations regarding the environment in which different embodiments may be implemented. Many modifications to the depicted environment can be made by those skilled in the art without departing from the scope of the present invention as described by the claims.

[0043] The queue distribution program 112 monitors the CPU core utilization rate (step 402). In step 402, the queue distribution program 112 continuously monitors the CPU core utilization rate of all CPU cores in the NVMe system using a monitoring daemon that collects information including the CPU core utilization rate and I / O queue resource availability and utilization rate of all available CPU cores. In one embodiment, the queue distribution program 112 uses a monitoring daemon that runs in parallel with the NVMe controller and the queue manager to monitor the queues, workloads, and CPU core availability established for all available CPU cores. In one embodiment, the queue distribution program 112 collects CPU core utilization rate data from the storage system configuration map and the storage system utilization rate table.

[0044] The queue distribution program 112 determines whether a CPU core is in an overloaded state (determination block 404). In one embodiment, the queue distribution program 112 determines whether one or more CPU cores are in an overloaded state and whether another set of one or more CPU cores is underutilized. In one embodiment, overutilization and underutilization are detected using predefined thresholds. In one embodiment, when a queue duplication situation occurs (for example, as shown in FIG. 3a, host A310 and host B312 are connected to the same CPU core), the queue distribution program 112 determines whether the CPU workload and load imbalance exceed predefined thresholds. For example, the threshold can be set such that if the utilization rate is greater than 80%, the CPU core is considered overutilized. In one embodiment, the predefined threshold is a system default. In another embodiment, the predefined threshold is received from the user at runtime.

[0045] In another embodiment, the queue distribution program 112 determines that one or more CPU cores are in an overloaded state by measuring the average utilization rate of each core over a certain period. In this embodiment, if the average utilization rate of the CPU core exceeds a certain period threshold, the queue distribution program 112 determines that the CPU core is in an overloaded state. For example, the threshold can be set such that if the core has exceeded a utilization rate of 50% for over one minute, the CPU core is considered overloaded. In one embodiment, the average utilization rate is a system default. In another embodiment, the average utilization rate is received from the user at runtime. In one embodiment, the period is a system default. In another embodiment, the period is received from the user at runtime.

[0046] In yet another embodiment, when the queue distribution program 112 determines that a CPU core is overloaded and the utilization rate of the core suddenly increases in a short period of time. In this embodiment, if the increase in the usage rate of the CPU core exceeds the threshold increase rate within the specified period, the queue distribution program 112 determines that the CPU core is overloaded. For example, the threshold may be that if the core usage rate increases by 30% within 10 seconds, the CPU core is overloaded. In one embodiment, the threshold increase rate is the system default. In another embodiment, the threshold increase rate is received from the user at runtime. In one embodiment, the specified period is the system default. In another embodiment, the specified period is received from the user at runtime.

[0047] In one embodiment, when CPU imbalance is confirmed based on the cumulative consumption rate, the queue distribution program 112 identifies the I / O queues connected to the imbalanced CPU cores and analyzes the I / O queues of the IOPS workloads with high bandwidth utilization rates. In one embodiment, the threshold for high bandwidth utilization is the system default. In another embodiment, the threshold for high bandwidth utilization is a value set by the user at runtime. Since the IOPS workload is sensitive to the CPU, the queue distribution program 112 collects this information and maps the CPU consumption for each I / O queue attached to the overloaded CPU core.

[0048] If the queue distribution program 112 determines that one or more CPU cores are detected to be overloaded and another set of one or more CPU cores is detected to be under-utilized (decision block 404 , yes branch), the queue distribution program 112 proceeds to step 406. If the queue distribution program 112 determines that one or more CPU cores are not detected to be overloaded or another set of one or more CPU cores is not detected to be under-utilized (decision block 404 , no branch), the queue distribution program 112 returns to step 402 to continue monitoring the CPU core usage rate.

[0049] The queue distribution program 112 sends an imbalance message (step 406). In one embodiment, the monitoring daemon sends a signal including the imbalance message to the NVMe controller. In one embodiment, the imbalance message includes a CPU core where an overload has been detected. In another embodiment, the imbalance message includes a CPU core detected to be underutilized. In yet another embodiment, the imbalance message includes both a CPU core where an overload has been detected and a CPU core detected to be underutilized. In some embodiments, the imbalance message includes the utilization rates of the cores where an overload has been detected and the cores detected to be underutilized. In one embodiment, the monitoring daemon uses the Admin Submission Queue in the CPU controller management core, such as the controller management core 210 from FIG. 2, to send a signal to the NVMe controller.

[0050] The queue distribution program 112 traverses the I / O queues (step 408). In one embodiment, upon receiving a CPU_IMBALANCE message from the monitoring daemon, the queue distribution program 112 traverses all the I / O queues connected to the overloaded CPU core and collects I / O statistics by accessing the data access map (bandwidth and Input / Output Operations per Second (IOPS) operations) maintained by the storage controller.

[0051] In one embodiment, the queue distribution program 112 examines all the other CPU cores in the storage system to determine which cores have additional bandwidth. In one embodiment, the queue distribution program 112 determines the utilization rates of all the CPU cores in the storage system to determine which cores are underutilized and which cores could potentially have new I / O queues assigned to them to balance the storage system.

[0052] The queue distribution program 112 analyzes all host I / O queues on the overloaded CPU core (step 410). In one embodiment, the queue distribution program 112 analyzes all I / O queues for the hosts that are part of the overloaded CPU core and captures other IOQ information. In one embodiment, the queue distribution program 112 uses the IOQ information to determine the available options for overload distribution. In one embodiment, the IOQ information includes the CPU core where the overload was detected and the CPU core detected as underutilized to determine the available options for overload distribution. In one embodiment, the IOQ information includes the utilization rates of the core where the overload was detected and the core detected as underutilized to determine the available options for overload distribution. In yet another embodiment, the IOQ information includes the CPU core where the overload was detected and the CPU core detected as underutilized, along with the utilization rates of the cores, to determine the available options for overload distribution.

[0053] The queue distribution program 112 selects a new IOQ that can accept additional workload (step 412). In one embodiment, the queue distribution program 112 selects a new IOQ recommended for I / O distribution. In one embodiment, the queue distribution program 112 selects a new IOQ based on the workload information collected from each IOQ in step 410. In one embodiment, the queue distribution program 112 selects a new IOQ based on the CPU core utilization rate for the new IOQ being less than a threshold. In one embodiment, the predetermined threshold is a system default. In another embodiment, the predetermined threshold is received from the user at runtime. In one embodiment, the queue distribution program 112 selects a new IOQ based on symmetric workload distribution of the workload of the I / O queues on the storage system. For example, assume that queue A1 and queue B1 are located on the same CPU core and are generating a high workload. Since the CPU cores associated with queue A1 and queue B1 are overloaded, the queue distribution program 112 will check all the I / O queues created by host A and host B. In this example, the queue distribution program 112 will then classify these I / O queues into existing CPUs and associated workloads. In this example, the queue distribution program 112 determines that IOQ A2 and B2 resident on core 2 have few queues and low CPU workload, so the queue distribution program 112 moves the workload of IOQ (either A1 or B1) to core 2.

[0054] In one embodiment, the queue distribution program 112 selects a plurality of IOQs that can each be used to balance the workload and assigns priorities to the IOQs based on the available workload. In one embodiment, the queue distribution program 112 selects the highest-priority available IOQ to recommend IOQ balance. In one embodiment, the highest-priority available IOQ is the IOQ connected to the CPU core with the lowest utilization rate. In another embodiment, the highest-priority available IOQ is determined by selecting the IOQ connected to the CPU core when there are no other IOQs connected to that core.

[0055] The queue distribution program 112 maps the new IOQ_ID to the NQN (step 414). In one embodiment, the queue distribution program 112 uses an IOQ manager to map the new IOQ_ID selected in step 412 to the NQN of a remote storage target, such as the storage system 330 of FIGS. 3a-3c.

[0056] The queue distribution program 112 sends an AER with an IOQ_ID to the specified NQN to shift the workload (step 416). In one embodiment, the queue distribution program 112 generates an AER message with a proposed new IOQ_ID for the specified NQN to recommend shifting the workload to this IOQ. In one embodiment, when the queue distribution program 112 makes a transfer decision for a new I / O workload, that information is sent as a signal to the management control unit of the NVMe controller. In one embodiment, the queue distribution program 112 sends an asynchronous notification of the queue duplication situation to the host via internal communication or via protocol-level communication (via the NVMe Asynchronous Event Request Command). In one embodiment, this message includes the IOQ_ID for which the storage system plans to move traffic to balance the CPU workload. Since the queue distribution program 112 has already established a new I / O queue with the new IOQ_ID, the queue distribution program 112 plans to send more traffic on the queue proposed by the host to obtain more performance and greater parallelism.

[0057] In one embodiment, the communication between the queue distribution program 112 and the host notification function can be performed via an out-of-band (OOB) application program interface (API) that implements the function of communicating between the host and the storage controller cluster system using an OOB protocol. For example, in FIG. 3b, signal 342 represents out-of-band communication. In another embodiment, the communication between the queue distribution program 112 and the host notification function can be via in-band communication using the NVMe standard. In this embodiment, the queue duplication information and the actuator signal are programmatically passed as part of the protocol frame. For example, in FIG. 3b, signal 341 represents in-band communication.

[0058] The queue distribution program 112 receives a response from the host (step 418). In one embodiment, the queue distribution program 112 receives a response from the host of the new IOQ selected in step 412. In the example of FIG. 3c, the host of the new IOQ is host B312. In one embodiment, the response may be either that the host accepts the recommendation or that the host rejects the recommendation.

[0059] The queue distribution program 112 determines whether the host has accepted the new IOQ_ID (decision block 420). In one embodiment, the queue distribution program 112 determines whether the host has accepted the recommendation. In one embodiment, when a signal is sent to the host, the host NVMe driver determines whether to continue the current I / O transmission policy or adopt the proposal from the queue distribution program 112 to prioritize a specific IOQ. In one embodiment, if the host determines to adopt the proposal from the queue distribution program 112, the IOQ path policy is adjusted by the host-side NVMe driver to send more traffic on the proposed IOQ_ID to obtain more performance. All new traffic from the server / host will be sent via the newly assigned IOQ_ID to the new CPU core, thus improving the performance of the host application.

[0060] In another embodiment, if the host can tolerate performance degradation, if the host can tolerate a decrease in the total amount in IOPS, or if the host does not want to change the IOQ policy for other reasons, the proposal is rejected and a signal is sent to notify the queue dispersion program 112 of the rejection. In one embodiment, when the queue dispersion program 112 receives the rejection signal, the queue dispersion program 112 sends an AER message to another host to shift its I / O workload from the overloaded CPU core. In this way, both the queue dispersion program and the host become parties to the decision, and workload dispersion is successfully achieved by sending a signal to the second host. For example, if queue A1 and queue B1 overlap and the queue dispersion program 112 determines to disperse the workload by shifting the load from queue A1 or queue B1, the queue dispersion program sends a signal to host A to use queue A2. If host A rejects the proposal, the queue dispersion program sends a signal to host B to shift the workload to B2. This process is repeated until the host accepts the request to change to the new IOQ. This serialization is performed to prevent a situation where a new imbalance occurs as a result of multiple hosts changing the priority IOQ simultaneously.

[0061] If the queue dispersion program 112 determines that the host has accepted the recommendation (decision block 420 , yes branch), the queue dispersion program 112 proceeds to step 422. In one embodiment, if the queue dispersion program 112 determines that the host has not accepted the recommendation ( Judgment block 420 , no branch), the queue dispersion program 112 returns to step 412 to select a different IOQ. In another embodiment, since the workload is not sensitive to IOPS, if the queue dispersion program 112 determines that the host has not accepted the recommendation ( Judgment block 420 , no branch), the queue dispersion program 112 ends this cycle.

[0062] The queue distribution program 112 receives an ACK from the host indicating that the IOQ change has been accepted (step 422). In one embodiment, if the queue distribution program 112 determines that the host has accepted the recommendation, the queue distribution program changes the host IOQ priority setting to send more workload on the queue with the new IOQ_ID. In one embodiment, the queue distribution program 112 receives an ACK signal in the acceptance message from the target. Thereby, the rebalancing cycle is completed.

[0063] In one embodiment, the queue distribution program 112 ends this cycle.

[0064] FIG. 5 is a block diagram showing the components of a storage system 110 suitable for the queue distribution program 112 according to at least one embodiment of the present invention. FIG. 5 shows a computer 500, one or more processors 504 (including one or more computer processors), a communication fabric 502, a memory 506 including a random access memory (RAM) 516 and a cache 518, a persistent storage 508, a communication unit 512, an I / O interface 514, a display 522, and an external device 520. FIG. 5 is merely illustrative of one embodiment and is not meant to impose any limitation with respect to the environments in which different embodiments may be implemented. Many modifications are possible to the illustrated environment.

[0065] As shown, the computer 500 operates on a communication fabric 502 that provides communication between a computer processor 504, a memory 506, a persistent storage 508, a communication unit 512, and an input / output (I / O) interface 514. The communication fabric 502 may be implemented in an architecture suitable for passing data or control information between the processor 504 (e.g., a microprocessor, a communication processor, and a network processor), the memory 506, the external device 520, and any other hardware components within the system. For example, the communication fabric 502 may be implemented with one or more buses.

[0066] Memory 506 and persistent storage 508 are computer-readable storage media. In the illustrated embodiment, memory 506 includes RAM 516 and cache 518. Generally, memory 506 can include any suitable volatile or non-volatile computer-readable storage media. Cache 518 is fast memory and improves the performance of processor 504 by holding recently accessed data, and data close to recently accessed data, from RAM 516.

[0067] The program instructions of queue distribution program 112 may be stored in persistent storage 508, or more generally, in any computer-readable storage media, for execution by one or more of each computer processor 504 via one or more memories of memory 506. Persistent storage 508 may be a magnetic hard disk drive, a solid state disk drive, a semiconductor memory device, a read-only memory ROM, an erasable programmable ROM (EPROM), flash memory, or any other computer-readable storage media capable of storing program instructions or digital information.

[0068] The media used by persistent storage 508 may be removable. For example, a removable hard drive may be used for persistent storage 508. Other examples include optical disks, magnetic disks, thumb drives, and smart cards, which are inserted into a drive for transfer to another computer-readable storage media that is also part of persistent storage 508.

[0069] In these examples, communication unit 512 enables communication with other data processing systems or devices. In these examples, communication unit 512 includes one or more network interface cards. Communication unit 512 may enable communication using either or both physical and wireless communication links. In the context of some embodiments of the present invention, various sources of input data may be physically remote from computer 500 such that input data is received and output may similarly be transmitted via communication unit 512.

[0070] I / O interface 514 enables input and output of data with other devices connectable to computer 500. For example, I / O interface 514 enables connection with (one or more) external devices 520 such as a keyboard, keypad, touch screen, microphone, digital camera, or other suitable input device or combinations thereof. Also, external device 520 may include a portable computer-readable storage medium such as, for example, a thumb drive, portable optical disk, portable magnetic disk, and memory card. Software and data (e.g., queue dispersion program 112) used to implement embodiments of the present invention may be stored on such portable computer-readable storage media and loaded into persistent storage 508 via I / O interface 514. I / O interface 514 also connects to display 522.

[0071] Display 522 realizes a mechanism for displaying data to a user and may be, for example, a computer monitor. Display 522 may also function as a touch screen such as that of a tablet computer display.

[0072] The programs described in this specification are identified based on the applications in which the programs are implemented in specific embodiments of the present invention. However, the names of specific programs in this specification are for convenience only, and thus, it should be understood that the present invention is not limited to being used only in specific applications that are identified or suggested or both by such names.

[0073] The present invention can be a system, a method, a computer program product, or a combination thereof. The computer program product may include a computer-readable storage medium storing computer-readable program instructions for causing a processor to execute aspects of the present invention.

[0074] The computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. The computer-readable storage medium can be, by way of example, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or a suitable combination thereof. A more specific example of the computer-readable storage medium can include a portable computer diskette, a hard disk, a RAM, a ROM, an EPROM (or flash memory), an SRAM, a CD-ROM, a DVD, a memory stick, a floppy disk, a punch card, a mechanically encoded device that records instructions in a raised structure within a groove, and suitable combinations thereof. As used herein, a computer-readable storage device should not be construed as being a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted through a wire.

[0075] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network, or a wireless network network or a combination thereof). The network is composed of copper wire transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, or edge servers or a combination thereof. The network adapter card or network interface of each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage in a computer-readable storage medium within each computing / processing device.

[0076] The computer-readable program instructions for carrying out the operations of the present invention may be source code or object code written in any combination of one or more programming languages, including assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or a combination of procedural programming languages such as the "C" programming language and object-oriented programming languages such as Smalltalk, C++. The computer-readable program instructions may be executable entirely on the user's computer as a stand-alone software package, or partially on the user's computer. Alternatively, it may be executable partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) can execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize and thereby execute aspects of the present invention.

[0077] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0078] These computer-readable program instructions can be provided to a general-purpose computer, a processor of a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus implement the functions / acts specified in one or more blocks of a flowchart, a block diagram, or both. These computer-readable storage media can also be stored in a computer-readable storage medium connectable to a computer, a programmable data processing apparatus, or other devices that function in a particular manner or a combination thereof, such that the computer-readable storage medium constitutes one of the manufactured articles containing instructions for implementing the aspects of the functions / acts specified in one or more blocks of a flowchart, a block diagram, or both.

[0079] Computer-readable program instructions, such as instructions for performing the functions / acts specified in one or more blocks of a flowchart, a block diagram, or both on a computer, other programmable apparatus, or other device, can also be loaded onto a computer, other programmable data processing apparatus, or other device to execute a series of operational steps on the computer, other programmable apparatus, or other device to produce a computer-implemented process.

[0080] The flowcharts and block diagrams in the figures illustrate the configuration, functionality, and operation of the implementation executable by systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions, which constitutes one or more executable instructions for implementing the specified logical function. In some alternative embodiments, the functions shown in the blocks may occur out of the order shown in the figures. For example, two blocks shown in succession may, in fact, be executed substantially simultaneously, or the blocks may be executed in the reverse order depending on the functions involved. It should also be noted that each block of the block diagram or flowchart diagram, or both, and combinations of blocks of the block diagram or flowchart diagram, or both, can be implemented by a special purpose hardware-based system that performs the specified functions or operations, or a combination of special purpose hardware and computer instructions.

[0081] The description of the various embodiments of the present invention has been presented for purposes of illustration, but is not intended to be exhaustive or to limit the invention to the disclosed embodiments. It will be apparent to those skilled in the art that many modifications and variations are possible without departing from the scope of the invention. The terms used herein have been chosen to best explain the principles of the embodiments, the practical application to techniques found in the marketplace, or the technical improvement, or to enable those skilled in the art to understand the embodiments described herein.

Claims

1. A computer-implemented method for load balancing at the storage level, comprising: monitoring, by one or more computer processors, a load level of a storage system, wherein the load level is a utilization rate of a plurality of CPU cores in the storage system; detecting, by the one or more computer processors, an overload state based on the utilization rate of one or more of the plurality of CPU cores exceeding a first threshold, wherein the overload state is caused by duplication of one or more I / O queues from each of a plurality of host computers accessing a first CPU core among the plurality of CPU cores in the storage system; selecting, by the one or more computer processors, a new I / O queue on a second CPU core among the plurality of CPU cores in the storage system in response to detecting the overload state, wherein the second CPU core has a utilization rate less than a second threshold; transmitting, by the one or more computer processors, a recommendation to a first host computer among the plurality of host computers, wherein the recommendation is to move I / O traffic from the first CPU core to the new I / O queue on the second CPU core to balance the load level of the storage system; A computer-implemented method comprising the steps of performing the above.

2. Recommending the new I / O queue on the second CPU core to the first host computer among the plurality of host computers to balance the load level of the storage system, receiving, by the one or more computer processors, a response from the first host computer among the plurality of host computers; in response to the response being a rejection of the recommendation, recommending, by the one or more computer processors, the new I / O queue on the second CPU core to a second host computer among the plurality of host computers to balance the load level of the storage system; The computer-implemented method according to claim 1, further comprising

3. The monitoring of the load level of the storage system further comprises using a daemon to collect the usage rate of the one or more CPU cores among the plurality of CPU cores, for the computer-implemented method according to claim 1 or claim 2.

4. Detecting the overloaded state collecting, by the one or more computer processors, CPU core usage rate data included in one or more storage system configuration maps and one or more storage system usage rate tables, wherein the one or more storage system configuration maps include an I / O queue configuration for each CPU core of the plurality of CPU cores; analyzing, by the one or more computer processors, the CPU core usage rate data collected in the one or more storage system configuration maps and the one or more storage system usage rate tables to determine the usage rate of each I / O queue in each CPU core of the plurality of CPU cores; The computer-implemented method according to any one of claims 1 to 3, further comprising

5. In response to detecting the overloaded state, selecting a new I / O queue on a second CPU core among the plurality of CPU cores in the storage system, wherein the second CPU core has a usage rate smaller than the second threshold, and the selecting further comprises performing symmetric load balancing of the workload on each I / O queue of the one or more I / O queues on the storage system, for the computer-implemented method according to any one of claims 1 to 4.

6. A computer program for load distribution at the storage level, the computer program comprising monitoring the load level of a storage system, wherein the load level is the usage rate of a plurality of CPU cores in the storage system; Detecting an overload state based on the fact that the usage rate of one or more of the plurality of CPU cores exceeds a first threshold, wherein the overload state is caused by the duplication of one or more I / O queues from each of a plurality of host computers accessing a first CPU core among the plurality of CPU cores in the storage system, and detecting; In response to detecting the overload state, selecting a new I / O queue on a second CPU core among the plurality of CPU cores in the storage system, wherein the second CPU core has a usage rate smaller than a second threshold, and selecting; Sending a recommendation to a first host computer among the plurality of host computers, wherein the recommendation is to move I / O traffic from the first CPU core to the new I / O queue on the second CPU core in order to balance the load level of the storage system, and sending; A computer program for causing a computer to execute.

7. In order to balance the load level of the storage system, recommending the new I / O queue on the second CPU core to the first host computer among the plurality of host computers is Receiving a response from the first host computer among the plurality of host computers; In response to the response being a rejection of the recommendation, in order to balance the load level of the storage system, recommending the new I / O queue on the second CPU core to a second host computer among the plurality of host computers; The computer program according to claim 6, further comprising.

8. Monitoring the load level of the storage system further includes using a daemon to collect the usage rate of the one or more of the plurality of CPU cores, the computer program according to claim 6 or claim 7.

9. Detecting the overload state is Collecting CPU core utilization data included in one or more storage system configuration maps and one or more storage system utilization rate tables, wherein the one or more storage system configuration maps include I / O queue configurations for each CPU core of the plurality of CPU cores, and collecting; Analyzing the CPU core utilization data collected in the one or more storage system configuration maps and the one or more storage system utilization rate tables to determine the utilization rate of each I / O queue in each CPU core of the plurality of CPU cores; The computer program according to claim 6, further comprising.

10. In response to detecting the overload state, selecting the new I / O queue on the second CPU core among the plurality of CPU cores in the storage system, wherein the second CPU core has the utilization rate smaller than the second threshold value, and the selecting further includes performing symmetric load balancing of the workload on each I / O queue of the one or more I / O queues on the storage system. The computer program according to claim 6.

11. A computer system for load distribution at the storage level, wherein the computer system includes One or more computer processors; One or more computer-readable storage media; Program instructions stored in the one or more computer-readable storage media and executed by at least one of the one or more computer processors, and the stored program instructions include Monitoring the load level of the storage system, wherein the load level is the utilization rate of a plurality of CPU cores in the storage system, and monitoring; Detecting an overload state based on the fact that the utilization rate of one or more CPU cores among the plurality of CPU cores exceeds a first threshold value, wherein the overload state is caused by the duplication of one or more I / O queues from each host computer of a plurality of host computers accessing the first CPU core among the plurality of CPU cores in the storage system, and detecting; In response to detecting the overload state, selecting a new I / O queue on a second CPU core among the plurality of CPU cores in the storage system, wherein the second CPU core has the utilization rate smaller than a second threshold; Sending a recommendation to a first host computer among the plurality of host computers, the recommendation being to move I / O traffic from the first CPU core to the new I / O queue on the second CPU core to balance the load level of the storage system; A computer system comprising instructions for performing.

12. Sending a recommendation to the first host computer among the plurality of host computers, the recommendation being to move I / O traffic from the first CPU core to the new I / O queue on the second CPU core to balance the load level of the storage system, is stored in the one or more computer-readable storage media; Receiving a response from the first host computer among the plurality of host computers; In response to the response being a rejection of the recommendation, recommending the new I / O queue on the second CPU core to a second host computer among the plurality of host computers to balance the load level of the storage system; Further comprising one or more program instructions for performing. The computer system according to claim 11.

13. Monitoring the load level of the storage system, wherein the load level is the utilization rate of the plurality of CPU cores in the storage system, and the monitoring further includes using a daemon to collect the utilization rate of the one or more CPU cores among the plurality of CPU cores. The computer system according to claim 11.

14. Detecting the overload state based on the fact that the usage rate of one or more of the plurality of CPU cores exceeds the first threshold value, wherein the overload state is caused by the duplication of one or more I / O queues from each host computer of a plurality of host computers accessing the first CPU core among the plurality of CPU cores in the storage system, and the detecting is stored in the one or more computer-readable storage media, Collecting CPU core usage rate data included in one or more storage system configuration maps and one or more storage system usage rate tables, wherein the one or more storage system configuration maps include I / O queue configurations for each of the plurality of CPU cores, and the collecting, Analyzing the CPU core usage rate data collected in the one or more storage system configuration maps and the one or more storage system usage rate tables to determine the usage rate of each I / O queue in each of the plurality of CPU cores, The computer system according to claim 11, further comprising one or more program instructions for executing.

15. In response to detecting the overload state, selecting the new I / O queue on the second CPU core among the plurality of CPU cores in the storage system, wherein the second CPU core has a usage rate smaller than the second threshold value, and the selecting further includes performing symmetric load balancing of the workload on each of the one or more I / O queues on the storage system. The computer system according to claim 11.

16. A computer-implemented method for load distribution at the storage level, In response to receiving a command for establishing an I / O queue pair from a host computer, allocating processor resources and memory resources in a storage system by one or more computer processors, wherein the storage system implements a Non-Volatile Memory Express over Fabrics (NVMe-oF) architecture, and the allocating, Detecting, by the one or more computer processors, an overloaded state on a first CPU core among a plurality of CPU cores in the storage system, wherein the overloaded state is an overlap of a plurality of host computers using the same I / O queue pair; In response to detecting the overloaded state, sending, by the one or more computer processors, a recommendation to a first host computer among the plurality of host computers, wherein the recommendation is to move I / O traffic from the first CPU core to a new I / O queue on a second CPU core to balance the load level of the storage system, and the second CPU core has a utilization rate lower than a second threshold; A computer-implemented method comprising the step of performing. [

17. ] In response to detecting the overloaded state, sending, by the one or more computer processors, a recommendation to the first host computer among the plurality of host computers, wherein the recommendation is to move I / O traffic from the first CPU core to the new I / O queue on the second CPU core to balance the load level of the storage system, and the sending is: Receiving, by the one or more computer processors, a response from the first host computer among the plurality of host computers; In response to the response being a rejection of the recommendation, recommending, by the one or more computer processors, to a second host computer among the plurality of host computers the new I / O queue on the second CPU core to balance the load level of the storage system; Further comprising: The computer-implemented method according to claim 16. [

18. ] Detecting, by the one or more computer processors, an overloaded state on the first CPU core among the plurality of CPU cores in the storage system, wherein the overloaded state is the overlap of the plurality of host computers using the same I / O queue pair, and the detecting is: collecting, by the one or more computer processors, CPU core utilization data included in one or more storage system configuration maps and one or more storage system utilization tables, wherein the one or more storage system configuration maps include I / O queue configurations for each CPU core of the plurality of CPU cores; analyzing, by the one or more computer processors, the CPU core utilization data collected in the one or more storage system configuration maps and the one or more storage system utilization tables to determine a utilization rate of each I / O queue in each CPU core of the plurality of CPU cores; The computer-implemented method according to claim 16 or claim 17, further comprising.

19. The computer-implemented method according to claim 18, wherein the overload state is based on the utilization rate of one or more of the plurality of CPU cores in the storage system exceeding a first threshold.

20. In response to detecting the overload state, transmitting, by the one or more computer processors, a recommendation to the first host computer among the plurality of host computers, the recommendation being to move I / O traffic from the first CPU core to a new I / O queue on the second CPU core to balance the load level of the storage system, the transmitting further comprising selecting the new I / O queue on the second CPU core among the plurality of CPU cores in the storage system based on the second CPU core having a utilization rate less than a second threshold. The computer-implemented method according to claim 16.

21. A computer system for load distribution at the storage level, the computer system comprising: one or more computer processors; one or more computer-readable storage media; program instructions stored in the one or more computer-readable storage media and executed by at least one of the one or more computer processors, the stored program instructions comprising In response to receiving a command to establish an I / O queue pair from a host computer, allocating processor resources and memory resources in a storage system, wherein the storage system implements a Non-Volatile Memory Express over Fabrics (NVMe-oF) architecture, the allocating, Detecting an overload state on a first CPU core among a plurality of CPU cores in the storage system, wherein the overload state is an overlap of a plurality of host computers using the same I / O queue pair, the detecting, In response to detecting the overload state, sending a recommendation to a first host computer among the plurality of host computers, wherein the recommendation is to move I / O traffic from the first CPU core to a new I / O queue on a second CPU core to balance the load level of the storage system, and the second CPU core has a utilization rate smaller than a second threshold, the sending, A computer system comprising instructions for performing.

22. In response to detecting the overload state, sending a recommendation to the first host computer among the plurality of host computers, wherein the recommendation is to move I / O traffic from the first CPU core to the new I / O queue on the second CPU core to balance the load level of the storage system, the sending is stored in the one or more computer-readable storage media, Receiving a response from the first host computer among the plurality of host computers, In response to the response being a rejection of the recommendation, recommending the new I / O queue on the second CPU core to a second host computer among the plurality of host computers to balance the load level of the storage system, The computer system according to claim 21, further comprising one or more program instructions for performing.

23. Detecting the overload state on the first CPU core among the plurality of CPU cores in the storage system, where the overload state is the duplication of a plurality of host computers using the same I / O queue pair, and the detecting is stored in the one or more computer-readable storage media, Collecting CPU core utilization data included in one or more storage system configuration maps and one or more storage system utilization rate tables, where the one or more storage system configuration maps include I / O queue configurations for each CPU core of the plurality of CPU cores, and the collecting, Analyzing the CPU core utilization data collected in the one or more storage system configuration maps and the one or more storage system utilization rate tables to determine the utilization rate of each I / O queue in each CPU core of the plurality of CPU cores, The computer system according to claim 21, further comprising one or more program instructions for executing the above.

24. The computer system according to claim 23, where the overload state is based on the utilization rate of one or more CPU cores among the plurality of CPU cores in the storage system exceeding a first threshold.

25. In response to detecting the overload state, sending a recommendation to the first host computer among the plurality of host computers, where the recommendation is to move I / O traffic from the first CPU core to a new I / O queue on the second CPU core to balance the load level of the storage system, and the sending is stored in the one or more computer-readable storage media, and based on the second CPU core among the plurality of CPU cores in the storage system having a utilization rate smaller than a second threshold, further comprising one or more program instructions for selecting the new I / O queue on the second CPU core among the plurality of CPU cores in the storage system. The computer system according to claim 21.

Citation Information

Patent Citations

  • Load distribution method and system for computer system

    JP2010079626A

  • Storage system and IO processing control method

    JP2019164510A

  • System and Methodology Providing Workload Management in Database Cluster

    US20090037367A1

  • Storage system and method for connection-based load balancing

    US20170171302A1

  • Storage network tiering

    US20180324250A1