Impact of uneven channel utilization on performance degradation in solid-state drives
By using an arbiter to prioritize and rearrange read requests in PCIe-based SSDs, the issue of uneven channel utilization is addressed, resulting in improved resource allocation and SSD performance.
Patent Information
- Application Number
- DE112015003568
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2014-09-26
- Filing Date
- 2015-08-26
- Publication Date
- 2025-06-05
- Estimated Expiration
- 2035-08-26
AI Technical Summary
In PCIe-based solid state drives, uneven channel utilization occurs when channels have different numbers of NAND chips, leading to inefficient resource allocation and potential bottlenecks in data processing.
An arbiter within the SSD allocates read requests to channels by prioritizing the weakest busy channel and rearranging pending read requests to ensure even channel utilization, thereby optimizing resource allocation and improving performance.
This approach ensures that resources are used efficiently, preventing bottlenecks and improving overall SSD performance by evenly distributing the load across channels.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
BACKGROUNDA solid state drive (SSD) is a data storage device that uses integrated circuitry as memory for persistently storing data. Many types of SSDs use NAND-based or NOR-based flash memory that retains data without power and is one type of nonvolatile memory technology.Communication interfaces may be used to couple SSDs to a host system that includes a processor. Such communication interfaces may include a Peripheral Component Interconnect Express (PCIe) bus. Further details on PCIe can be found in the publication published November 10, 2010 by PCI-SIG, entitled "PCI Express Base Specification Revision 3.0". The most important advantage of SSDs communicating over the PCI bus is improved performance, and such SSDs are referred to as PCIe SSDs.US 2013 / 0 262 745 A1 describes a nonvolatile memory system including a memory controller that receives commands from a host and identifies commands that can be executed in parallel. The order in which the commands are received is recorded so that the responses to the host can be in the same order in which the commands were received.U.S. Pat. No. 8,341,300 B1 describes a memory system having nonvolatile memory devices (NVMDs) coupled to memory channels to share buses, and a memory controller coupled to the memory channels and communicating between the plurality of NVMDs.BRIEF DESCRIPTION OF THE DRAWINGSReferring now to the drawings, wherein like reference numerals represent corresponding parts throughout: FIG. 1 is a block diagram of a computing environment in which a solid state drive is coupled to a host via a PCIe bus; FIG. 2 is another block diagram illustrating how an arbiter allocates read requests in an incoming queue to channels of a solid state drive, in accordance with certain embodiments; FIG. 3 is a block diagram illustrating allocation of read requests in a solid state drive prior to beginning prioritization of the most lightly populated channel and reordering of host commands, in accordance with certain embodiments; FIG. 4 is a block diagram illustrating allocation of read requests in a solid state drive after prioritizing the most sparse channel and rearranging host commands, in accordance with certain embodiments; FIG. 5 is a first flow diagram for preventing non-uniform channel utilization in solid state drives, in accordance with certain embodiments; FIG. 6 illustrates a second flow diagram for preventing non-uniform channel utilization in solid state drives, in accordance with certain embodiments; and FIG. 7 is a block diagram of a computing device, in accordance with certain embodiments.DETAILED DESCRIPTIONIn the following description, reference is made to the accompanying drawings that form a part hereof and illustrate several embodiments. It is to be understood that other embodiments may be utilized and structural and functional changes may be made.The increased performance of PCIe SSDs may be primarily due to the number of channels implemented in the PCIe SSDs. For example, in certain embodiments, certain PCIe SSDs may provide improved internal bandwidth through an 18 channel extended design.In a PCIe-based solid state drive, the PCIe bus from the host to the solid state drive may have a high bandwidth (e.g., 40 gigabytes / second). The PCIe-based solid state drive may have a plurality of channels, each channel having a relatively lower bandwidth than the bandwidth of the PCIe bus. For example, in a solid state drive with 18 channels, each channel may have a bandwidth of about 200 megabytes / second.In certain situations, the number of NAND chips coupled to each channel is an equal number, and in such situations, in the case of random but uniform read requests, the channels may be used approximately the same amount by the host, i.e., each channel is used to process read requests approximately the same amount over a period of time. It should be noted that in many situations, over 95% of the requests from the host to the solid state drive may be read requests, while less than 5% of the requests from the host to the solid state drive may be write requests, and proper mapping of read requests to channels in solid state drives may be important.However, in certain situations, at least one of the channels may have a different number of NAND chips coupled to the channel than the other channels. Such a situation may occur when the number of NAND chips is not a multiple of the number of channels. For example, if there are 18 channels and the number of NAND chips is not a multiple of 18, then at least one of the channels must have a different number of NAND chips coupled to the channel than the other channels. In such a situation, channels coupled to a larger number of NAND chips may be used more heavily than channels coupled to a smaller number of NAND chips. It is assumed that each NAND chip in the solid state drive is of identical construction and has the same storage capacity.In the case of non-uniform utilization of channels, some channels may get into the residue more than others, and the PCIe bus must wait before completing the response to the host until the residue is refurbished.Certain embodiments provide mechanisms to prevent uneven utilization of channels even when at least one of the channels has a different number of NAND chips coupled to the channel than the other channels. This is accomplished by predominantly loading the weakest busy channel with read requests destined for the weakest busy channel and rearranging the processing of pending read requests waiting in a queue in the solid state drive for execution. Since resources are allocated with read requests when loading a read request into a channel by loading the weakest exhausted channels, resources are used only when needed and therefore efficiently used. Thus, certain embodiments improve performance of SSDs.FIG. 1 illustrates a block diagram of a computing environment 100 in which a solid state drive 102 is coupled to a host 104 via a PCIe bus 106, in accordance with certain embodiments. The host 104 may include at least one processor.In certain embodiments, an arbiter 108 is implemented in firmware in the solid state drive 102. In other embodiments, arbiter 108 may be implemented in hardware or software, in any combination of hardware, firmware, or software. Arbiter 108 allocates read requests received from host 104 over PCIe bus 106 to one or more channels of a plurality of channels 110 a, 110 b,..., 110 nof solid state drive 102.In certain embodiments, channels 110 a...110 nare coupled to a plurality of non-volatile memory chips, such as NAND chips, NOR chips, or other suitable non-volatile memory chips. In alternative embodiments, other types of memory chips may also be used, such as phase change memory (PCM) based chips, 3D cross point memory, resistive memory, nanowire memory, ferroelectric transistor random access memory (FeRAM), magnetoresistive random access memory (MRAM), memory including memristor technology, spin torque transfer (STT) MRAM, or other suitable memory.For example, in certain embodiments, channel 110 ais coupled to NAND chips 112 a...112 p, channel 110 bis coupled to NAND chips 114 a...114 q, and channel 110 nis coupled to NAND chips 114 a...114 r. Each of the NAND chips 112 a... 112 p, 114 a... 114 q, 114 a... 114 ris of identical construction. At least one of the channels of the plurality of channels 110 a...110 nhas a different number of NAND chips coupled to the channel than other channels, such that there is a possibility of uneven utilization of the channels 110 a...110 nwhen the read requests from the host 104 are random and uniform.In certain embodiments, the solid state drive 102 may be capable of storing multiple terabytes of data or even more, and the plurality of NAND chips 112 a...112 p, 114 a...114 q, 116 a...116 r, each storing multiple gigabytes of data or even more, may be found in the solid state drive 102. The PCIe bus 106 may have a maximum bandwidth (i.e., data transfer capacity) of 4 gigabytes per second. In certain embodiments, the plurality of channels 110 a...110 nmay be eighteen in number, and each channel may have a maximum bandwidth of 200 megabytes per second.In certain embodiments, arbiter 108 examines the plurality of channels 110 a...110 nin sequence one after the other and, after examining all of the plurality of channels 110 a...110 n, loads the weakest busy channel with read requests destined for the channel to increase the load of the at least busy channel in an attempt to perform equal utilization of the plurality of channels.FIG. 2 illustrates another block diagram 200 of the solid state drive 102 illustrating how the arbiter 108 allocates read requests in an incoming queue 202 to channels 110 a...110 nof the solid state drive 102, according to certain embodiments.Arbiter 108 maintains incoming queue 202, where incoming queue 202 stores read requests received from host 104 over PCIe bus 106. The read requests arrive in an order in the incoming queue 202 and are initially maintained in the same order as the order of arrival of the read requests in the incoming queue 202. For example, a request arriving first may be for data stored in NAND chips coupled to channel 110 band a second request arriving next may be for data stored in NAND chips coupled to channel 110 a. In such a situation, the request arriving first is at the first location of the incoming queue 202, and the request arriving next is the next element in the incoming queue 202.Arbiter 108 also maintains, for each channel 110a...110b, a data structure in which an identification of outstanding read requests processed by the channel is maintained. For example, data structures 204 a, 204 b,...204 nstore the identification of the pending reads processed by the plurality of channels 110 a, 110b,....110n. The outstanding read requests for a channel are the read requests that have been loaded into the channel and that are processed by the channel, i.e., the NAND chips coupled to the channel are used to fetch data corresponding to the read requests that have been loaded into the channel.The solid state drive 102 also maintains a variety of hardware, firmware, and software resources, such as buffers, caches, memory, various data structures, etc. (as represented by reference numeral 206) used when a read request is loaded into a channel. In certain embodiments, arbiter 108 prevents unnecessary blocking of resources by reserving resources at the time of loading read requests into the least-used channel.Thus, FIG. 2 illustrates certain embodiments in which arbiter 108 maintains incoming queue 202 of read requests and also maintains data structures 204 a...204 ncorresponding to the pending reads being processed by each channel 110 a...110 nof solid state drive 102.FIG. 3 illustrates a block diagram illustrating allocation of read requests in an example solid state drive 300 prior to beginning prioritization of the most lightly populated channel and reordering of host commands, in accordance with certain embodiments. The most lightly populated channel has the least number of read requests that are processed by the channel compared to other channels.The example solid state drive 300 has three channels: channel A 302, channel B 304, and channel C 306. Channel A 302 has outstanding read operations 308 denoted by reference numerals 310, 312, 314, i.e., there are three read requests (denoted "read operation" 310, 312, 314) for data stored in NAND chips coupled to channel A 302. Channel B 304 has outstanding reads 316, indicated by reference numeral 318, and channel C 306 has outstanding reads 320, indicated by reference numerals 322, 324.The incoming queue of read requests 326 includes ten read commands 328, 330, 332, 334, 336, 338, 340, 342, 344, 346, where the command at the first location of the incoming queue 326 is the "read A" command 328, and the command at the last location of the incoming queue 326 is the "read B" command 346.FIG. 4 illustrates a block diagram illustrating allocation of read requests in a solid state drive 300 after prioritizing the most lightly populated channel and rearranging host commands, in accordance with certain embodiments.In certain embodiments, arbiter 108 examines the incoming queue of read requests 326 (as shown in FIG. 3 ) and the pending reads processed by the channels as shown in data structures 308, 316, 320. Arbiter 108 then loads channel B 304 (which has only one outstanding read request 318 in FIG. 3 ) on the at least busy channel B 304 (which is read B commands) with commands 340, 344 (which are read B commands) selected out of order from the incoming queue of read requests 326 (as shown in FIG. 3 ).FIG. 4 illustrates the situation after loading the least-busy channel B 304 with commands 340, 344. In FIG. 4, reference numerals 402 and 404 represent commands 340, 344 of FIG. 3 that have now been loaded into channel B 304 for processing in pending reads 316 being processed for channel B 304.Thus, channels 302, 304, and 306 are more evenly loaded by loading the weakest loaded one of three channels 302, 304, 306 with corresponding read requests selected out of order from the queue of read requests 326. It should be noted that none of the commands 328, 330, 332, 334, 336, 338 that were in the incoming queue 326 prior to the command 340 may be loaded into the channel B 304, as the commands 328, 330, 332, 334, 336, 338 are read requests for data accessed via the channel A 302 or the channel C 306. It should also be noted that there is only one arbiter 108 and a plurality of channels, so that the arbiter 108 examines the outstanding reads 308, 316, 320 on channels 302, 304, 306 one at a time. Of course, channels 302, 304, 306 may inform arbiter 108 when channels 302, 304, 306 complete processing of certain read requests, and arbiter 108 may keep track of the outstanding read requests on channels 302, 304, 306 from such information provided by channels 302, 304, 306.In addition, arbiter 108, when implemented using a microcontroller, is a serialized processor. A NAND chip (e.g., NAND chip 112a) has an inherent property that allows only a read request for it. The channel (e.g., channel 110a) for the NAND chip has a "busy" state until the read request for the NAND chip is completed. It is under the responsibility of arbiter 108 not to schedule a new read operation while a channel is busy. Once the channel is no longer busy, arbiter 108 must issue the next command to the NAND chip. To improve channel utilization, in certain embodiments arbiter 108 "interrogates" the "lightly-utilized" channel (i.e., channels used to process relatively fewer read requests) more often than the "heavily-utilized" channels (i.e., channels used to process relatively more read requests), so that reordered read commands are issued to the lightly-utilized channels as soon as possible. This is important because the time to execute a new read command is on the order of 100 microseconds, while arbiter 108 takes approximately the same time to scan all 18 channels and reorder the read commands.FIG. 5 illustrates a first flow diagram 500 for preventing non-uniform channel utilization in solid state drives, in accordance with certain embodiments. The operations illustrated in FIG. 5 may be performed by arbiter 108, which performs operations within solid state drive 102.Control starts at block 502 where the arbiter 108 starts the read processing load (i.e., the bandwidth used) on the first channel 110 aof a plurality of channels 110 a, 110 b,... 110n. Control passes to block 504, in which the scheduler 108 determines whether the read processing load on the channel 110n has been determined. If not ("no" branch 505), arbiter 108 determines the read processing load on the next channel and control returns to block 504. The read processing load may be determined by examining a number of pending read requests in the pending read data structure 204 a...204 n, or via other mechanisms.If a determination is made at block 504 that the read processing load on channel 110n has been determined ("yes" branch 507), control passes to block 508 where it is determined which of the plurality of channels has the lowest processing load and the channel with the lowest processing load is referred to as channel X.From block 508, control passes to block 509, where a determination is made as to whether channel X is busy or not busy, where a channel that is busy cannot handle additional read requests and a channel that is not busy can handle additional read requests. The determination of whether channel X is busy or idle is necessary because a NAND chip coupled to channel X has an inherent property that it allows only a read request for it. Channel X for the NAND chip has a state of "busy" until the read request for the NAND chip is completed.If it is determined at block 509 that channel X is not busy (reference numeral 509a), then control passes to block 510, where arbiter 108 selects one or more read requests destined for channel X that have accumulated in "incoming queue of read requests" 202 such that the available bandwidth of channel X is used as fully as possible, which selection may result in reordering pending requests in "incoming queue of read requests" 202. Arbiter 108 allocates resources for the one or more selected read requests and sends (at block 512) the one or more read requests to channel X for processing.If it is determined at block 509 that channel X is busy (reference numeral 509b), then the process waits until channel X is no longer busy.In alternative embodiments, instead of determining the channel having the lowest processing load, a relatively lightly loaded channel (i.e., a channel having a relatively low processing load of the plurality of channels) may be determined. In certain embodiments, the read requests may be transmitted primarily to the relatively lightly loaded channel. It should be noted that the lightly loaded channel arbiter 108 does not schedule another read request until the lightly loaded channel is acknowledged as "unoccupied.".It may be noted that host read requests continue to accumulate (at block 514) in the "inbound queue of read requests" data structure 202 during execution of operations 502, 504, 505, 506, 507, 508, 510, 512.Thus, FIG. 5 illustrates certain embodiments for selecting the weakest-out-of-load channel and rearranging queue elements in the incoming queue of read requests to select corresponding read requests to load into the weakest-out-of-load channel.FIG. 6 illustrates a second flow diagram 600 for preventing uneven channel utilization in solid state drives, in accordance with certain embodiments. The operations illustrated in FIG. 6 may be performed by arbiter 108, which performs operations within solid state drive 102.Control starts at block 602, in which a solid state drive 102 receives a plurality of read requests from a host 104 via a PCIe bus 106, where a plurality of channels 110 a...110 nin the solid state drive have identical bandwidths. Although the channels 110a... 110 nmay have identical bandwidths, in actual scenarios, one or more of the channels 110 a...110 nmay not fully utilize the bandwidth.An arbiter 108 in the solid state drive 102 determines (at block 604) which of a plurality of channels 110 a...110 nin the solid state drive 102 is a lightly loaded channel (in certain embodiments, the lightly loaded channel is the most lightly loaded channel). Resources for processing one or more read requests destined for the particular lightly loaded channel are allocated (at block 606), the one or more read requests having been received from the host 104.Control passes to block 608 where the one or more read requests are placed in the particular lightly loaded channel for processing. After placing the one or more read requests in the particular lightly loaded channel for processing, the particular lightly loaded channel is used as fully as possible during processing.Thus, FIGS. 1-6 illustrate certain embodiments for preventing uneven utilization of channels in a solid state drive by selecting out-of-order read requests from an incoming queue and loading the out-of-order selected read requests into the channel that is relatively lightly or the weakest utilized.The described operations may be implemented as a method, apparatus, or computer program product using standard programming and / or development techniques for generating software, firmware, hardware, or any combination thereof. The described operations may be implemented as code stored in a "computer readable storage medium," where a processor may read and execute the code from the computer readable storage medium. The computer readable storage medium includes at least one of electronic circuitry, storage materials, inorganic materials, organic materials, biological materials, a protective shell, a housing, a coating, and hardware. A computer readable storage medium may include, but is not limited to, magnetic storage media (e.g., hard disk drives, floppy disks, tape, etc.), optical storage (CD-ROMs, DVDs, optical disks, etc.), volatile and nonvolatile memory devices (e.g., EEPROMs, ROMs, PROMs, RAMs, DRAMs, SRAMs, flash memory, firmware, programmable logic, etc.), solid state devices (SSDs), etc. The code implementing the described operations may be further implemented in hardware logic implemented in a hardware device (e.g., an integrated circuit chip, a programmable gate array (PGA), an application specific integrated circuit (ASIC), etc.). Further, the code implementing the described operations may be implemented in "transmit signals", where transmit signals may propagate through space or through a transmission medium such as fiber optics, copper wire, etc. The transmission signals in which the code or logic is encoded may further include a wireless signal, satellite transmission, radio waves, infrared signals, Bluetooth, etc. The program code embedded on a computer readable storage medium may be transmitted as transmission signals from a transmitting station or a computer to a receiving station or a computer. A computer readable storage medium includes not only transmission signals. It will be appreciated by those skilled in the art that many modifications may be made to this configuration and that the article of manufacture may comprise a suitable information carrier medium known in the art.Computer program code for carrying out operations for aspects of certain embodiments may be written in any combination of one or more programming languages. Blocks of the flowchart and block diagrams may be implemented by computer program instructions.FIG. 7 illustrates a block diagram of a system 700 that includes both the host 104 (the host 104 includes at least one processor) and the solid state drive 102, in accordance with certain embodiments. For example, in certain embodiments, the system 700 may be a computer (e.g., a laptop computer, a desktop computer, a tablet, a cellular telephone, or any other suitable computing device) that includes the host 104 and the solid state drive 102 in the system 700. For example, in certain embodiments, system 700 may be a laptop computer that includes hard disk drive 102.The system 700 may include circuitry 702, which in certain embodiments may include at least one processor 704. The system 700 may also include memory 706 (e.g., a volatile storage device) and storage 708. The data store 708 may include the solid state drive 102 or other drives or devices including a non-volatile memory device (e.g., EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, firmware, programmable logic, etc.). The data storage 708 may also include a magnetic disk drive, an optical disk drive, a tape drive, etc. The data store 708 may include an internal data storage device, an attached data storage device, and / or a network access data storage device. The system 700 may include program logic 710 including code 712 that may be loaded into memory 706 and executed by the processor 704 or circuitry 702. In certain embodiments, program logic 710 including code 712 may be stored in data store 708. In certain embodiments, program logic 710 may be implemented in circuitry 702. Thus, although FIG. 7 illustrates the program logic 710 separate from the other elements, the program logic 710 may be implemented in the memory 706 and / or circuitry 702. The system 700 may also include a display 714 (e.g., a liquid crystal display (LCD), a light emitting diode (LED) display, a cathode ray tube (CRT) display, a touch screen display, and any other suitable display). The system 700 may also include one or more input devices 716, such as a keyboard, a mouse, a joystick, a trackpad, or any other suitable input device. Other components or devices beyond those shown in FIG. 7 may also be found in system 700.Certain embodiments may relate to a method of providing computer instructions by a person or automatic processing, wherein computer readable code is incorporated into a computer system, the code, together with the computer system, being capable of performing the operations of the described embodiments.The terms "(any) embodiment," "embodiment," "embodiments," "the embodiment," "the embodiments," "one or more embodiments," "some embodiments," and "one (particular) embodiment" refer to "one or more (but not all) embodiments," unless expressly stated otherwise.The terms "comprising", "having", and variants thereof mean "comprising, without being limited thereto" unless expressly stated otherwise.The enumeration of elements does not mean that each or all of the elements exclude or exclude each other unless expressly stated otherwise.The terms "a", "an" and "the" mean "one or more" unless expressly stated otherwise.Devices in communication with each other need not be in continuous communication with each other unless expressly stated otherwise. In addition, devices in communication with each other may communicate directly with each other or indirectly with each other through one or more agents.A description of the embodiments having multiple components in communication with each other does not mean that all of these components are required. On the contrary, a variety of optional components are described to illustrate the wide variety of possible embodiments.Further, although process steps, method steps, algorithms, or the like may be described in an order, such processes, methods, and algorithms may be configured to operate in alternative orders. In other words, a sequence or order of steps that may be described does not necessarily indicate a need for the steps to be executed in that order. The steps of processes described herein may be performed in any order that is feasible. Further, some steps may be performed simultaneously.It will be appreciated that when a single device or article is described herein, more than one device / article (whether co-operating or not) may be used instead of a single device / article. Similarly, it will be appreciated that when more than one device or item is described herein (whether or not they are cooperating), a single device or item may be used in place of the plurality of devices or items, or a different number of devices / items may be used in place of the specified number of devices or programs. The functionality and / or features of a device may alternatively be realized by one or more devices not expressly described as having such functionality or features. Thus, other embodiments need not include the device itself.At least certain operations that may be illustrated in the figures represent certain events that occur in a particular order. In alternative embodiments, certain operations may be performed, modified, or removed in a different order. In addition, steps may be added to the logic described above and still correspond to the described embodiments. Further, operations described herein may occur in sequence, or certain operations may be processed in parallel. Moreover, operations may be performed by a single processing unit or by distributed processing units.The foregoing description of various embodiments has been presented for purposes of illustration and description. They are not intended to be exhaustive or to be limited to the specific forms disclosed. Many modifications and changes are possible in view of the above teaching.ExamplesThe following examples are directed to further embodiments.Example 1 is a method in which an arbiter in a solid state drive determines which of a plurality of channels in the solid state drive is a lightly loaded channel compared to other channels. Resources are allocated for processing one or more read requests destined for the particular lightly loaded channel, the one or more read requests having been received from a host. The one or more read requests are placed in the particular lightly loaded channel for processing.In Example 2, the subject matter of claim 1 may include that the particular lightly loaded channel is a weakest loaded channel of the plurality of channels, wherein after placing the one or more read requests in the particular weakest loaded channel for processing, the particular weakest loaded channel is used as fully as possible during processing.In Example 3, the subject matter of claim 1 can include that the one or more read requests are included in a plurality of read requests destined for the plurality of channels, wherein an order of processing the plurality of read requests is modified for processing by placing the one or more read requests in the particular lightly loaded channel.In Example 4, the subject matter of claim 3 can include that modifying the order of processing the plurality of requests processes the one or more read requests destined for the particular lightly loaded channel with priority over other requests.In Example 5, the subject matter of claim 1 can include that the solid state drive receives the one or more read requests from the host via a Peripheral Component Interconnect Express (PCIe) bus, wherein each of the plurality of channels in the solid state drive has an identical bandwidth.In Example 6, the subject matter of claim 5 may include that a sum of bandwidths of the plurality of channels corresponds to a bandwidth of the PCIe bus.In Example 7, the subject matter of claim 1 can include that at least one of the plurality of channels is coupled to a different number of NAND chips than other channels of the plurality of channels.In Example 8, the subject matter of claim 1 can include that if the one or more read requests are not placed in the particular lightly loaded channel for processing, then read performance on the solid state drive decreases by over 10% over another solid state drive in which all channels are coupled to an equal number of NAND chips.In Example 9, the subject matter of claim 1 can include that the allocating the resources for processing is after the arbiter in the solid state drive determines which of the plurality of channels in the solid state drive is the lightly loaded channel.In Example 10, the subject matter of claim 1 may include the arbiter polling relatively lightly loaded channels more often than relatively heavily loaded channels to issue reordered read requests primarily to the relatively lightly loaded channels.In Example 11, the subject matter of claim 1 can include associating a data structure that holds pending reads processed by the channel with each of the plurality of channels and holding the one or more read requests received from the host in an incoming queue of read requests received from the host.Example 12 is an apparatus comprising a plurality of non-volatile memory chips, a plurality of channels coupled to the plurality of non-volatile memory chips, and an arbiter for controlling the plurality of channels, the arbiter configured to: determine which of the plurality of channels is a lightly loaded channel compared to other channels; allocate resources to process one or more read requests determined for the determined lightly loaded channel, wherein the one or more read requests were received from a host; and arrange the one or more read requests in the determined lightly loaded channel for processing.In Example 13, the subject matter of claim 12 may include that the nonvolatile memory chips include NAND chips, wherein the particular lightly loaded channel is a weakest loaded channel of the plurality of channels, wherein after placing the one or more read requests in the particular weakest loaded channel, to process the particular weakest loaded channel is used as fully as possible during processing.In Example 14, the subject matter of claim 12 may include that the one or more read requests are included in a plurality of read requests destined for the plurality of channels, wherein an order of processing the plurality of read requests is modified for processing by placing the one or more read requests in the particular lightly loaded channel.In Example 15, the subject matter of claim 14 can include that modifying the order of processing the plurality of requests processes the one or more read requests destined for the particular lightly loaded channel with priority over other requests.In Example 16, the subject matter of claim 12 can include the device receiving the one or more read requests from the host via a Peripheral Component Interconnect Express (PCIe) bus, wherein each of the plurality of channels in the device has an identical bandwidth.In Example 17, the subject matter of claim 16 may include that a sum of bandwidths of the plurality of channels corresponds to a bandwidth of the PCIe bus.In Example 18, the subject matter of claim 12 can include that the nonvolatile memory chips include NAND chips, wherein at least one of the plurality of channels is coupled to a different number of NAND chips than other channels of the plurality of channels.In Example 19, the subject matter of claim 12 may include that the nonvolatile memory chips include NAND chips, wherein if the one or more read requests are not placed in the particular lightly loaded channel for processing, the read power on the device decreases by over 10% over another device in which all channels are coupled to an equal number of NAND chips.In Example 20, the subject matter of claim 12 can include that the allocating the resources for processing is after the arbiter determines in the device which of the plurality of channels in the device is the lightly loaded channel.In Example 21, the subject matter of claim 12 may include the arbiter polling relatively lightly loaded channels more often than relatively heavily loaded channels to issue reordered read requests primarily to the relatively lightly loaded channels.In Example 22, the subject matter of claim 12 may include associating a data structure that holds pending reads processed by the channel with each of the plurality of channels and holding the one or more read requests received from the host in an incoming queue of read requests received from the host.Example 23 is a system comprising a solid state drive, a display, and a processor coupled to the solid state drive and the display, wherein the processor sends a plurality of read requests to the solid state drive, and wherein the solid state drive performs operations in response to the plurality of read requests, the operations comprising: determining which of a plurality of channels in the solid state drive is a lightly loaded channel compared to other channels in the solid state drive; allocating resources to process one or more read requests selected from the plurality of read requests, wherein the one or more read requests are determined for the determined lightly loaded channel; and arranging the one or more read requests in the determined lightly loaded channel for processing.In Example 24, the subject matter of claim 23 further includes wherein the solid state drive further includes a plurality of non-volatile memory chips including NAND or NOR chips, wherein the lightly loaded channel is a weakest loaded channel of the plurality of channels, and wherein after placing the one or more read requests in the particular weakest loaded channel, to process the particular weakest loaded channel is used as fully as possible during processing.In Example 25, the subject matter of claim 23 further includes modifying an order of processing the plurality of requests for processing by placing the one or more read requests in the particular lightly loaded channel.
Claims
A method comprising: determining (604), by an arbiter in a solid state drive, which of a plurality of channels in the solid state drive is a lightly loaded channel as compared to other channels; allocating (606) resources to process one or more read requests destined for the particular lightly loaded channel, wherein the one or more read requests have been received from a host; and arranging (608) the one or more read requests in the particular lightly loaded channel for processing; wherein the one or more read requests are included in a plurality of read requests destined for the plurality of channels, and wherein the one or more read requests destined for the particular lightly loaded channel are processed for processing with priority over other requests by placing the one or more read requests in the particular lightly loaded channel.The method of claim 1, wherein the particular lightly loaded channel is a weakest loaded channel of the plurality of channels, and wherein after placing the one or more read requests in the particular weakest loaded channel for processing, the particular weakest loaded channel is used as fully as possible during processing.The method of claim 1, further comprising: receiving, by the solid state drive, the one or more read requests from the host via a Peripheral Component Interconnect Express bus, wherein each of the plurality of channels in the solid state drive has an identical bandwidth.The method of claim 3, wherein a sum of bandwidths of the plurality of channels corresponds to a bandwidth of the PCIe bus.The method of claim 1, wherein at least one of the plurality of channels is coupled to a different number of NAND chips than other channels of the plurality of channels.The method of claim 1, wherein if the one or more read requests are not placed in the particular lightly loaded channel for processing, then read power on the solid state drive decreases by over 10% over another solid state drive in which all channels are coupled to an equal number of NAND chips.The method of claim 1, wherein allocating the resources for processing is after the arbiter in the solid state drive determines which of the plurality of channels in the solid state drive is the lightly loaded channel.The method of claim 1, wherein the arbiter interrogates relatively lightly loaded channels more often than relatively heavily loaded channels to issue reordered read requests primarily to the relatively lightly loaded channels.The method of claim 1, further comprising: associating a data structure that holds pending reads processed by the channel with each of the plurality of channels; and holding the one or more read requests received from the host in an incoming queue of read requests received from the host.An apparatus comprising: a plurality of non-volatile memory chips (112, 114); a plurality of channels (110) coupled to the plurality of non-volatile memory chips; and an arbiter (108) for controlling the plurality of channels, the arbiter configured to: determine (604) which of the plurality of channels is a lightly loaded channel as compared to other channels; allocate (606) resources to process one or more read requests destined for the particular lightly loaded channel, wherein the one or more read requests were received from a host; and arrange (608) the one or more read requests in the particular lightly loaded channel for processing; wherein the one or more read requests are included in a plurality of read requests destined for the plurality of channels, and wherein the one or more read requests destined for the particular lightly loaded channel are processed for processing with priority over other requests by placing the one or more read requests in the particular lightly loaded channel.The apparatus of claim 10, wherein the non-volatile memory chips comprise NAND chips, wherein the lightly loaded channel is a weakest loaded channel of the plurality of channels, and wherein after placing the one or more read requests in the particular weakest loaded channel for processing, the particular weakest loaded channel is used as fully as possible during processing.The apparatus of claim 10, wherein the plurality of read requests is received from the host.The apparatus of claim 10, wherein the apparatus receives the one or more requests from the host via a Peripheral Component Interconnect Express bus, each of the plurality of channels having an identical bandwidth.The apparatus of claim 13, wherein a sum of bandwidths of the plurality of channels corresponds to a bandwidth of the PCIe bus.The apparatus of claim 10, wherein the non-volatile memory chips comprise NAND chips, and wherein at least one of the plurality of channels is coupled to a different number of NAND chips than other channels of the plurality of channels.The apparatus of claim 10, wherein the non-volatile memory chips comprise NAND chips, and wherein if the one or more read requests are not placed in the particular lightly loaded channel for processing, read performance decreases by over 10% over another device in which all channels are coupled to an equal number of NAND chips.The apparatus of claim 10, wherein the allocating the resources for processing is after the arbiter determines which of the plurality of channels is the lightly loaded channel.The apparatus of claim 10, wherein the arbiter interrogates relatively lightly loaded channels more often than relatively heavily loaded channels to issue reordered read requests primarily to the relatively lightly loaded channels.The apparatus of claim 10, wherein the arbiter is further configured to: associate a data structure that holds pending reads processed by the channel with each of the plurality of channels; and hold the one or more read requests received from the host in an incoming queue of read requests received from the host.A system, comprising: a solid state drive; a display (714); and a processor (704) coupled to the hard disk drive and the display, wherein the processor sends a plurality of read requests to the solid state drive, and wherein the solid state drive performs operations in response to the plurality of read requests, the operations comprising: determining (604) which of a plurality of channels in the solid state drive is a lightly loaded channel compared to other channels in the solid state drive; allocating (606) resources to process one or more read requests selected from the plurality of read requests, wherein the one or more read requests are destined for the particular lightly loaded channel; and arranging (608) the one or more read requests in the particular lightly loaded channel for processing; wherein the one or more read requests are included in a plurality of read requests destined for the plurality of channels, and wherein the one or more read requests destined for the particular lightly loaded channel are processed for processing with priority over other requests by placing the one or more read requests in the particular lightly loaded channel.The system of claim 20, wherein the solid state drive further comprises a plurality of non-volatile memory chips comprising NAND or NOR chips, wherein the lightly loaded channel is a weakest loaded channel of the plurality of channels, and wherein after placing the one or more read requests in the particular weakest loaded channel for processing, the particular weakest loaded channel is used as fully as possible during processing.
Citation Information
Patent Citations
Memory System with Command Queue Reordering
US20130262745A1
Systems for sustained read and write performance with non-volatile memory
US8341300B1