Balanced transfer over interface

By reordering transactions and data groups in the NVMe SSD system, the problem of interface imbalance was solved, interface saturation and traffic balance were achieved, performance and reliability were optimized, and system scalability was supported.

CN121785518APending Publication Date: 2026-04-03SANDISK TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

In existing technologies, cross-interface traffic balancing is unbalanced in non-volatile memory (NVMe) solid-state drive (SSD) systems, leading to idle or inefficient interface operations, which affects performance and reliability.

Method used

By reordering transactions or data groups, based on factors such as namespace identifier (ID), zone ID in the partition namespace (ZNS) driver, commit and completion ID, physical and virtual functions, interface saturation and traffic balance are achieved, and control and scheduling are performed using the host interface module (HIM) and flash interface module (FIM).

Benefits of technology

Performance was optimized, system reliability was maintained, and scalability was supported, ensuring full saturation of each interface and avoiding bottlenecks and uneven data distribution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121785518A_ABST
    Figure CN121785518A_ABST
Patent Text Reader

Abstract

Randomly sending transactions across interfaces may result in idle interfaces and generally inefficient operations. Reordering transactions or data packets involves potentially changing the order of packet transmissions to a host device. The reordering may ensure that each interface between the host device and the data storage device is saturated. The saturation is achieved by the reordering, such that packet transmission across the interfaces is balanced. The balancing may be based on any number of factors, such as namespace identification (ID), zone ID in a Zoned namespace (ZNS) driver, commit and completion ID, physical and virtual functions, and host address, just by several examples. This reordering and thus balancing may facilitate packet-level fairness and enable integration of asymmetric systems. This may optimize performance, maintain reliability, and support scalability.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Background of this disclosure

[0002] This disclosure of the field

[0003] The implementation scheme disclosed herein generally relates to improving cross-interface traffic balancing.

[0004] Description of related technologies

[0005] Non-volatile memory (NVM) fast (NVMe) solid-state drives (SSDs) connect to host devices via a peripheral component interconnect (PCIe) interface. This interface is used to satisfy the NVMe protocol while attempting to achieve maximum performance. To serve host commands, NVMe requires the use of interfaces for different tasks: reading commands, reading pointers, reading data, and, on some products, reading mapping tables.

[0006] In the context of storage devices using PCIe and NVMe interfaces, traffic balancing optimizes performance, improves efficiency, and ensures fair utilization of resources. Some key reasons why traffic balancing is important include optimizing throughput, avoiding bottlenecks, effectively utilizing multiple channels, enhancing scalability, improving reliability, reducing latency, maximizing NVMe parallelism, supporting dynamic workloads, and meeting application requirements.

[0007] In summary, traffic balancing in PCIe and NVMe-based storage systems optimizes performance, maintains reliability, and supports scalability. Traffic balancing ensures that the full capacity of high-speed interfaces is effectively utilized, thus contributing to overall responsiveness and a reliable storage infrastructure. However, achieving traffic balancing is always a challenge.

[0008] Therefore, there is a need in this field for improved traffic balancing across interfaces.

[0009] Overview of this disclosure

[0010] Sending transactions randomly across interfaces can lead to idle interfaces and generally inefficient operations. Reordering transactions or data packets involves potentially altering the order in which packets are transmitted to the host device. This reordering ensures that each interface between the host device and the data storage device is saturated. Achieving this saturation through reordering balances packet transmission across these interfaces. This balance can be based on any number of factors, such as namespace identifiers (IDs), zone IDs in partition namespace (ZNS) drives, commit and completion IDs, physical and virtual functions, and host addresses, to name a few. This reordering and thus balancing contributes to packet-level fairness and enables the integration of asymmetric systems. Doing so optimizes performance, maintains reliability, and supports scalability.

[0011] In one embodiment, a data storage device includes: a memory device; and a controller coupled to the memory device, wherein the controller is configured to: classify transactions to be sent to one or more host devices, wherein the transactions are at the group level; reorder the transactions based on the classification; and transmit the reordered transactions to one or more host devices.

[0012] In another embodiment, a data storage device includes: a memory device; and a controller coupled to the memory device, wherein the controller includes: a host interface module (HIM) including a reordering buffer and a completion and interrupt synchronization module, wherein the HIM is configured to maintain a first interface between the controller and a first host device, and wherein the HIM is configured to maintain a second interface between the controller and a second host device; a flash interface module (FIM) coupled to the memory device; and a command scheduler coupled between the HIM and the FIM, wherein the controller is configured to balance traffic between the first host device and the second host device, wherein balancing includes ensuring saturation of the first and second interfaces.

[0013] In another embodiment, a data storage device includes: a component for storing data; and a controller coupled to the component for storing data, wherein the controller is configured to: receive transaction packets from the component for storing data; classify the transaction packets; place the transaction packets in queues of a plurality of queues; send the transaction packets to a host device; and maintain saturation on the interface between the controller and the host device and on the interface between the controller and other host devices. Attached Figure Description

[0014] To gain a more detailed understanding of the features described above, the present disclosure can be described in more detail by referring to embodiments, some of which are illustrated in the accompanying drawings. However, it should be noted that the drawings illustrate only typical embodiments of the present disclosure and should not be construed as limiting the scope of the disclosure, as other equivalent embodiments are permissible.

[0015] Figure 1 This is a schematic block diagram illustrating a storage system in which a data storage device can be used as a storage device for a host device, according to certain embodiments.

[0016] Figure 2 This is a schematic diagram of a non-volatile memory (NVM) fast (NVMe) solid-state drive (SSD) system according to one implementation scheme.

[0017] Figure 3This is a schematic diagram of a multi-tenant system based on an implementation plan.

[0018] Figure 4 This is a schematic diagram of a data storage system based on an implementation plan.

[0019] Figure 5 It is a schematic diagram of the reordering logic based on an implementation plan.

[0020] Figure 6 This is a schematic diagram of two host systems based on an implementation plan.

[0021] Figure 7 This is a flowchart illustrating the transaction balancing process according to one implementation scheme.

[0022] For ease of understanding, the same reference numerals are used where possible to denote common elements in the figures. It is contemplated that elements disclosed in one embodiment may be advantageously used in other embodiments without being specifically listed. Detailed Implementation

[0023] In the following text, reference is made to embodiments of this disclosure. However, it should be understood that this disclosure is not limited to the specifically described embodiments. Rather, any combination of the following features and elements (whether or not different embodiments are involved) is contemplated to realize and practice this disclosure. Furthermore, while embodiments of this disclosure may achieve advantages over other possible solutions and / or over the prior art, whether a particular advantage is achieved by a given embodiment does not limit this disclosure. Therefore, the following aspects, features, embodiments, and advantages are merely illustrative and should not be considered as elements or limitations of the appended claims unless expressly stated in the claims. Similarly, reference to “this disclosure” should not be construed as a generalization of any inventive subject matter disclosed herein and should not be considered as elements or limitations of the appended claims unless expressly stated in the claims.

[0024] Sending transactions randomly across interfaces can lead to idle interfaces and generally inefficient operations. Reordering transactions or data packets involves potentially altering the order in which packets are transmitted to the host device. This reordering ensures that each interface between the host device and the data storage device is saturated. Achieving this saturation through reordering balances packet transmission across these interfaces. This balance can be based on any number of factors, such as namespace identifiers (IDs), zone IDs in partition namespace (ZNS) drives, commit and completion IDs, physical and virtual functions, and host addresses, to name a few. This reordering and thus balancing contributes to packet-level fairness and enables the integration of asymmetric systems. Doing so optimizes performance, maintains reliability, and supports scalability.

[0025] Figure 1 This is a schematic block diagram illustrating a storage system 100 having a data storage device 106 that can be used as a storage device for a host device 104, according to some embodiments. For example, the host device 104 may utilize non-volatile memory (NVM) 110 included in the data storage device 106 to store and retrieve data. The host device 104 includes host dynamic random access memory (DRAM) 138. In some examples, the storage system 100 may include multiple storage devices, such as the data storage device 106, which may operate as a storage array. For example, the storage system 100 may include multiple data storage devices 106 configured as a redundant array of inexpensive / independent disks (RAID) that collectively serve as a large-capacity storage device for the host device 104.

[0026] Host device 104 can store data to or retrieve data from one or more storage devices (such as data storage device 106). Figure 1 As shown, host device 104 can communicate with data storage device 106 via interface 114. Host device 104 can include any device from a wide range of devices, including computer servers, network attached storage (NAS) units, desktop computers, notebook computers (i.e., laptops), tablet computers, set-top boxes, handsets (such as so-called "smart" phones, so-called "smart" boards), televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, or other devices capable of sending or receiving data from data storage devices.

[0027] Host DRAM 138 may optionally include a main memory buffer (HMB) 150. HMB 150 is part of host DRAM 138 allocated to data storage device 106 for use by the controller 108 of data storage device 106. For example, controller 108 may store mapped data, buffer commands, logical-to-physical (L2P) tables, metadata, etc., in HMB 150. In other words, HMB 150 can be used by controller 108 to store data that would typically be stored in volatile memory 112, buffer 116, or the controller 108's internal memory (such as static random access memory (SRAM)). In an example where data storage device 106 does not include DRAM (i.e., optional DRAM 118), controller 108 may utilize HMB 150 as DRAM for data storage device 106.

[0028] Data storage device 106 includes a controller 108, an NVM 110, a power supply 111, volatile memory 112, an interface 114, a write buffer 116, and optional DRAM 118. In some examples, data storage device 106 may include components not shown for clarity. Figure 1 Additional components are shown in the diagram. For example, data storage device 106 may include a printed circuit board (PCB) to which components of data storage device 106 are mechanically attached, and the PCB includes conductive traces for electrically interconnecting components of data storage device 106, etc. In some examples, the physical dimensions and connector configuration of data storage device 106 may conform to one or more standard form factors. Some example standard form factors include, but are not limited to, 3.5″ data storage devices (e.g., HDDs or SSDs), 2.5″ data storage devices, 1.8″ data storage devices, peripheral component interconnect (PCI), PCI expansion (PCI-X), PCI fast (PCIe) (e.g., PCIe x1, x4, x8, x16, PCIe mini-cards, mini PCI, etc.). In some examples, data storage device 106 may be directly coupled (e.g., directly soldered or inserted into a connector) to the motherboard of host device 104.

[0029] Interface 114 may include one or both of a data bus for exchanging data with host device 104 and a control bus for exchanging commands with host device 104. Interface 114 may operate according to any suitable protocol. For example, interface 114 may operate according to one or more of the following protocols: Advanced Technology Attachment (ATA) (e.g., Serial ATA (SATA) and Parallel ATA (PATA)), Fibre Channel Protocol (FCP), Small Computer System Interface (SCSI), Serial Attached SCSI (SAS), PCI and PCIe, Non-Volatile Memory Express (NVMe), OpenCAPI, GenZ, Cache Coherent Interface Accelerator (CCIX), Open Channel SSD (OCSSD), etc. Interface 114 (e.g., a data bus, a control bus, or both) is electrically connected to controller 108, providing an electrical connection between host device 104 and controller 108, thereby allowing data exchange between host device 104 and controller 108. In some examples, the electrical connection of interface 114 may also allow data storage device 106 to receive power from host device 104. For example, as Figure 1 As shown, power supply 111 can receive power from host device 104 via interface 114.

[0030] NVM 110 may include multiple memory devices or memory cells. NVM 110 may be configured to store and / or retrieve data. For example, a memory cell of NVM 110 may receive data from controller 108 and messages instructing the memory cell to store data. Similarly, a memory cell may receive messages from controller 108 instructing the memory cell to retrieve data. In some examples, each memory cell in the memory cell may be referred to as a die. In some examples, NVM 110 may include multiple dies (i.e., multiple memory cells). In some examples, each memory cell may be configured to store a relatively large amount of data (e.g., 128MB, 256MB, 512MB, 1GB, 2GB, 4GB, 8GB, 16GB, 32GB, 64GB, 128GB, 256GB, 512GB, 1TB, etc.).

[0031] In some examples, each memory cell may include any type of non-volatile memory device, such as flash memory device, phase-change memory (PCM) device, resistive random access memory (ReRAM) device, magnetoresistive random access memory (MRAM) device, ferroelectric random access memory (F-RAM), holographic memory device, and any other type of non-volatile memory device.

[0032] NVM 110 may include multiple flash memory devices or memory cells. The NVM flash memory devices may include NAND- or NOR-based flash memory devices and may store data based on the charge in the floating gate of the transistors contained in each flash memory cell. In an NVM flash memory device, the flash memory device may be divided into multiple dies, each of which includes multiple physical or logical blocks, which may be further divided into multiple pages. Each of the multiple blocks within a particular memory device may include multiple NVM cells. Rows of NVM cells may be electrically connected using word lines to define pages within the multiple pages. A corresponding cell in each page of the multiple pages may be electrically connected to a corresponding bit line. Furthermore, the NVM flash memory device may be a 2D or 3D device and may be a single-level cell (SLC), multi-level cell (MLC), three-level cell (TLC), or four-level cell (QLC). Controller 108 may write data to and read data from the NVM flash memory device at the page level and erase data from the NVM flash memory device at the block level.

[0033] Power supply 111 can provide power to one or more components of data storage device 106. When operating in standard mode, power supply 111 can use power provided by an external device (such as host device 104) to power one or more components. For example, power supply 111 can use power received from host device 104 via interface 114 to power one or more components. In some examples, power supply 111 may include one or more power storage components configured to provide power to one or more components when operating in a shutdown mode, such as when power is no longer received from external devices. In this way, power supply 111 can be used as an onboard backup power source. Some examples of one or more power storage components include, but are not limited to, capacitors, supercapacitors, batteries, etc. In some examples, the amount of electrical energy that can be stored by one or more power storage components may vary with the cost and / or size (e.g., area / volume) of one or more power storage components. In other words, as the amount of electrical energy stored by one or more power storage components increases, the cost and / or size of one or more power storage components also increases.

[0034] Controller 108 may use volatile memory 112 to store information. Volatile memory 112 may include one or more volatile memory devices. In some examples, controller 108 may use volatile memory 112 as a cache. For example, controller 108 may store cached information in volatile memory 112 until the cached information is written to NVM 110. Figure 1 As shown, volatile memory 112 can consume power received from power supply 111. Examples of volatile memory 112 include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static RAM (SRAM), and synchronous dynamic RAM (SDRAM (e.g., DDR1, DDR2, DDR3, DDR3L, LPDDR3, DDR4, LPDDR4, etc.)). Similarly, optional DRAM 118 can be used to store mapped data, buffered commands, logical-to-physical (L2P) tables, metadata, cached data, etc. In some examples, data storage device 106 does not include optional DRAM 118, making data storage device 106 DRAM-free. In other examples, data storage device 106 includes optional DRAM 118.

[0035] Controller 108 can manage one or more operations of data storage device 106. For example, controller 108 can manage reading data from and / or writing data to NVM 110. In some embodiments, when data storage device 106 receives a write command from host device 104, controller 108 can initiate a data storage command to store data into NVM 110 and monitor the progress of the data storage command. Controller 108 can determine at least one operating characteristic of storage system 100 and store at least one operating characteristic in NVM 110. In some embodiments, when data storage device 106 receives a write command from host device 104, controller 108 temporarily stores the data associated with the write command in internal memory or write buffer 116 before sending the data associated with the write command to NVM 110. Controller 108 may include circuitry or a processor configured to execute programs for operating data storage device 106.

[0036] Controller 108 may include optional second volatile memory 120. Optional second volatile memory 120 may be similar to volatile memory 112. For example, optional second volatile memory 120 may be SRAM. Controller 108 may allocate a portion of the optional second volatile memory to host device 104 as a controller memory buffer (CMB) 122. CMB 122 may be directly accessed by host device 104. For example, host device 104 may utilize CMB 122 to store one or more commit queues that are normally maintained in host device 104, rather than maintaining one or more commit queues in host device 104. In other words, host device 104 may generate commands and store the generated commands (with or without associated data) in CMB 122, wherein controller 108 accesses CMB 122 to retrieve the stored generated commands and / or associated data.

[0037] As mentioned above, traffic balancing offers numerous benefits. For example, PCIe and NVMe interfaces are designed to provide high-speed data transfer between data storage devices and host devices for optimizing throughput. Traffic balancing helps distribute data traffic evenly across multiple channels or pathways, thereby maximizing the overall throughput of the data storage system.

[0038] To avoid bottlenecks, uneven distribution of data traffic can lead to bottlenecks in certain channels or pathways. Traffic balancing helps prevent congestion and ensures that no single channel becomes a limiting factor, thus allowing data storage devices to operate at their full potential.

[0039] To effectively utilize multiple channels, PCIe-based data storage devices typically have multiple channels or channels to support parallel data transmission. Traffic balancing ensures efficient use of each channel, preventing some channels from being underutilized while others are overloaded.

[0040] For enhanced scalability, traffic balancing becomes increasingly important as storage systems scale in terms of capacity and performance. Traffic balancing allows storage infrastructure to scale horizontally, concurrently utilizing multiple PCIe / NVMe devices while maintaining optimal performance.

[0041] To improve reliability, flow balancing contributes to storage system reliability by preventing uneven wear on different components. Balanced flow distribution reduces the likelihood that specific components may experience hotspots due to overuse.

[0042] In data storage systems, minimizing latency is beneficial for responsiveness. Traffic balancing ensures that data is distributed evenly, preventing some data paths from experiencing higher latency due to congestion.

[0043] To maximize NVMe parallelism, NVMe is designed to leverage the inherent parallelism in memory, particularly NAND flash memory. Proper traffic balancing enables efficient use of NVMe queues and parallelism, ensuring the concurrent processing of multiple input / output (I / O) operations.

[0044] For systems that support dynamic workloads, storage workloads can change dynamically, and traffic balancing allows the data storage system to adapt to changing conditions. Traffic balancing ensures that data storage devices can handle different levels of read and write requests without causing performance imbalances.

[0045] To meet application needs, different applications and workloads have different storage performance requirements. Traffic balancing helps meet the specific requirements of different applications, thus allowing the storage infrastructure to support a wide range of usage scenarios.

[0046] The disclosures in this paper address the problem of achieving balanced delivery at the packet level, which is beneficial for certain applications. For example, traffic balancing is advantageous in multi-host asymmetric systems when each host device has a different maximum performance.

[0047] Figure 2This is a schematic diagram of an NVMe SSD system 200 according to one implementation scheme. Generally, system 200 includes a host device and a data storage device. The host device includes local DRAM with HMB. The data storage device includes a controller with a PCIe interface. This interface has buffers for data storage. The controller is responsible for command fetching, pointer fetching, data fetching, and table fetching, etc.

[0048] More specifically, the data storage device uses NVMe over a PCIe interface. This interface sits between the host device and the data storage device, and generally, for the NVMe protocol, there are several types of transfers originating from the interface, such as fetch commands or data pointers. There are also data transfers of the user data itself, and L2P tables or other tables potentially stored in the HMB. This disclosure will discuss balancing transfers over the host interface—fetch commands, fetch pointers, fetch data, fetch tables, and other transfers—to maximize performance.

[0049] If the interface is used only for data retrieval, there may be scenarios where there are not enough commands in the data storage device, and therefore the interface with the memory device (e.g., NAND) will be unsaturated. Because the interface is always busy with only data transfers, a balance needs to be struck and bottlenecks need to be avoided. The goal is to figure out how to achieve a better balance between the various types of transfers occurring on top of the interface to maximize performance. Generally, this disclosure is more relevant to multi-host interfaces and focuses on low-level packetization. However, it should be understood that this disclosure relates to single-host interfaces, particularly where a single host includes one or more virtual and / or physical functions.

[0050] This disclosure discusses reordering logic responsible for balancing traffic across interfaces to ensure each interface is fully saturated. This balancing can be based on namespace IDs, zone IDs in ZNS drives, commit and completion IDs, PCIe physical and virtual functions, and / or host addresses. It should be understood that balancing can also be based on other standards. Special logic is used to synchronize completion and interrupt messages toward the host device. Fairness in packet-level processing can be achieved through this disclosure. This balancing enables the integration of asymmetric systems and is beneficial for optimizing performance, maintaining reliability, and supporting scalability. Figure 3 The concept of this system is described.

[0051] Figure 3This is a schematic diagram 300 of a multi-tenant system according to one implementation. A single storage device is connected to multiple host systems via a switch. The storage device is connected to the switch using a single-port interface. The maximum throughput on the interface between the storage device and the switch is greater than the throughput on the interface between a specific host device and the switch. In a balanced system, the throughput on the interface between the storage device and the switch equals the total throughput on all interfaces between the host devices and the switch. The data storage device needs to ensure that each host device's interface is fully saturated.

[0052] exist Figure 3 In the example, there are four host devices. A switch exists between the host devices and the data storage device. For example, there might be a scenario where the interface between the storage device and the switch could be a Generation 5 interface with four channels, and the interface between host A and the switch could be any other interface. This interface could also be a Generation 5 by 4 interface, but for example, it could be a Generation 5 by 1 interface or some other interface. Any speed can be operating on the interfaces. Different interfaces should be considered when maximizing performance. Therefore, the data storage device also needs to ensure sufficient tasks, activity, or transmissions on all interfaces. Ensuring only one interface is saturated is insufficient, as there could be scenarios where one host interface is saturated while all other host devices are idle, which is undesirable. For efficient operation, all interfaces should be saturated, not just one or a few.

[0053] Figure 4 A high-level block diagram of the data storage system that implements the present disclosure is shown. Figure 4 This is a schematic diagram 400 of a data storage system according to an implementation plan. Figure 4 The reordering logic within the controller's Host Interface Module (HIM) is illustrated. This reordering logic is responsible for classifying transactions and sending them to host devices in different orders, thus balancing traffic across all host device interfaces. The reordering logic uses a reordering buffer to perform balancing by reordering transmissions, achieving interface-wide balance and maximizing performance. Additionally, the completion and interruption synchronization logic handles internal completion and interruption messages, storing them internally and sending them to the host devices only after the data transmission associated with those messages is complete.

[0054] More specifically, relative to Figure 4The system comprises multiple host devices on one side and a data storage device on the other side, having a memory device (e.g., NAND or DRAM) and a device controller responsible for interaction between the host devices and the memory device. The HIM includes reordering logic, which includes a reordering buffer and completion and interrupt synchronization modules.

[0055] Regarding balancing, there are two main parts that will be discussed. The first part is reordering the buffers, and the other part is completion and interrupt synchronization, which will be relative to... Figure 5 Further discussion. This balancing could be based on several factors, such as namespace IDs, ensuring fairness across each namespace by the data storage device. Another possibility is that balancing could be based on zone IDs in the storage device, on commit queues, or on completion queues. This balancing could be based on physical or virtual functions. This balancing could be based on LBAs. The foregoing is merely illustrative, as other bases for balancing are envisioned. For fairness to be achieved, balancing needs to be based on something.

[0056] Figure 5 The reordering logic is described in detail. Figure 5 This is a schematic diagram 500 illustrating the reordering logic arrangement according to an implementation scheme. All components in the device controller remain identical. Encryption and decryption, data paths, direct memory access (DMA), flash interface logic, etc., are present. NVMe components send data transactions to the reordering logic. The reordering logic parses the transactions, classifies them, and arranges them in the appropriate queues. The round-robin logic is responsible for retrieving entries from the queues and sending them to the host device in a fair manner.

[0057] This classification can be based on the following parameters: namespace ID; zone ID in the ZNS drive; commit and completion ID; PCIe physical and virtual functions; and / or host device address. Other parameters may also be considered.

[0058] Based on classification, data is placed in specific queues. For example, one queue could be a namespace ID queue, and another could be a submission queue. For each delivery type, a queue can exist, and a classifier is responsible for categorizing the packets by type and then arranging them in the appropriate queue. Only then is balancing done using a round-robin algorithm, etc. Weighted round-robin algorithms can be applied.

[0059] Another part of the logic is for completion, which includes the completion queue and the interruption of the publication to the host device. Special care should be taken with those transmissions because the logic reorders them, and therefore cannot assume that packets have been sent so that completion can be sent to the host device. This is not allowed in this case because there might be a scenario where the ordering logic decides to reorder the transmissions and takes time until that logic decides to send the packets to the host device, and thus the completion publication logic should be aware of this.

[0060] Data transfer associated with this command is delayed, and therefore the completion message associated with this command should also be delayed. Otherwise, race conditions may occur, and therefore, this logic is responsible for synchronization. The reordering buffer adds a delay before submitting data to the host device, while the NVMe logic assumes that the data has already been transferred. Therefore, the NVMe logic can transfer completion and interruption messages, but the reordering logic cannot immediately transfer this message to the host device. This logic tracks the relevant data and only transfers the completion / interruption message to the host device when the data transfer associated with the command is complete.

[0061] An example where fairness is based on host device addresses will now be described. It should be understood that the host device addresses are merely an example, and other bases can be used to determine fairness. Additionally, the main memory space is 32 bits, with half allocated to host A and the other half to host B. The memory devices are responsible for balancing packets among these host devices at a packet granularity. Figure 6 An example is presented. Figure 6 This is a schematic diagram 600 of two host systems according to one implementation scheme. Combined with Tables I and II, there are examples of fairness based on LBA.

[0062] In this example, host A will handle the higher address, and host B will handle the lower address. For example, 1K needs to be transferred from address A to host B, and Table I shows the internal logic implementation and the order in which the device controller wants to send the transfers to the host devices. Therefore, Table I shows the internal order before the reordering logic.

[0063] Table I

[0064]

[0065]

[0066] Table I shows the traffic sent by the NVMe component to the reorder buffer. Four addresses are associated with host A, and then four addresses are associated with host B. What will happen in this scenario is that all those packets, in the order shown in Table I, will be sent to host A, and host A will be saturated while host B will be idle, which is undesirable. Because a 4KB transaction is published to host A, and only then is a 4KB transaction published to host B, the transmission at the packet level is unbalanced for the purposes of Table I. Previously, Table I showed the traffic to be monitored on the host device interface.

[0067] A better approach is reordering, as shown in Table II. Table II illustrates the balanced traffic visible to the host devices after the reordering logic. This logic is responsible for the reordering. After reordering, the first packet will go to host A, then the second packet will go to host B, and so on. This process ensures that the interface is utilized to maximize performance. It can be seen that fairness is achieved at the packet granularity in Table II, while Table I does not.

[0068] Table II

[0069] address size destination 0xA000 1KB Host A 0x8000_0000 1KB Host B 0xA400 1KB Host A 0x8000_0400 1KB Host B 0xA800 1KB Host A 0x8000_0800 1KB Host B 0xAC00 1KB Host A 0x8000_0C00 1KB Host B

[0070] In yet another implementation, fairness granularity can be configured as, for example, packet granularity, transmit size, etc. A 1KB transmit size is just an example.

[0071] Figure 7 This is a flowchart 700 illustrating a transaction balancing process according to one embodiment. The process involves receiving transaction packets from a storage device at block 702, sorting the transaction packets at block 704, and placing the transaction packets into a queue at block 706. Blocks 702, 704, and 706 are repeated sequentially as more transaction packets are received from the storage device. Once any packet is in a queue, the data storage device determines at block 708 whether the controller is ready to send the packet to any host device. This determination involves determining whether any interface between the host device and the controller is available for packet transmission. If no interface is available, block 708 continues looping. If at least one interface is ready, at block 710, the controller selects which interface is ready by selecting which host device to send the packet to. Then, at block 712, an appropriate queue is selected for the selected host device, and then at block 714, the selected packet is sent from the selected queue to the selected host device.

[0072] Packet-level fairness is achieved through balancing, enabling the integration of asymmetric systems. Balancing is accomplished by saturating the host interface. Balancing transmissions over the interface optimizes performance, maintains reliability, and supports scalability.

[0073] In one embodiment, a data storage device includes: a memory device; and a controller coupled to the memory device, wherein the controller is configured to: classify transactions to be sent to one or more host devices, wherein the transactions are at the group level; reorder the transactions based on the classification; and transmit the reordered transactions to the one or more host devices. The classification is based on one or more of the following: namespace identifier (ID), zone ID, commit ID, completion ID, peripheral component interconnect (PCIe) physical function, PCIe virtual function, and host device address. The reordering includes placing the classified transactions into different queues. The controller is configured to select a queue for transmission from the different queues. The controller is configured to delay the issuance of the completion of a command corresponding to a transaction until the transaction for the command has been transmitted. The one or more host devices are multiple host devices, and wherein the controller is configured to ensure interface saturation between the controller and the multiple host devices. The controller includes a host interface module (HIM), and wherein the HIM includes a reordering buffer and a completion and interrupt synchronization module. The reordering buffer utilizes a round-robin technique to determine when to transmit the reordered transactions. The reordering buffer is configured to add a delay before reordered transactions are transmitted to the host device. One or more host devices comprise multiple host devices, wherein a first host device among the multiple host devices has a different saturation level compared to a second host device among the multiple host devices.

[0074] In another embodiment, a data storage device includes: a memory device; and a controller coupled to the memory device, wherein the controller includes: a host interface module (HIM) including a reordering buffer and a completion and interrupt synchronization module, wherein the HIM is configured to maintain a first interface between the controller and a first host device, and wherein the HIM is configured to maintain a second interface between the controller and a second host device; a flash interface module (FIM) coupled to the memory device; and a command scheduler coupled between the HIM and the FIM, wherein the controller is configured to: balance traffic between the first host device and the second host device, wherein balancing includes ensuring saturation of the first and second interfaces. The first and second interfaces are asymmetric. The controller is configured to synchronize completion and interrupt messages toward the first and second host devices. The first host device is a physical function, and the second host device is a virtual function. The reordering buffer is configured to maintain multiple queues for placing classified transactions. The reordering buffer is configured to add a delay before transmitting data to the first or second host device. The completion and interruption synchronization module is configured to track data transfers over the first and second interfaces to avoid race conditions.

[0075] In another embodiment, a data storage device includes: a component for storing data; and a controller coupled to the component for storing data, wherein the controller is configured to: receive transaction packets from the component for storing data; classify the transaction packets; place the transaction packets in queues among a plurality of queues; send the transaction packets to a host device; and maintain saturation on the interface between the controller and the host device and on the interface between the controller and other host devices. The controller is configured to determine which host device among the host devices and other host devices should send data to. The controller is configured to select which queue among the plurality of queues should send data from.

[0076] While the foregoing relates to embodiments of this disclosure, other and further embodiments of this disclosure may be designed without departing from the basic scope of this disclosure, the scope of which is defined by the appended claims.

Claims

1. A data storage device, comprising: Memory devices; and A controller, coupled to the memory device, wherein the controller is configured to: Transactions to be sent to one or more host devices are categorized at the group level; The transactions are reordered based on the classification. as well as The reordered transactions are transmitted to the one or more host devices.

2. The data storage device according to claim 1, wherein the classification is based on one or more of the following: namespace identifier (ID), zone ID, commit ID, completion ID, peripheral component interconnect (PCI) fast (PCIe) physical function, PCIe virtual function, and host device address.

3. The data storage device of claim 1, wherein the reordering includes placing the categorized transactions into different queues.

4. The data storage device of claim 3, wherein the controller is configured to select a queue for the transmission from the different queues.

5. The data storage device of claim 4, wherein the controller is configured to delay the release of a command corresponding to a transaction until after the transaction for the command has been transmitted.

6. The data storage device of claim 1, wherein the one or more host devices are multiple host devices, and wherein the controller is configured to ensure interface saturation between the controller and the multiple host devices.

7. The data storage device of claim 1, wherein the controller includes a host interface module (HIM), and wherein the HIM includes a reordering buffer and a completion and interrupt synchronization module.

8. The data storage device of claim 7, wherein the reordering buffer utilizes a loop technique to determine when to transmit the reordered transactions.

9. The data storage device of claim 8, wherein the reordering buffer is configured to add a delay before transmitting the reordered transactions to the host device.

10. The data storage device of claim 1, wherein the one or more host devices comprise a plurality of host devices, and wherein a first host device among the plurality of host devices has a different saturation level compared to a second host device among the plurality of host devices.

11. A data storage device, comprising: Memory devices; and A controller coupled to the memory device, wherein the controller includes: A host interface module (HIM) includes a reordering buffer and a completion and interrupt synchronization module, wherein the HIM is configured to maintain a first interface between the controller and a first host device, and wherein the HIM is configured to maintain a second interface between the controller and a second host device. A flash memory interface module (FIM) coupled to the memory device; and A command scheduler, coupled between the HIM and the FIM, wherein the controller is configured to: Balance the traffic between the first host device and the second host device, wherein the balancing includes ensuring that the first interface and the second interface are saturated.

12. The data storage device according to claim 11, wherein the first interface and the second interface are asymmetric.

13. The data storage device of claim 11, wherein the controller is configured to synchronize completion and interrupt messages toward the first host device and the second host device.

14. The data storage device of claim 11, wherein the first host device is a physical function and the second host device is a virtual function.

15. The data storage device of claim 11, wherein the reordering buffer is configured to maintain a plurality of queues for placing classified transactions.

16. The data storage device of claim 11, wherein the reordering buffer is configured to add a delay before transmitting data to the first host device or the second host device.

17. The data storage device of claim 11, wherein the completion and interruption synchronization module is configured to track data transfers over the first interface and the second interface to avoid race conditions.

18. A data storage device, comprising: Components used for storing data; and A controller, coupled to the component for storing data, wherein the controller is configured to: Receive transaction packets from the component used for storing data; The transaction groups are classified; The transaction groups are placed in queues within multiple queues; Send the transaction group to the host device; as well as Maintain saturation on the interface between the controller and the host device, as well as on the interface between the controller and other host devices.

19. The data storage device of claim 18, wherein the controller is configured to determine which of the host devices and other host devices should send data to.

20. The data storage device of claim 19, wherein the controller is configured to select which of the plurality of queues should send data from.