NVME Packetization on a Virtual Output Queue Mapping Structure

By integrating SSDs in network switches and leveraging NVMe-oF protocol and virtual queueing technology, the problem of inefficient storage data transmission in software-defined infrastructure is solved, and a high-performance and reliable storage solution is achieved.

CN113994321BActive Publication Date: 2025-06-27HEWLETT PACKARD ENTERPRISE DEV LP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201980097393.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-06-25
Publication Date
2025-06-27
Estimated Expiration
2039-06-25

AI Technical Summary

Technical Problem

The prior art is difficult to achieve high-performance storage data transmission in software-defined infrastructure, especially in hardware resource virtualization and distributed environments.

Method used

By integrating SSD into network switches, the NVMe-oF protocol is used to realize lossless network storage, combining virtual queues and rate control mechanisms to optimize the access and management of storage devices.

Benefits of technology

It realizes efficient and reliable storage data transmission in a software-defined infrastructure, improves fault tolerance and high availability of storage devices, and enhances the adaptability and performance of network additional storage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113994321B_ABST
    Figure CN113994321B_ABST
Patent Text Reader

Abstract

Provided is a network infrastructure device (e.g., a network switch) for use by a remote application host, the device integrating solid-state drive (SSD) storage and using the Non-Volatile Memory Express (NVMe) data transfer protocol. A high-availability configuration of a network switch using direct rate control (RC) feedback for multiple submission queues mapped to SSD storage is provided. Structurally, NVMe (NVMe-oF) is an implementation of the NVMe protocol on a network fabric. Access to SSDs on the network fabric can be controlled using egress queue congestion accounting (associated with a single egress output) and a direct RC feedback signal between a source node receiving input / output commands from a remote host for an integrated SSD device. In some implementations, the direct RC feedback signal uses a hardware-based signal. In some implementations, the direct RC feedback signal is implemented in hardware logic (silicon chip logic) within the internal switching fabric of the network switch.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Some information technology departments in companies have started to build their computer infrastructure to be as software-defined as possible. Generally, such software-defined infrastructure sometimes relies on hyper-converged infrastructure (HCI), in which different functional components are integrated into a single device. One aspect of HCI is that hardware components can be virtualized into software-defined and logically isolated computational, storage, and network representations of the computer hardware infrastructure. HCI and virtualization of hardware resources can allow for flexible allocation of computing resources. For example, configuration changes can be applied to the infrastructure, and the underlying hardware simply adapts to the new software-implemented configuration. Some companies may further use HCI to achieve virtualized computers by fully defining the capabilities of the computer in software. Each virtualized computer (e.g., defined by software) can then utilize a portion of one or more physical computers (e.g., the underlying hardware). A recognized result of virtualization is that physical computing, storage, and network capacity can be utilized more efficiently across the organization.

[0002] Non-Volatile Memory Express (NVMe) is a data transfer protocol that is commonly used to communicate with solid-state drives (SSDs) over the Peripheral Component Interconnect Express (PCIe) communication bus. There are many different types of data transfer protocols for different uses in computer systems. Each transfer protocol can exhibit different characteristics in terms of speed and performance, and thus each protocol can be suitable for different uses. NVMe is an example of a data protocol that can be used to achieve high-speed data transfer between a host system and an SSD. NVMe is typically used in computers that require high-performance read and write operations on SSDs. Utilizing NVMe-based storage that can support high-performance reads and writes within a software-defined infrastructure that further leverages HCI hardware can represent a useful and adaptable configuration of the infrastructure network.

[0003] The specification for running NVMe (NVMe-oF) over fabrics was initiated in 2014. One goal of the specification is to extend NVMe to fabrics such as Ethernet, Fibre Channel, and InfiniBand, or any other suitable storage fabric technology. Access to SSD drives over a network fabric via NVMe-oF can allow software-defined storage capacity (e.g., a portion of a larger hardware storage capacity) to be extended for access. This extension for access can: a) allow access to a large number of NVMe devices; and b) extend the physical distance between devices (e.g., within a data center). The extension can include increasing the distance at which another computing device can access an NVMe storage device. Due to the nature of storage objectives, storage protocols are typically lossless protocols. If the protocol used for storage is lossy (lossy is the opposite of lossless), correct data storage is likely to exhibit unacceptably slow performance (e.g., due to packet transmission retries), and may even become corrupted (e.g., data inaccuracies) and thus unusable in a real-world computer environment. Therefore, NVMe-oF traffic over a network fabric is implemented as lossless. NVMe-oF network packets can be transmitted over the network along with other traffic. Thus, NVMe-oF traffic on an intermediate device (e.g., a network switch providing the network fabric between a host device and a storage device) may be on the same physical transmission medium (e.g., an optical cable or a cable) as other types of data. BRIEF DESCRIPTION OF THE DRAWINGS

[0004] When read in conjunction Figure 1 with the following detailed description, the present disclosure can be better understood. It should be emphasized that, according to industry standard practice, the various features are not drawn to scale. In fact, the dimensions or positions of functional attributes may be repositioned or combined based on design, security, performance, or other factors known in the field of computer systems. Additionally, for some functions, whether internally or among each other, the processing order can be changed. That is, some functions may not be implementable using serial processing and may thus be performed in a different order than shown or may be performed in parallel with each other. For a detailed description of the various examples, reference will now be made to the drawings, in which:

[0005] Figure 1 is a functional block diagram showing an example of a network infrastructure device such as a switch / router implemented according to one or more disclosures;

[0006] Figure 2A is a functional block diagram showing an example of a high-availability switch implemented according to one or more disclosures;

[0007] Figure 2Bis a functional block diagram showing an example of a high-availability switch that includes an SSD integrated within a high-availability switch as an example of a storage-capability-enhanced switch implemented according to one or more disclosures;

[0008] Figure 3 is a block diagram showing the logical layer of communication between a host application and an NVMe storage device implemented according to one or more disclosures, and the block diagram includes abstract considerations of the local data communication bus implementation of NVMe and a table of how these abstract considerations differ for a network-based implementation of NVMe;

[0009] Figure 4 is a block diagram showing a high-level example of an internal switching fabric connection that maps ingress inputs from multiple source nodes (through fabric nodes) to a target NVMe queue pair associated with a target SSD (including a direct feedback control loop) implemented according to one or more disclosures, where the SSD device is integrated within a network infrastructure device such as a network switch / router;

[0010] Figure 5 is an example process flow diagram depicting an example method for automatically applying rate control via a direct feedback control loop that can be integrated into a communication fabric locally implemented on a network infrastructure device with an integrated SSD device;

[0011] Figure 6 is an example computing device implemented according to one or more disclosures, the computing device having a hardware processor and accessible machine-readable instructions stored on a machine-readable medium (e.g., disk, memory, firmware, or instructions that may be implemented directly in hardware logic), the accessible machine-readable instructions being usable to implement Figure 5 the example method;

[0012] Figure 7 represents a computer network infrastructure implemented according to one or more disclosures, the computer network infrastructure being usable to implement all or part of the disclosed NVMe-oF queue mapping to integrate SSD storage into a network infrastructure device; and

[0013] Figure 8 illustrates a computer processing device that can be used to implement the functions, modules, processing platforms, execution platforms, communication devices, and other methods and processes of the present disclosure. Detailed Description

[0014] Illustrative examples of the subject matter claimed below will now be disclosed. For the sake of clarity, not all features of actual implementations are described for each example implementation in this specification. It should be understood that numerous implementation-specific decisions may be made in developing any such actual example to achieve the specific goals of the developer, such as compliance with system-related and business-related constraints, which will vary from one implementation to another. Moreover, it should be understood that such development efforts, even if complex and time-consuming, would be routine for those of ordinary skill in the art who would benefit from this disclosure.

[0015] As explained in more detail below, this disclosure provides an implementation of network-attached storage where storage devices can be directly integrated into network infrastructure components. In one example implementation, an SSD can be integrated into a line card of a high-availability network switch. A line card represents a plug-in component of a network switch that has an integrated communication fabric within a single (or directly connected) chassis. In a typical scenario, line cards are used to provide additional network ports for a network switch. However, those of ordinary skill in the art who would benefit from this disclosure will be able to recognize that integrating storage (specifically, NVMe storage on an SSD) into a network switch can provide several benefits over previously available network-attached storage implementations. For example, providing a line card that replaces a network port with an SSD storage device can allow the network switch to utilize virtual queues to map multiple ingress inputs to a single egress output associated with the SSD. Additionally, implementing rate control (RC) and queue management capabilities for components that communicate directly with the network switch's communication fabric can provide more efficient and reliable storage capabilities. Further, implementing control flow elements directly in the silicon (e.g., hardware logic) associated with the switch communication fabric can represent a greater improvement over existing capabilities for network-attached storage.

[0016] The fault tolerance and high availability (HA) capabilities of network-attached storage devices can be enhanced using an implementation of storage directly within a network switch, in part because the number of hops (e.g., intermediate devices) between a remote host application and the storage device communicating with that remote host application can be reduced. Put simply, according to some disclosed implementations, the storage device (e.g., an NVMe SSD within a network switch) has been repositioned to a "closer" location relative to the remote host application. Specifically, the amount of physical network traversed between the remote host application and the storage device has been reduced (and the number of intermediate devices can also be reduced). These and other capabilities for improving the functionality of computer systems and computer networks are explained in more detail throughout this disclosure.

[0017] Storage devices attached to a network infrastructure can be accessed simultaneously by many different remote devices. However, at a given moment, only a single remote device may be allowed to write to a specific location within the storage device. Thus, a many-to-one mapping can be implemented to allow for apparent concurrent access to the storage device. An example of allowing multiple input entries to interact with a single output can rely on the mapping of multiple virtual output queues to a single output. In some cases, multiple input network ports can provide input to separate virtual queues, which are then mapped to a single physical output network port. In this example, each individual virtual queue can be assigned a portion of the physical bandwidth that the physical network output port is capable of supporting. Software control and other rate control (RC) mechanisms can be used to control the flow rate of each virtual queue so that the physical port does not become overloaded beyond its functional capabilities.

[0018] Some virtual queue implementations can include quality of service (QoS) attributes such that higher-priority traffic is allocated bandwidth prior to lower-priority traffic. Thus, if a bottleneck (congestion) situation occurs at runtime, lower-priority traffic will be delayed (or dropped) to support higher-priority traffic. In a network where packets are dropped (and then, if possible, recovered from these dropped packets later), a traffic flow in which packet loss can occur is referred to as a "lossy" traffic flow. The opposite of a "lossy" traffic flow is referred to as a "lossless" traffic flow. In a lossless traffic flow, data packets arrive at their intended destination without significant packet loss.

[0019] To achieve a lossless traffic flow over a natively lossless communication technology (e.g., Ethernet), different techniques for oversubscription (or over-allocation) of bandwidth can be used. In one example, a data stream designed to support a communication speed of 1 MB / s can be allocated 5 MB / s of bandwidth. Since the bandwidth allocated for this example traffic flow is five times the bandwidth it is designed to use, the probability of data congestion can be greatly reduced and a lossless traffic flow can be produced. Of course, in most cases, such over-allocation will represent the reservation of computer resources that will remain idle most of the time to allow these resources to be available on demand. A disadvantage of oversubscription techniques is a scenario in which multiple traffic flows are similarly over-allocated within a network segment. This allocation of rarely used resources can result in sub-optimal utilization of that network segment (e.g., there will be wasted bandwidth to support the over-allocation).

[0020] NVMe-oF traffic on a network fabric and other communication fabrics such as an internal switching fabric is designed to be lossless, assuming that reads and writes are expected to be lossless. Any computer user can understand that reading and writing data to a drive or other storage device of a computer should result in all reads and writes being successfully processed. For example, when an operation to save data to a disk is discarded because the write is considered optional (e.g., lost in the communication stream), a college student is unlikely to accept that pages are missing from their term paper.

[0021] NVMe-oF network packets used to perform read and write operations (collectively referred to herein as "data transfer" operations) can be exposed to network infrastructure devices such as network switches or routers that must have the ability to handle NVMe-oF traffic without losing NVMe-oF network packets. The network infrastructure devices can use methods such as the Priority Flow Control (PFC) standard 802.1Qbb to support lossless port queues. It can be difficult to use lossless queues in network devices because NVMe-oF network packets need to be separated from non-NVMe-oF network packets. Unlike NVMe-oF network packets, non-NVMe-oF network packets can form a network data stream that is more resilient to network packet loss and thus can operate without using lossless queues. In this case, a lossless queue is a temporary storage location where network packets can be stored until they are transmitted to a receiver and the transmission is acknowledged. A lossless queue can utilize more computing power or memory resources to operate in network infrastructure devices that may have a limited amount of computing and memory resources. It may not be feasible to use lossless queues for all network packets in network infrastructure devices that may have limited available computing resources.

[0022] Current methods of separating NVMe-oF network packets from non-NVMe-oF network packets can be processed on network infrastructure devices configured to allow NVMe-oF network packets to be processed without loss. One method of separating NVMe-oF network packets from non-NVMe-oF network packets can be to configure the network infrastructure device to assume that network packets originating from a particular Internet Protocol (IP) address or having a particular destination IP address are NVMe-oF packets. For example, an integrated SSD in a network switch can have an associated IP address for lossless processing. Network packets from the IP address defined as the source of the NVMe-oF network packets can be routed to a lossless queue, while non-NVMe-oF network packets can be routed to other queues that may not need to provide lossless processing. As new sources of NVMe-oF network packets are added to the network, or existing sources of NVMe-oF network packets are removed, the network infrastructure device can be adjusted to correctly process network packets originating from the new or updated IP addresses. For large-scale deployments of NVMe devices to be accessed via NVMe-oF over a network fabric, manually updating the configuration of the network infrastructure device in response to network changes may not be desirable. Thus, the disclosed techniques for integrating SSDs within a network switch using virtual output queues can be implemented to support automatic configuration changes based on other changes detected within the network infrastructure without human intervention. For example, a plug-and-play type of automatic adjustment configuration of the device can be implemented.

[0023] The network infrastructure device can be programmed (e.g., using software controls or more preferably using a hardware feedback loop for rate control) to allow the lossless queue to accept the additional communication when the corresponding SSD is able to accept the additional communication. As described below with reference to Figure 4As discussed in more detail, a direct hardware feedback control mechanism based on each queue and source rate control (RC) can be implemented. Thus, congestion at the SSD or unavailability of the SSD (e.g., due to failure or removal of the SSD) can cause a missing signal to be provided to the source input, such that the source input automatically reacts to the inability of the target (i.e., the SSD) to accept additional communication (e.g., data transfer operations in the form of storage read requests or storage write requests). Overall, this can enhance the HA capabilities, as the source input can automatically redirect data to another still-available backup device of the unavailable SSD. A typical configuration of an HA device is designed to include at least one primary / backup pair of sub-devices that may encounter failures. Thus, a higher-level HA device is not affected in its operational capabilities related to its rules within the corporate infrastructure network due to a single sub-device failure. For example, an HCI network switch with the disclosed feedback loop thus has enhanced HA capabilities. A timing mechanism can also be implemented to allow for transient congestion scenarios where the sub-device has not failed but has just become temporarily busy. In such cases, the disclosed RC mechanism can provide a signal that allows additional communication with the associated device after a slight delay. Many different implementations are possible and can depend on different design considerations, including the type of data storage being implemented or the QoS techniques described above.

[0024] Now referring to Figure 1 , a network infrastructure device 100 such as a switch / router 105 is shown in block diagram form. Generally, a router has two types of network element components that are organized onto separate planes shown as a control plane 110 and a data plane 115. Additionally, a typical switch / router 105 can include processing resources and local data storage 120. Depending on the capabilities of the particular switch / router 105, there can be different types of processing resources and local storage (for internal device use). Generally speaking, a higher-capacity switch / router 105 implementation will include a large amount of processing resources and memory, while a simpler (e.g., low-capacity) device will contain fewer internal resources. As described throughout this disclosure, local storage for internal device use should not be confused with attachable or integrated storage devices for network use (e.g., SSDs).

[0025] The control plane 110 (e.g., in a router) can be used to maintain a routing table (or a single consolidated routing table) that lists which route should be used to forward data packets and through which physical interface connection (e.g., output ports 160 to 169). The control plane 110 can perform this function by using internally pre-configured instructions (called static routing) or by dynamically learning routes using a routing protocol. Static and dynamic routes can be stored in one or more routing tables. The control plane logic can then strip non-essential instructions from the table and construct a Forwarding Information Base (FIB) for use by the data plane 115.

[0026] The router can also use a forwarding plane (e.g., part of the data plane 115) that contains different forwarding paths (e.g., forwarding path A 116 or forwarding path Z 117) for information from different ports or different destination addresses. Typically, the router forwards data packets between incoming (e.g., ports 150 - 159) and outgoing interface connections (e.g., ports 160 - 159). The router uses the information contained in the packet header to forward the data packet to the correct network type, which matches an entry in the FIB provided by the control plane 110. Ports are typically bi-directional and are shown as "input" or "output" in this example to illustrate the message flow through the routing path. In some network implementations, the router (e.g., switch / router 105) can have interfaces for different types of physical layer connections, such as copper cables, fiber optics, or wireless transmissions. A single router can also support different network layer transmission standards. Each network interface can be used to enable data packets to be forwarded from one transmission system to another. The router can also be used to connect two or more logical groups of computer devices called subnets, each logical group having a different network prefix.

[0027] Also shown Figure 1 in the figure, the bi-directional arrow 107 indicates that the control plane 110 and the data plane 115 can work in a coordinated manner to achieve the overall capabilities of the switch / router 105. Similarly, the bi-directional arrow 125 indicates that the processing and local data storage 120 can interface with the control plane 110 to provide processing and storage support for the capabilities assigned to the control plane 110. The bi-directional arrow 130 indicates that the processing and local data storage 120 can also interface with the data plane 115 as needed.

[0028] As Figure 1As shown, the control plane 110 includes several example functional control boxes. Depending on the capabilities of the specific implementation of the network infrastructure device 100 (e.g., switch / router 105), additional control boxes are possible. Box 111 indicates that the control plane 110 may have associated build information regarding the software version of the control code currently executing on the network infrastructure device 100. Additionally, this software version may include configuration settings for determining how the network infrastructure device 100 and its associated control code perform different functions.

[0029] Many different configuration settings of the software and the device itself are possible, and describing each is beyond the scope of this disclosure. However, the disclosed implementation of integrating SSDs into switches and the corresponding handling of NVMe-oF network packets can be implemented in one or more subsystems of the network infrastructure device 100. Additionally, in some implementations such as Figures 2A - 2B As shown, the network infrastructure device 100 (e.g., switch / router 105 or HA switches 200A and 200B) may consist of multiple devices with different HA configurations. One or more devices in the network infrastructure device 100 may be configured to implement the automatic detection and routing of NVMe-oF network packets and feedback loops to provide an RC for mapping multiple virtual output queues to a single SSD (or port) within the network infrastructure device 100.

[0030] Continue Figure 1 Continuing, box 111 indicates that the switch / router 105 (as an example of the network infrastructure device 100) and the control plane 110 may know different types of routing information and connection information. Box 112 indicates that information storage can be accessed from the control plane 110 and the information storage includes a forwarding table or NAT information as appropriate. Box 113 indicates that the control plane 110 may also know forwarding decisions and other processing information. Although Figure 1 These logical capabilities within the control plane 110 are shown, they may actually be implemented outside the control plane 110 but be accessible to the control plane 110.

[0031] Now refer to Figure 2A, an example of a high-availability switch 205A is shown in block diagram 200A. The high-availability switch 205A is shown as having two controllers. Controller 1 (210) is identified as the "active" controller, and controller 2 (215) is identified as the "standby" controller. As explained in more detail below, high-availability switches such as high-availability switch 205 can have any number of controllers and typically have at least two. In some configurations, the controllers form a primary / backup pair with a dedicated active controller and a dedicated standby controller. In a primary / backup configuration, the primary device performs all network functions, and if a failover condition is reached, the standby device will wait to become active. Failover can be automatic or manual and can be implemented for different components within a higher-level HA device. Generally speaking, the concept of high-level failover refers to the active and standby components switching roles so that the standby device becomes the active device and the active device (sometimes after a restart or replacement) becomes the standby device. In the case of SSD devices integrated into a network switch, one SSD can act as the primary device in a redundant SSD pair, and these SSDs are kept in sync with data writes so that when the primary SSD becomes unavailable (for various reasons) automatically, the backup device of the redundant pair can take over (e.g., the backup device is a hot backup device).

[0032] The high-availability switch 205A also includes a plurality of communication cards (e.g., card slot 1 (221), card slot 2 (222), card slot 3 (223), and card slot N (225)), and each communication card can have a plurality of communication ports configured to support network communication. Card slots such as card slot 1 (221) can also be referred to as "line cards" and have a plurality of bidirectional communication ports (as well as management ports (not shown)). Card slot 1 (221) is shown as having port 1-1 (241) and port 1-2 (242), and can represent a "card" inserted into a slot (e.g., a communication bus connection) in the backplane (e.g., a communication bus) of the high-availability switch 205A. Other connections and connection types are also possible (e.g., cable connections, NVMe devices, etc.). Additionally, in Figure 2A it, card slot 2 (222) is shown as having port 2-1 (243) and port 2-2 (244); card slot 3 (223) is shown as having port 3-1 (245), 3-2 (246), and port 3-N (247); card slot N (225) is shown as having port X (248) and port Y (249).

[0033] To support communication between a controller (e.g., active and / or standby controller) in a switch and client devices connected to the switch, multiple communication client applications can be executed on a given switch. The client applications executed on the switch can assist in communication with the connected clients and hardware configuration on the switch (e.g., ports of line cards). In some cases, the client applications are referred to as "listeners", partly because they "listen" for communications or commands and then process what is received. For the high-availability switch 205A, an example client application is Client 1 (230-1), which is shown as supporting communication from either the active controller or the standby controller to devices connected via slot 1 (221).

[0034] Figure 2A A second example client application in [the switch] is Client 2 (230-2), which is shown as supporting communication from either controller to both slot 2 (222) and slot 3 (223). Finally, Client Z (230-Z) is shown as supporting communication from both controllers to slot N (225). The dashed lines from the standby controller 2 to the client applications in block diagram 200 indicate that the standby controller can be communicatively coupled to the communication slots via the client applications, but due to its standby state, may not transmit significant data. The solid lines from the active controller 1 to the client applications in block diagram 200 indicate the active state where more communication may occur. In some examples, a single client can be configured to support more than one (or even a part of one) communication slot (line card), as shown where Client 2 (230-2) concurrently supports both slot 2 (222) and slot 3 (223). The upper limit on the number of slots a client supports can be an implementation decision based on the performance characteristics of the switch or other factors and its internal design.

[0035] Reference Figure 2B, Block diagram 200B shows HA switch 205B as a variant of the aforementioned HA switch 205A. As shown, in region 255 (outlined by the dashed box), HA switch 205B integrates multiple SSD components, which can be used to provide network-attached storage for remote devices. As shown, SSD devices can be used in place of the communication ports of HA switch 205B. Specifically, communication card slot 2 (252) integrates SSD 2-1 (250-1) and SSD 2-2 (250-2). Depending on the implementation specification, SSD 2-1 (250-1) can be paired with SSD 2-2 (250-2) as a redundant pair of storage devices, or can be implemented independently of each other. Since both SSD 2-1 (250-1) and SSD 2-2 (250-2) are on card slot 2 (252), it may be necessary to provide a redundant pair where the primary and backup devices of the redundant pair are not on the same line card. Specifically, an SSD can be paired with an SSD on a different line card to achieve redundancy. Any implementation is possible. One possible benefit of having inputs and outputs (or redundant pairs) on the same line card is that communication between devices on the same line card does not have to traverse the chassis structure (i.e., device-to-device communication will be local to the line card structure). Of course, different implementation criteria can be considered to determine the best implementation for a given application solution.

[0036] Also as shown in example HA switch 205B, a line card can communicate with any number of integrated SSD components. Specifically, region 255 shows that SSD 3-1, SSD 3-2, and SSD 3-N (all denoted by reference numeral 251) can be integrated with (or connected to) card slot 3 (253). In this example, client 2 (230-2) can be adapted to communicate with a line card having integrated SSD components, and other computing devices (e.g., outside region 255) may not be aware of the detailed implementation within region 255. That is, the implementation of the SSD components integrated within HA switch 205B can be transparent to external devices and other components of HA switch 205B. Although client 2 (230-2) is shown as a potential software (or firmware) module in block diagram 200B, the functionality of client 2 (230-2) can be fully (at least partially) implemented within the hardware logic (i.e., silicon-based logic) of HA switch 205B. Given the benefits of this disclosure, those of ordinary skill in the art will recognize that many different implementations of software, firmware, and hardware logic can be used to implement the disclosed technique of providing integrated SSDs (with RC control loops) for network switches.

[0037] Now referring to Figure 3 , a block diagram of logic layer 300 representing communication between a host application and an NVMe storage device is shown. Figure 3It also includes Table 1 (350), which lists the abstract considerations of the local data communication bus implementation of NVMe according to one or more disclosures and how these abstract considerations are different for the network-based implementation of NVMe (e.g., NVMe-oF). All devices (whether the device is local or remote) can have an NVMe Qualified Name (NQN) for addressing the device.

[0038] The logical layer 300 includes a top layer 305 representing a remote application, which can utilize NVMe storage to store data as part of the processing of the application. In this context, a remote application refers to a host application that executes remotely (e.g., via a network) from the storage device storing the data of the host application. The host application executing as a local application or a remote application can have multiple internal abstractions and is unaware of the actual storage physical location updated when the application performs read or write operations.

[0039] Layer 310 indicates that an application interacting with NVMe-oF storage can interface with a host-side transport abstraction layer, which can present itself to the application in a manner similar to a device driver used to interface with locally attached storage. The host-side transport abstraction can provide fabric communication or network connection communication to perform read or write operations on behalf of the application. Layer 315 indicates that a host-initiated command (e.g., which may include data for writing) can transition from layer 310 to layer 320, and a response command (e.g., a response to a host-initiated command such as a read request) can transition in the opposite direction. Layer 315 can be a local bus communication implementation or can interface with remote physical storage (e.g., remotely across a network) based on a network protocol.

[0040] When communicating across a network, possible data encapsulation for transmission can be automatically provided to achieve network communication efficiency. Layer 320 indicates that a device controller (local or remote) can receive information from layer 315 and interface to the storage device. For example, layer 315 can include a data request (e.g., read or write) flowing to the physical device in a first direction and a response command (e.g., data) flowing in a second direction to return from the physical device to the application. In some disclosed implementations, the device controller represented at layer 320 can include communication via a switch fabric to map virtual queues (possibly many) to a single device (e.g., an SSD). Layer 325 represents an NVMe SSD or other NVMe storage device that stores or provides the actual data obtained by the remote application.

[0041] Table 1 (350) shows different abstract considerations for implementations that communicate with NVMe storage over a local bus versus implementations that communicate remotely with NVMe storage (e.g., NVMe-oF). As shown, when the communication is to a local NVMe device (e.g., a device connected to a communication bus rather than over a network), the device identifier (i.e., for an SSD) can include bus / device / function attributes, while in remote communication, the device identifier can be represented by the NQN referenced above. The queuing basis for remote data access versus local data access can also be different. In a local implementation, a memory queuing mechanism can be utilized, while for a remote storage implementation, queuing can utilize discovery and connection commands as part of performing network communication operations. In a local implementation, discovery techniques can include data bus enumeration (e.g., looking for devices on the data bus), while a remote implementation can include using network / fabric-based messages to discover available storage devices. For data transfer, a local implementation can use physical memory page addresses (PRP) or scatter-gather lists (SGL) to provide data to the storage device, while a remote implementation can utilize SGLs with additional keys to perform data transfer operations. These are just examples of potential differences that are illustrated to explain that the abstraction layers can be different (sometimes slightly) to make the data access for applications as transparent as possible. That is, the application does not need to know where the actual data is physically stored and can operate normally regardless of the actual implementation details. These abstraction layers can further allow a system administrator to implement software configuration of the actual physical hardware and thus allow for greater flexibility in providing application services to an organization.

[0042] Reference Figure 4 , shows an example block diagram that shows an overview of a possible internal routing 400, which can be used by a network infrastructure device such as a HA switch 205B having an integrated SSD component as shown Figure 2B . In this example, the concept of a node can be used to describe a logical subsystem of a network infrastructure device or the concept of a logical processing block implemented inside or outside of a network infrastructure device. In this example, multiple source nodes 435 can receive network packets and act as ingress inputs for multiple queues contained within the source nodes 435. Each queue within a source node 435 can be coupled to rate control (RC) logic 485. Each source node 435 can be connected to multiple fabric nodes 440 via a fabric load balancer (FLB) 490. In this example, the fabric nodes 440 can represent chips on a line card that also contains ports for network communication or an integrated SSD component for providing storage to remote applications (e.g., a multi-queue SSD 482).

[0043] The connections from the source node 435 to the multiple fabric nodes 440 form multiple alternative paths 455, where network packets can be sent to the fabric nodes 440. The fabric nodes 440 may also have multiple internal queues for receiving network packets sent from the source node 435. The fabric nodes 440 may also have a load balancing mechanism (e.g., load balancer (LB) 495) for routing the received packets to the queues within the fabric nodes 440. The fabric nodes 440 may be commonly coupled to the destination node 450 (e.g., implemented to include NVMe submission / completion queue pairs 470). The coupling may be achieved through a connection mechanism such as the cross-connect communication bus 465 shown. The destination node 450 may be implemented as multiple NVMe submission / completion queue pairs 470 and have an egress output 481 to interface to one or more multi-queue SSDs 482. For purposes of brevity and clarity, Figure 4 only two source nodes 435, fabric nodes 440, and destination nodes 450 are shown, but there may be many different types of nodes, which may be configured according to the concepts of the present disclosure to provide an integrated internal switching fabric to further provide network attached storage for network infrastructure devices such as switch / router 105 or HA switch 205A or 205B, as described above.

[0044] In some example implementations, the NVMe protocol may be used as an interface to the SSD (e.g., multi-queue SSD 482) via the PCIe bus or via NVMe-oF. Each multi-queue SSD 482 may have 64K queue pairs, where each pair includes a submission queue for reads and writes and a completion queue for data transfer status. Simply put, each pair has a submission queue for commands and a completion queue for responses. Additionally, since 64K queue pairs may be more than the queue pairs required for any particular implementation, some implementations match a single queue pair of the SSD component to the corresponding core on the processing device. Thus, in the case where the processing device has 96 cores, each of these 96 cores may represent the source node 435. Thus, if Figure 4 this example of 96 implemented queue pairs is illustrated in the style of, N would be equal to 96, and there would be 96 instances of the egress output 481 of each multi-queue SSD 482 from each of the 96 instances of the NVMe queue pairs 450.

[0045] To ensure lossless queue processing and improve the HA implementation, each NVMe queue pair 450 is shown as including a sub-component called egress queue congestion accounting 480. The egress queue congestion accounting 480 represents a measure of the congestion of the egress queue, which is used to determine when (or if) a particular egress queue is unable to maintain the expected performance to complete data transfer requests. For example, if a particular egress queue exceeds a threshold number of queued (i.e., in-progress) requests, the threshold exceedance may indicate an indication of a performance problem. Alternatively, if a particular egress queue does not complete a large number of outstanding requests within a predetermined period of time, congestion may become a problem.

[0046] As shown by the direct feedback control 460, the egress queue congestion accounting 480 can be communicatively coupled to the input RC 485 of each NVMe queue pair 450. In Figure 4 , for clarity, not all instances of the feedback control 460 are shown. However, the feedback control 460 can be implemented between each instance of the egress queue congestion accounting 480 and each instance of the RC 485. Each instance of the egress queue congestion accounting 480 can be used to monitor the ability to accept new network packets at the SSD 482. For example, signals can be provided to many corresponding RC 485 instances to indicate that the transmission is available for egress output 481. Alternatively, if no availability signal is provided, the RC 485 can be notified that the transmission to this NVMe queue pair should be delayed or rerouted to another queue. Thus, if the storage device is unavailable (e.g., due to a fault, removal, congestion, etc.), the RC 485 can immediately become aware of the unavailability. Accordingly, the packet transmission to the destination will only occur in conjunction with the availability signal at the RC 485 of that destination.

[0047] This type of implementation using the RC control feedback 460 should not be confused with credit-based implementations used in some networks, partly because Figure 4 the RC control feedback 460 represents a per-queue direct feedback control mechanism. In some implementations, the RC control feedback 460 can be implemented using direct connections or direct wired connections set within a silicon-based switching fabric to provide hardware signals between the egress queue congestion accounting 480 and the RC 485. Finally, according to some implementations, internal switching fabric connectivity 499 can be provided for all components shown within the dashed outline as shown. That is, all components shown within the outline of the dashed box corresponding to the internal switching fabric connection can have a direct connection (or be implemented within or as part of the communication fabric) to the communication fabric inside the network infrastructure device (e.g., switch / router 105, HA switch 205A, or HA switch 205B). The exact position of the dashed box is Figure 4 approximate in

[0048] In one example implementation, a network infrastructure device including internal fabric connectivity 499 may receive a first network transmission from a remote application communicatively coupled to the network infrastructure device via a network. The first network transmission may be a data transfer operation (e.g., associated with fetching or storing data) using the Non-Volatile Memory Express over Fabrics (NVMe-oF) protocol. After receiving (or as part of) the network transmission, information retrieved from the network transmission may be queued to form one or more submission commands for routing and processing using the internal fabric connectivity 499 of the network infrastructure device. The internal fabric connectivity 499 may be configured using software or implemented using hardware logic. Each submission command may then be associated with a submission queue (e.g., NVMe queue pair 1 450, as shown in Figure 4 ) of an NVMe storage driver that interfaces with both the internal fabric connectivity 499 and an NVMe storage device (e.g., multi-Q NVMe SSD 482). The association of the submission command with the submission queue may be implemented using virtual output queues, each of which includes a completion queue of the NVMe storage driver. See, for example, Figure 4 NVMe queue pair 1 450 to NVMe queue pair N 450.

[0049] Continuing with the example, a second network transmission may be received at the network infrastructure device. The second network transmission may be received after the processing of the first network transmission has been completed, or may occur while the first network transmission is being processed (e.g., in a queue for processing). When a large amount of data or a large number of small amounts of data are being processed simultaneously (e.g., a heavy data transfer load), it is more likely that overlapping processing of multiple network transmissions providing data transfer operations will occur. Therefore, it may be necessary to control the flow of multiple concurrently processed data transfer operations. The techniques of the present disclosure include using a direct rate control (RC) feedback signal between egress congestion accounting and an associated source node (e.g., the input of the internal fabric connectivity 499) to control the flow of the second network transmission (and all other subsequent network transmissions). In this way, the source node input may be directed to available queues for processing (as opposed to being directed to queues with backlogs). An example is shown as the direct feedback control signal 460 in Figure 4 shown. As explained below with reference to Figure 5 , the RC availability signal may indicate that a queue pair is available and the lack of an availability signal may indicate that a queue pair is unavailable (e.g., busy or inoperable). Note that if a storage device encounters a failure, the techniques of the present disclosure may seamlessly route subsequent data transfer operations to a different device due to the lack of an availability signal for the queue pair associated with the failed device. Thus, the techniques of the present disclosure may address performance improvements as well as improved availability of data transfer operations.

[0050] Reference Figure 5 , as an example method 500, provides a process flow diagram that depicts an example of the logic applied to route multiple virtual output queues (egress inputs) to a single egress output associated with an SSD. Starting at block 510, a switch (or other network infrastructure device) establishes a connectivity fabric within the switch to store and retrieve data from SSD components. In this example, the protocol used to interface with the SSD components is NVMe, which can be associated with NVMe-oF network packets. Block 520 indicates that an input is received at the source node queue. Decision 530 indicates that a determination of rate control (RC) availability can be made. In some implementations, RC availability can be determined by a signal from an RC control feedback loop, as discussed above with reference to Figure 4 . If RC availability is affirmative (the "yes" branch of decision 530), the flow continues to block 570, where the input is provided from the source node to the appropriately associated NVMe queue pair via the internal switching fabric. The NVMe queue pair is further associated with the egress output and the SSD components.

[0051] Alternatively, if there is no RC availability signal at decision 530 (the "no" branch of decision 530), different implementations can implement a lossless communication flow and high availability support. One possible implementation is indicated at block 540, where the input can be held for a period of time before an RC availability signal is received. This possible implementation is further reflected by a configurable retry loop 535, where multiple retries can be performed at decision 530 to check for an RC availability signal. Block 550 indicates that if the retry loop (or the initial source input) has not been processed via block 570, the input command and data can optionally be redirected to a different source input node (and ultimately to a different egress output). Block 560 indicates that information can optionally be provided "upstream" of the remote device that is providing an input command or data to an unavailable destination, so that the upstream application can be notified that additional communication can be redirected to a different destination.

[0052] Each of the optional actions described above for Figure 5 can be selected based on configuration information that can be provided by a system administrator and can be further adjusted based on the runtime information of the overall network and the device performance. In some implementations, the configuration option can indicate that no retries (as indicated by retry loop 535) will be performed and the input should be blocked immediately based on the absence of an RC availability signal (as indicated at decision 530). Additionally, in a system experiencing optimal throughput, no optional actions may be invoked (eg, the RC availability signal persists).

[0053] Now refer to Figure 6, shows an example computing device 600 implemented according to one or more disclosed examples. The computing device 600 has a hardware processor 601 and accessible machine-readable instructions stored on a machine-readable medium and / or hardware logic 602, which can be used to perform automatic NVMe-oF network packet routing. As an example, Figure 6 shows a computing device 600 configured to execute the flow of method 500. However, the computing device 600 can also be configured to execute the flow of other methods, techniques, functions, or processes described in this disclosure. In Figure 6 this example, the machine-readable storage medium 602 includes instructions that cause the hardware processor 601 to execute the blocks 510 - 570 discussed above with reference to Figure 5 . Different implementations of method 500 are possible, including hardware logic configured on a chip to implement all or part of method 500 in combination with an overall implementation of the disclosed technology to provide an integrated SSD within a network infrastructure device.

[0054] A machine-readable storage medium (such as Figure 6 602) can include volatile and non-volatile removable and non-removable media, and can be any electronic, magnetic, optical, or other physical storage device that contains or stores executable instructions, data structures, program modules, or other data accessible by a processor, such as firmware, erasable programmable read-only memory (EPROM), random access memory (RAM), non-volatile random access memory (NVRAM), optical discs, solid-state drives (SSD), flash memory chips, etc. The machine-readable storage medium can be a non-transitory storage medium, where the term "non-transitory" does not include transitory propagated signals.

[0055] Figure 7 represents a computer network infrastructure 700 according to one or more disclosed embodiments. The computer network infrastructure 700 can be used to implement all or part of the disclosed automatic NVMe-oF network packet routing technology to directly integrate storage within a network infrastructure device. The network infrastructure 700 includes a set of networks in which embodiments of this disclosure can operate. The network infrastructure 700 includes a customer network 702, a network 708, a cellular network 703, and a cloud service provider network 710. In one embodiment, the customer network 702 can be a local private network, such as a local area network (LAN) including various network devices (including but not limited to switches, servers, and routers).

[0056] Each of these networks can include wired or wireless programmable devices and use any number of network protocols (e.g., TCP / IP) and connection technologies (e.g., network or )Perform operations. In another embodiment, the customer network 702 represents an enterprise network, which may include or be communicatively coupled to one or more local area networks (LANs), virtual networks, data centers, and / or other remote networks (e.g., 708, 710). In the context of the present disclosure, the customer network 702 may include one or more high-availability switches or network devices using methods and techniques such as those described above. Specifically, the computing resources 706B and / or the computing resources 706A may be configured to incorporate network infrastructure devices of storage devices (e.g., 707A and 707B).

[0057] As Figure 7 shown, the customer network 702 can be connected to one or more client devices 704A-E and allow the client devices 704A-E to communicate with each other and / or with the cloud service provider network 710 via the network 708 (e.g., the Internet). The client devices 704A-E can be computing systems, such as a desktop computer 704B, a tablet computer 704C, a mobile phone 704D, a laptop computer (shown as wireless) 704E, and / or other types of computing systems generally shown as the client device 704A.

[0058] The network infrastructure 700 may also include other types of devices commonly referred to as the Internet of Things (IoT) (e.g., edge IoT devices 705), which may be configured to send and receive information via the network to access cloud computing services or interact with a remote web browser application (e.g., to receive configuration information).

[0059] Figure 7 It is also shown that the customer network 702 includes local computing resources 706A-C, which may include servers, access points, routers, or other devices configured to provide local computing resources and / or facilitate communication between networks and devices. For example, the local computing resources 706A-C may be one or more physical local hardware devices, such as the HA switches outlined above. The local computing resources 706A-C may also facilitate communication between other external applications, data sources (e.g., 707A and 707B), and services and the customer network 702.

[0060] The network infrastructure 700 also includes a cellular network 703 for use with mobile communication devices. The mobile cellular network supports mobile phones and many other types of mobile devices, such as laptop computers, etc. The mobile devices in the network infrastructure 700 are shown as the mobile phone 704D, the laptop computer 704E, and the tablet computer 704C. Mobile devices such as the mobile phone 704D can interact with one or more mobile provider networks when the mobile device is moving, typically interacting with multiple mobile network towers 720, 730, and 740 to connect to the cellular network 703.

[0061] Figure 7 It shows that the customer network 702 is coupled to the network 708. The network 708 may include one or more computing networks available today, such as other LANs, wide area networks (WANs), the Internet, and / or other remote networks, for transmitting data between the client devices 704A-D and the cloud service provider network 710. Each computing network within the network 708 may include wired and / or wireless programmable devices operating in the electrical and / or optical domains.

[0062] In Figure 7 it, the cloud service provider network 710 is shown as a remote network (e.g., a cloud network) capable of communicating with the client devices 704A-E via the customer network 702 and the network 708. The cloud service provider network 710 acts as a platform for providing additional computing resources to the client devices 704A-E and / or the customer network 702. In one embodiment, the cloud service provider network 710 includes one or more data centers 712 having one or more server instances 714. The cloud service provider network 710 may also include one or more frames or clusters (and cluster groups) representing scalable computing resources that may benefit from the techniques of the present disclosure. Additionally, cloud service providers typically offer near-perfect uptime availability and may use the disclosed techniques, methods, and systems to provide that level of service.

[0063] Figure 8 It shows a computing device 800 that can be used to implement the functions, modules, processing platforms, execution platforms, communication devices, and other methods and processes of the present disclosure, or used in conjunction therewith. For example, Figure 8 the computing device 800 shown may represent a client device or a physical server device, and depending on the level of abstraction of the computing device, includes one or more (multiple) hardware or virtual processors. In some cases (without abstraction), as Figure 8 shown, the computing device 800 and its elements each relate to physical hardware. Alternatively, in some cases, one, more, or all elements may be implemented using an emulator or a virtual machine as the level of abstraction. In any case, regardless of the level of abstraction from the physical hardware, the computing device 800 at its lowest level can be implemented on physical hardware.

[0064] Also as Figure 8 shown, the computing device 800 may include: one or more input devices 830, such as a keyboard, mouse, touchpad, or sensor reader (e.g., a biometric scanner); and one or more output devices 815, such as a display, audio speaker, or printer. Some devices may also be configured as input / output devices (e.g., a network interface or a touchscreen display).

[0065] The computing device 800 may also include a communication interface 825 communicatively coupled to the processor 805, such as a network communication unit that may include wired communication components and / or wireless communication components. The network communication unit may utilize any of a variety of proprietary or standardized network protocols (such as Ethernet, TCP / IP, to name just a few) to enable communication between devices. The network communication unit may also include one or more transceivers that utilize Ethernet, power line communication (PLC), WiFi, cellular, and / or other communication methods.

[0066] As Figure 8 shown, the computing device 800 includes processing elements, such as a processor 805 that includes one or more hardware processors, where each hardware processor may have a single or multiple processor cores. As described above, each of the multiple processor cores may be paired with an NVMe queue pair to facilitate the implementation of the present disclosure. In one embodiment, the processor 805 may include at least one shared cache that stores data (e.g., computing instructions) used by one or more other components of the processor 805. For example, the shared cache may be local cache data stored in memory for faster access by components that make up the processing elements of the processor 805. In one or more embodiments, the shared cache may include one or more mid-level caches, such as level 2 (L2), level 3 (L3), level 4 (L4), or other levels of cache, last level cache (LLC), or a combination thereof. Examples of processors include, but are not limited to, a central processing unit (CPU), a microprocessor. Although Figure 8 not shown, the processing elements that make up the processor 805 may also include one or more other types of hardware processing components, such as a graphics processing unit (GPU), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), and / or a digital signal processor (DSP).

[0067] Figure 8Illustrated is that the memory 810 is operable and communicatively coupled to the processor 805. The memory 810 can be a non-transitory medium configured to store various types of data. For example, the memory 810 can include one or more storage devices 820, including non-volatile storage devices and / or volatile memories. A volatile memory such as random access memory (RAM) can be any suitable non-permanent storage device. The non-volatile storage device 820 can include one or more disk drives, optical drives, solid state drives (SSDs), tape drives, flash memories, read only memories (ROMs), and / or any other type of memory designed to retain data for a period of time after a power-off or shutdown operation. In some cases, if the allocated RAM is not sufficient to accommodate all working data, the non-volatile storage device 820 can be used to store overflow data. The non-volatile storage device 820 can also be used to store programs that are loaded into the RAM when the programs are selected for execution.

[0068] As is known to those of ordinary skill in the art, software programs can be developed, coded, and compiled in various computing languages for various software platforms and / or operating systems, and subsequently loaded and executed by the processor 805. In one embodiment, the compilation process of a software program can convert program code written in a programming language into another computer language such that the processor 805 can execute the program code. For example, the compilation process of a software program can generate an executable program that provides the processor 805 with encoded instructions (e.g., machine code instructions) to complete a specific, non-general, particular computing function.

[0069] After the compilation process, the encoded instructions can then be loaded from the storage device 820, from the memory 810, into the processor 805 and / or embedded within the processor 805 (e.g., via a cache or on-board ROM) as computer-executable instructions or process steps. The processor 805 can be configured to execute the stored instructions or process steps in order to execute the instructions or process steps to transform the computing device into a non-general, specific, specially-programmed machine or apparatus. The stored data (such as the data stored by the storage device 820) can be accessed by the processor 805 during the execution of the computer-executable instructions or process steps to direct one or more components within the computing device 800.

[0070] The user interface (e.g., output device 815 and input device 830) may include a display, a position input device (such as a mouse, touchpad, touch screen, etc.), a keyboard, or other forms of user input and output devices. The user interface components may be communicatively coupled to the processor 805. When the output device is a display or includes a display, the display may be implemented in various ways, including by a liquid crystal display (LCD), a cathode ray tube (CRT), or a light emitting diode (LED) display, such as an organic light emitting diode (OLED) display. As is known to those of ordinary skill in the art, the computing device 800 may include Figure 8 other components known in the art that are not explicitly shown, such as sensors, power supplies, and / or analog-to-digital converters.

[0071] Throughout this specification and the claims, certain terms are used to refer to particular system components. As will be understood by those skilled in the art, different parties may refer to components by different names. This document is not intended to distinguish between components that differ in name but not in function. In this disclosure and the claims, the terms "including" and "comprising" are used in an open-ended manner and should therefore be interpreted as "including but not limited to...". Additionally, the term "couple" or "couples" is intended to mean an indirect or direct wired or wireless connection. Thus, if a first device is coupled to a second device, that connection may be made through a direct connection or through an indirect connection via other devices and connections. The statement "based on" is intended to mean "at least partially based on". Thus, if X is based on Y, then X may be a function of Y and any number of other factors.

[0072] The foregoing discussion is intended to illustrate the principles and various implementations of the present disclosure. Once the above disclosure is fully understood, many variations and modifications will become apparent to those skilled in the art. It is intended that the following claims be interpreted to include all such variations and modifications.

Claims

1. A method for communication, comprising: Receiving, at a network infrastructure device, a first network transmission provided by a remote application communicatively coupled to the network infrastructure device via a network, the first network transmission being associated with a first data transfer operation using the non-volatile memory express over fabrics (NVMe-oF) protocol; Queuing the first network transmission into one or more submission commands for providing to an internal switching fabric of the network infrastructure device, the one or more submission commands being associated with a plurality of respective submission queues of an NVMe storage driver interfacing with the internal switching fabric and a non-volatile memory express (NVMe) storage device, wherein the plurality of respective submission queues are mapped to at least a first virtual output queue and a second virtual output queue of the NVMe storage driver, the first virtual output queue being mapped to a first completion queue of the NVMe storage driver, and the second virtual output queue being mapped to a second completion queue of the NVMe storage driver; Receiving, at the network infrastructure device, a second network transmission, the second network transmission being associated with a second data transfer operation using the NVMe-oF protocol; and Controlling the enqueuing and processing of the flow of the second network transmission using an egress queue congestion accounting and a direct rate control (RC) feedback signal between the network infrastructure device and a source node receiving at least one of the first network transmission or the second network transmission.

2. The method according to claim 1, wherein the direct RC feedback signal is a hardware-based signal.

3. The method according to claim 2, wherein the hardware-based signal is implemented using hardware logic implemented within the internal switching fabric.

4. The method according to claim 1, wherein the network infrastructure device is a network switch.

5. The method according to claim 1, wherein the network infrastructure device is a high-availability network switch.

6. The method according to claim 5, wherein the high-availability network switch provides high-availability access to a plurality of integrated solid-state drives including the NVMe storage device.

7. The method according to claim 1, wherein the first network transmission is received at a first port of the network infrastructure device, and the second network transmission is received at a second port of the network infrastructure device different from the first port.

8. The method according to claim 1, wherein the plurality of corresponding submission queues includes N submission queues, N being an integer greater than 1; the first completion queue and the second completion queue are part of a plurality of completion queues, the plurality of completion queues including N completion queues; the plurality of corresponding submission queues and the plurality of completion queues are correspondingly arranged in N NVMe queue pairs, each of the N NVMe queue pairs including a submission queue in the plurality of corresponding submission queues and a completion queue in the corresponding plurality of completion queues; and the N NVMe queue pairs are mapped to a single egress output interfacing with the NVMe storage device, and the first virtual output queue and the second virtual output queue are mapped to the single egress output.

9. A network switch, comprising: A first input port for receiving a first plurality of submission commands from a first remote host, the first plurality of submission commands being based on the Non-Volatile Memory Express (NVMe) protocol; A first plurality of submission queues in which the first plurality of submission commands are queued, the first plurality of submission queues being mapped to a single egress output via an internal switching fabric and a virtual output queue to provide data transfer operations using an integrated NVMe-based storage device; A first completion queue and a second completion queue, each completion queue being associated with a corresponding one of the first plurality of submission queues, wherein a plurality of response commands are queued for transmission to the first remote host using the NVMe protocol; And An NVMe storage driver for interfacing the first plurality of submission commands from the first plurality of submission queues to the integrated NVMe-based storage device via the internal switching fabric, wherein the control flow for the first plurality of submission queues includes direct rate control (RC) feedback between the single egress output and the first plurality of submission queues.

10. The network switch according to claim 9, wherein the NVMe storage driver is configured to determine whether at least one of the first completion queue and the second completion queue is available to control the flow of submission commands based on a signal from the direct RC feedback.

11. The network switch according to claim 10, wherein the NVMe storage driver controls the flow of the submission commands by: Obtaining a submission command from the first virtual output queue in response to determining that the first completion queue is available; or Obtaining the submission command from the second virtual output queue in response to determining that the second completion queue is available.

12. The network switch according to claim 11, wherein the NVMe storage driver is configured to: Dequeue the submission command obtained from the first virtual output queue or the second virtual output queue; and Execute the submission command by performing a data transfer operation using the integrated NVMe-based storage device.

13. The network switch according to claim 12, wherein the NVMe storage driver is configured to: In response to determining that the submission command obtained from the first virtual output queue has been processed, enqueue a response command in the first completion queue in response to the submission command; or In response to determining that the submission command obtained from the second virtual output queue has been processed, enqueue the response command in the second completion queue.

14. The network switch according to claim 10, wherein: Information from the first completion queue or the second completion queue is provided to the remote host via a response network message using the NVMe protocol.

15. The network switch according to claim 9, further comprising: A second input port for receiving a second plurality of submission commands from a second remote NVMe host; A third output port for transmitting response commands to the second remote NVMe host; A second plurality of submission queues in which the second plurality of submission commands received via the second input port are enqueued; A third virtual output queue associated with the second plurality of submission queues and mapped to the first completion queue; And A fourth virtual output queue associated with the second plurality of submission queues and mapped to the second completion queue, wherein the NVMe storage driver is configured to control the flow of the second plurality of submission commands using egress queue congestion accounting and direct RC feedback associated with a respective one of the first completion queue and the second completion queue.

16. The network switch according to claim 9, further comprising: A high availability (HA) configuration of a plurality of integrated NVMe-based storage devices, wherein multiple pairs of the plurality of integrated NVMe-based storage devices provide redundancy for respective members of each pair.

17. The network switch according to claim 16, wherein at least one pair of the multiple pairs of the plurality of integrated NVMe-based storage devices includes a first SSD on a first line card of the network switch and a second SSD on a second line card of the network switch, and the first line card is inserted into a different card slot of the network switch than the second line card.

18. A non-transitory computer-readable medium comprising computer-executable instructions stored thereon that, when executed by a processor in a network switch, cause the processor to: Receive, at a network infrastructure device, a first network transmission provided from a remote application communicatively coupled to the network infrastructure device via a network, the first network transmission being associated with a data transfer operation using the non-volatile memory express over fabric (NVMe-oF) protocol; Enqueue the first network transfer into one or more first submission commands for providing to an internal switching fabric of the network infrastructure device, the one or more first submission commands being associated with a plurality of respective submission queues of an NVMe storage driver interfacing with the internal switching fabric and a non-volatile memory express (NVMe) storage device, wherein the plurality of respective submission queues are included in a first virtual output queue and a second virtual output queue of the NVMe storage driver, the first virtual output queue being mapped to a first completion queue of the NVMe storage driver, and the second virtual output queue being mapped to a second completion queue of the NVMe storage driver; Receive a second network transfer at the network infrastructure device, the second network transfer being associated with a data transfer operation using the NVMe-oF protocol; and Use egress queue congestion accounting and a direct rate control (RC) feedback signal between the source node receiving at least one of the first network transfer or the second network transfer to control the flow of the second network transfer to be enqueued into the one or more submission commands.

19. The non-transitory computer-readable medium according to claim 18, wherein the direct RC feedback signal is a hardware-based signal.

20. The non-transitory computer-readable medium according to claim 19, wherein the hardware-based signal is implemented using hardware logic implemented within an internal switching fabric of the network switch.

Citation Information

Patent Citations

  • METHOD AND APPARATUS TO ENABLE INDIVIDUAL NON VOLATLE MEMORY EXPRESS (NVMe) INPUT / OUTPUT (IO) QUEUES ON DIFFERING NETWORK ADDRESSES OF AN NVMe CONTROLLER

    CN108351813A

  • Processing Method of Data Redundancy and Computer System Thereof

    US20190155772A1