Non-volatile memory switch with host isolation

By implementing multi-host isolation logic in NVM switches, detecting and isolating wrong hosts, the transaction delay or blocking problems caused by host failures in dual-active multi-host configurations are solved, and stable and efficient NVM device and host communication is achieved.

CN110825555BActive Publication Date: 2025-05-23MARVELL WORLD TRADE LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN201910727578.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-08-05
Filing Date
2019-08-07
Publication Date
2025-05-23
Estimated Expiration
2039-08-07

AI Technical Summary

Technical Problem

In dual active multi-host configurations, problems such as host failure or clock loss may cause transaction delays or blockages between other hosts and NVM devices, and it is difficult for the prior art to effectively isolate and deal with these issues.

Method used

An NVM switch is designed to detect and isolate the error host by implementing multi-host isolation logic in the switch, clearing its running commands, flushing its data, and ensuring the correct communication of error reports.

Benefits of technology

It realizes effective isolation and handling of host failures or clock loss in dual active multi-host configuration, ensuring stable communication and data integrity between NVM devices and hosts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN110825555B_ABST
    Figure CN110825555B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to a non-volatile memory switch with host isolation. The NVM switch has been designed to allow multiple hosts to simultaneously and independently access a single-port NVM device. While this active-active multi-host usage configuration allows for a variety of uses of low-cost single-port NVM devices, a problem with one of the hosts may delay or prevent transactions between other hosts and the NVM device. Although the logic of the switch is shared across hosts, the NVM switch includes logic to isolate the activities of multiple hosts. When the switch detects a problem with one host (the "error host"), the switch clears the running commands of the error host and flushes the data of the error host. Similarly, the NVM switch ensures the correct communication of error reports from attached NVM devices to multiple hosts.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This disclosure claims the benefit of priority to U.S. Provisional Application Serial No. 62 / 715,713, filed on August 7, 2018, and entitled “NVMe Protocol Switch Host Isolation in Active-Active Configuration,” the contents of which are incorporated herein by reference in their entirety. Background Art

[0003] The present disclosure relates generally to the field of computer architecture, and more particularly to input / output devices and operations.

[0004] High-performance computing environments are increasingly using non-volatile memory (NVM), such as flash memory, for storage solutions. In addition to traditional storage device interfaces optimized for rotating media technologies, host controller interfaces optimized for NVM are also being used. The NVM Express (NVMe) specification is an extensible host controller interface specification for NVM that leverages the Peripheral Component Interconnect Express (PCI-Express). Architecture. BRIEF DESCRIPTION OF THE DRAWINGS

[0005] Embodiments of the present disclosure may be better understood by referring to the accompanying drawings.

[0006] Figure 1 A block diagram of an example system including an NVM switch having logic and program code for multi-host to single-port host isolation is depicted.

[0007] Figure 2 is an example diagram of NVM switch logic to isolate hosts in an active-active configuration with a single-port NVM device.

[0008] Figure 3 is a flow chart of example operations for isolating an erroneous completion status detected in a host response.

[0009] Figure 4 is a flow chart of example operations for a NVM switch to isolate host data errors detected by the switch.

[0010] Figure 5 is a flow chart of example operations for a NVM switch to isolate data errors detected in a read completion packet.

[0011] Figure 6 is a flow diagram of example operations for an NVM switch to isolate a link down event from one of a plurality of hosts.

[0012] Figure 7is a flow diagram of example operations for an NVM switch to isolate a reset from one of a plurality of connected hosts.

[0013] Figure 8 is a flow chart of example operations for a NVM switch to propagate error reports across hosts. DETAILED DESCRIPTION

[0014] The following description includes example systems, methods, techniques, and program flows that embody various aspects of the present disclosure. However, it should be understood that the present disclosure can be practiced without these specific details. For example, the present disclosure relates to PCIe in an illustrative example. Various aspects of the present disclosure can also be applied to another similar interconnect architecture or specification for highly scalable high-speed communication with non-volatile memory or solid-state devices. In other cases, well-known instruction instances, protocols, structures, and techniques are not shown in detail to avoid confusing the description.

[0015] Overview

[0016] Dual-port NVM devices (e.g., solid-state drives (SSDs)) allow multiple hosts to access flash memory at a low cost. NVM switches have been designed to allow multiple hosts to simultaneously and independently access single-port NVM devices. While this active-active, multi-host usage configuration allows for various uses of low-cost single-port NVM devices (e.g., storage virtualization, redundant array of independent disks (RAID) functions, etc.), a problem with one of the hosts (e.g., failure, loss of reference clock, etc.) can delay or prevent transactions between other hosts and the NVM device. Although the logic of the switch is shared between the hosts, the NVM switch includes logic to isolate the activities of multiple hosts. When the switch detects a problem with one host (the "error host"), the switch clears the running commands of the error host and flushes the data of the error host. Similarly, the NVM switch ensures the correct communication of error reports from the connected NVM devices to multiple hosts.

[0017] Example diagram

[0018] Figure 1 A block diagram of an example system is depicted that includes an NVM switch with logic and program code for multi-host to single-port host isolation. An example system (e.g., a server rack) includes a backplane interconnect 117. Multiple hosts including host 119 and host 121 are connected to the backplane interconnect 117. NVM package 102 is also connected to the backplane interconnect. Hosts 119, 121, and NVM package 102 are all connected to the backplane interconnect 117 via a PCIe connector.

[0019] NVM package 102 includes NVM switch 101, single-port NVM device 120, and another single-port NVM device 122. NVM devices can be solid-state devices of various configurations. Figure 1 , NVM device 120 is depicted as having NVM controller 103, addressing logic 107, and flash memories 109A-109D.

[0020] The NVM switch 101 facilitates the hosts 119, 121 simultaneously and independently with the NVM devices 120, 122. The hosts 119, 121 maintain command queues and completion queues in their local memory. The hosts 119, 121 transmit messages and requests (e.g., doorbell messages) to the NVM devices 120, 122 via the NVM switch. The NVM controller 103 retrieves commands from the command queues of the hosts 119, 121 via the NVM switch 101. In response to the messages from the hosts 119, 121. The NVM controller 103 writes completion packets to the completion queues of the hosts 119, 121 via the NVM switch 101. Since each of the NVM devices 120, 122 is a single port, the NVM switch 101 presents a single requester to each of the NVM devices 120, 122. The NVM switch 101 includes logic to route packets appropriately to the hosts 119, 121 and isolate errors between the connected hosts.

[0021] Figure 2 2 is an example diagram of NVM switch logic to isolate hosts in a dual-active configuration with a single-port NVM device. NVM switch 200 includes an interconnect interface 207 and a non-volatile memory device interface 215. These interfaces can be the same but can be different. For example, interface 207 can be a PCIe interface with a greater number of lanes than interface 215. This example diagram assumes that the host (root complex from the perspective of switch 200) is linked via interconnect interface 207 and the NVM memory device is attached to switch 200 through interface 215. Switch 200 includes switch configuration and management logic 203, transaction management logic 201, and direct data path logic 205. The term "logic" refers to a circuit arrangement for implementing a task or function. For example, the logic for determining whether a value matches can be a circuit arrangement that uses a mutually exclusive NOR gate and an AND gate for an equal comparison. Configuration and management logic 203 guides management and configuration commands to attached NVM devices. Direct data path logic 205 allows read-completed writes to traverse switch 200 with little or no delay. Transaction management logic 201 prevents hosts from affecting each other. This includes at least isolating transactions that affect the performance of an event to the host where the event occurred (the "faulty host"), propagating error reports according to host-specific error reporting settings, and monitoring errors in direct data path logic 205.

[0022] Transaction management logic 201 includes registers and logic to facilitate functions for preventing events from one host from affecting the transactions of another host. Transaction management logic includes queue 202, timing source 231, host identifier logic 209, reservation logic 211, error isolation logic 213, queue 204 and queue 206. Queues 202, 204, 206 can be 32-bit registers, 64-bit registers, or different types of memory elements suitable for the physical space available for switch 200. Queue 202 stores incoming packets from the host. Examples of incoming data packets include doorbell messages, commands to read from host memory, and completion responses. Since switch 200 accommodates multiple hosts, reservation logic 211 reserves different areas of attached memory or backend devices to hosts to prevent hosts from overwriting each other. When a host establishes a connection with a backend device via switch 200, reservation logic 211 creates and maintains a mapping to the reserved storage space of each host. Reservation logic 211 can utilize available private namespace functions to achieve reservation. Another responsibility of the switch 200 is to present a single host to the backend device, because the backend device is a single port. This hides multiple hosts on the other end of the switch 200 from the backend device. The host identifier logic 209 and queues 204, 206 are used to ensure consistency of communication between the backend device and the host, even though the backend device is presented as a single host. The host identifier logic 209 associates the first host identifier with the queue 204 and the second host identifier with the queue 206. The implementation can add additional queues based on the number of hosts connected to the backend device through the NVM switch. The host identifier logic 209 copies the subfield value from the header of the input read type packet to one of the queues 204, 206 corresponding to the detected host identifier (e.g., requester identifier or node identifier). These copied values ​​will be used to determine which host is the correct requester. Using the reserved space, the host identifier logic 209 can copy the length, address, and sort tag fields to match the subsequent read completion packets that write the data returned in response to the read type packet. The host identifier logic 209 then resets the host identifier in the incoming packet to conform to the expected host identifier of the backend device (eg, root complex 0) before allowing the incoming packet to flow to the backend device.

[0023] When the backend device returns a read completion packet, the backend device writes the completion packet to the requester through the switch. The read completion packet will have a host identifier reset by the switch 200. When the backend device writes the completion packet to the completion write queue 221 of the direct data path logic 205, the host identifier 209 determines which of the queues 204, 206 has an entry that matches at least the stored fields (e.g., length, sort tag, address).

[0024] By retaining the host / requester at the switch, the error isolation logic 213 can prevent host events from affecting each other. The error isolation logic 213 can cause appropriate packets to be cleared from the queue 202 based on the detection of a problem event for one of the hosts. The error isolation logic 213 can also clear the completion packets corresponding to the failed or disconnected host from the direct data path logic 205. In addition, the error isolation logic 213 can switch the NVM switch 200 to use the internal timing source 231 in response to detecting the loss of the clock reference of one of the adapters. The switch 200 switches to the timing source 231 for processing and communicating packets from the adapter that has lost the reference clock.

[0025] Although Figure 2 Discrete logic blocks are depicted, but the different blocks are not necessarily physical boundaries of the microchip. For example, an NVM switch may include a processor that is logically part of multiple of the depicted logic blocks.

[0026] Figures 3 to 8 The flowchart in depicts example operations related to handling error conditions to maintain host isolation. The description refers to a switch as performing the example operations. The switch performs at least some of the operations according to program instructions (e.g., firmware) stored on the NVM switch.

[0027] Figure 3 1 is a flow chart of example operations for isolating an erroneous completion status detected in a host response. After a backend device requests to read data from a host memory, the host provides a packet with read completion data to the backend device. The packet includes a field for a completion status that can indicate an erroneous completion status. The NVM switch reads the data packet before transmitting the completion data to the backend device.

[0028] In block 301, the switch detects an error code in the completion status of the host response to a read from the backend device. The error code may indicate a completion-based error, a poisoned payload notification, and an internal parity error or an error correction code error. The switch may compare the bit at the position corresponding to the completion status with a predefined error code or look up the completion status value in a completion status table.

[0029] At block 303, the switch determines whether the completion status indicates a completion-based error. Examples of completion-based errors may be Completer Abort (CA), Unsupported Request (UR), and Completion Timeout.

[0030] If the error code in the Completion Status field is completion based, the switch modifies the Completion Status in the Host Response at block 305. The switch changes the error code in the Completion Status to indicate a completion with a Completer Abort (CA) before allowing the Host Response to be sent to the backend device identified in the Host Response.

[0031] At block 307, the switch determines whether the completion status field indicates a poisoned payload. If the completion status field indicates a poisoned payload, control flows to block 311. At block 311, the switch transmits the host response with the poisoned payload indication to the backend device. Otherwise, control flows to block 309.

[0032] If the completion status indicates that an internal parity error or ECC error was detected at the host, then at block 309, the switch discards the corrupted data and triggers a data path data integrity error mode in the switch. In this mode, all requests from the backend device to a particular host are dropped, and read requests are completed with an error. For example, the switch may set a value in a register associated with a port of the backend device. When the switch receives a request from the backend device, the switch determines the corresponding host identifier. If the host identifier matches the value associated with the port of the backend device having the determined host, the request is discarded.

[0033] Figure 4 is a flow chart of example operations of an NVM switch to isolate host data errors detected by the switch. The switch also supports data integrity in an attempt to avoid inserting delays into the NVM transaction path. If the data integrity error mode is activated, the switch evaluates the data parity of the read packets and the read completion packets traversing the switch.

[0034] The switch detects a data parity error in the completion data of the host response at block 401. The detection may be detecting a bit set by a data link layer component.

[0035] The switch modifies the read packet host response to indicate the poisoned payload in the completion status field at block 403. The switch may use the stored poisoned payload code to propagate the parity error detection.

[0036] At block 405, the switch transmits the modified host response to the backend device.

[0037] Figure 5 1 is a flow chart of example operations for an NVM switch to isolate data errors detected in a write transaction. The direct data path logic of the NVM switch may be configured to check for parity or ECC errors in data in a write transaction from a backend device.

[0038] The switch detects parity errors or uncorrectable ECC errors in write data from a backend device at block 501. The switch may check the data link layer bits to detect internal parity errors or uncorrectable ECC errors.

[0039] In block 503, the switch discards the write data. The discarding of the write data is discarding the entire write transaction issued by the backend device.

[0040] At block 505, the switch triggers a data path data integrity error mode. In this mode, all requests from backend devices to a particular host will be dropped, and read requests will be completed with an error.

[0041] Figure 6 6 is a flow chart of an example operation of an NVM switch to isolate a link disconnection event from one of a plurality of hosts. In box 601, the switch detects a link disconnection of a host previously linked to the switch. A data link layer component or a component of a PCIe core detects the link disconnection. The link disconnection indication includes an identifier of the corresponding host. In box 603, the switch triggers a transaction flush of all running transactions targeting the host with the link disconnected from an NVM device attached to the switch. In box 604, the switch stops service to the link disconnected port. In order to pause traffic to the link disconnected port, the switch aborts all outstanding commands associated with the host to be processed in the NVM device. In box 605, the switch initiates a reset of a component corresponding to the host with the link disconnected. In box 607, the switch reestablishes or attempts to reestablish a link with the host.

[0042] Figure 7 Flowchart of example operations for isolating a reset of an NVM switch from one of a plurality of connected hosts. The switch processes reset events on a link-by-link basis. Reset events may be warm resets (PERST), hot resets, disabled links, and functional level resets (FLR). In block 701, the switch detects a reset command from a host to an endpoint (EP) core implementing a lower layer protocol (e.g., data link layer, physical layer). The switch detects the reset because the reset signal or command generates an interrupt to the switch. In block 703, the switch discards unissued submission queue entries received via the EP core. In block 705, the switch aborts issued but uncompleted commands associated with the EP core. The switch determines that there is a running read command lacking a corresponding read completion in a write queue sent to the host. The switch then sends an abort command for each of these incomplete transactions to the corresponding backend device. In block 707, the switch triggers a transaction flush to the EP core. The switch asserts a signal to cause the EP core to clear data packets traversing the EP core. In block 709, the switch deletes the queue associated with the EP core being reset. In block 711, the switch initiates a reset of the EP core.

[0043] Figure 8 8 is a flow chart of example operations for an NVM switch to propagate error reports across hosts. Error reporting settings include basic error reporting settings and advanced error reporting (AER) settings set in registers of back-end devices. In box 801, the switch detects error reporting settings from a host during enumeration and configuration operations. In box 803, the switch determines whether the error reporting settings are different between hosts. If the error reporting settings are the same between hosts, the switch transmits an instance of the error reporting settings to the attached back-end device in box 804. If the error reporting settings are different, the switch stores each instance of the error reporting settings. In box 805, the host associates a corresponding host identifier with each instance of the error reporting settings. This will be used to ensure host-specific compliance of error reports from back-end devices. In box 807, the switch transmits the instance of the error reporting settings provided by the host with more settings to the attached back-end device.

[0044] At some later point, the switch may detect an error report from a backend device in block 809. In block 810, the switch determines whether there is a stored instance of an error reporting setting that the error report does not comply with. For example, a stored instance of an error reporting setting may indicate that the error report should be a completed state, while the detected error report is an error message. In block 813, the switch transmits the error report to all hosts because the error report complies with the instance of the error reporting setting. If an inconsistency is detected for one instance of the error reporting setting, then in block 811 the switch communicates the error report to the hosts associated with the instance of the error reporting setting, and the error report is consistent with the error reporting setting. In block 815, the switch derives and communicates error reports for other instances of the error reporting setting. The switch extracts information from the detected error report and generates an error report with the information based on the instance of the error reporting setting.

[0045] The flowcharts are provided to aid understanding of the description and are not intended to limit the scope of the claims. The flowcharts depict example operations that may vary within the scope of the claims. Additional operations may be performed; fewer operations may be performed; operations may be performed in parallel; and operations may be performed in a different order. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, may be implemented by program code. The program code may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable machine or device.

[0046] As will be appreciated, aspects of the present disclosure may be embodied as systems, methods, or program codes / instructions stored in one or more machine-readable media. Thus, aspects may take the form of hardware, software (including firmware, resident software, microcode, etc.), or a combination of software and hardware aspects, which may all be generally referred to herein as "circuits," "modules," or "systems." Functionality presented as a single module / unit in the example illustrations may be organized differently depending on any of the platform (operating system and / or hardware), application ecosystem, interface, programmer preference, programming language, administrator preference, etc.

[0047] Any combination of one or more machine-readable media may be utilized. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable storage medium may be, for example but not limited to, a system, apparatus, or device that employs any one or a combination of electronic, magnetic, optical, electromagnetic, infrared, or semiconductor technologies to store program code. More specific examples (a non-exhaustive list) of machine-readable storage media would include the following: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a machine-readable storage medium may be any tangible medium that may contain or store a program used by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable storage medium is not a machine-readable signal medium.

[0048] A machine-readable signal medium may include a propagated data signal with machine-readable program code embodied therein, for example, in baseband or as part of a carrier wave. Such a propagated signal may take any of a variety of forms, including but not limited to electromagnetic, optical, or any suitable combination thereof. A machine-readable signal medium may be any machine-readable medium that is not a machine-readable storage medium and that can communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0049] The program code / instructions may also be stored in a machine-readable medium which may direct a machine to operate in a specific manner such that the instructions stored in the machine-readable medium produce an article of manufacture including instructions for implementing the functions / actions specified in the flowchart and / or block diagram blocks or blocks.

[0050] Multiple instances may be provided as a single instance for the components, operations, or structures described herein. Finally, the boundaries between the various components, operations, and data stores are arbitrary to some extent, and specific operations are shown in the context of a specific illustrative configuration. Other functional allocations are contemplated and may fall within the scope of the present disclosure. Typically, the structures and functions presented as separate components in the example configuration may be implemented as combined structures or components. Similarly, the structures and functions presented as single components may be implemented as separate components. These and other variations, modifications, additions, and improvements may fall within the scope of the present disclosure.

[0051] the term

[0052] As used herein, unless expressly stated otherwise, the term "or" is inclusive. Thus, the phrase "at least one of A, B, or C" is satisfied by any element from the set {A, B, C} or any combination thereof, including multiples of any element.

Claims

1. A non-volatile memory switch, include: Multiple interconnection interfaces; processor; Multiple queues; A machine-readable medium having program code stored thereon, the program code being executable by the processor to cause the switch to: reserving, to different root complexes, different memory spaces on a set of one or more single-ported non-volatile memory devices accessible via at least a second interconnect interface of the plurality of interconnect interfaces; presenting the different root complexes as a single root complex to the set of one or more single-ported non-volatile memory devices, including modifying identifiers within memory transaction messages received by the non-volatile memory switch from the different root complexes to a common value indicative of the single root complex, each identifier indicating a corresponding root complex from which a memory transaction message having the identifier was received; maintaining an association of memory transactions with corresponding ones of the different root complexes using at least a subset of the plurality of queues; as well as Based at least in part on the maintained association, an error condition in one of the different root complexes is prevented from affecting memory transactions in another of the different root complexes.

2. The non-volatile memory switch of claim 1 , wherein the program code that maintains the association of memory transactions comprises program code executable by the processor to cause the non-volatile memory switch to: associating a first queue of the plurality of queues with a first root complex of the root complexes, and associating a second queue of the plurality of queues with a second root complex of the root complexes; for each read packet of the memory transaction identifying the first root complex as a requestor, storing information from a header of the read packet that can be used to identify corresponding read completion data in the first queue; as well as For each read packet of the memory transaction identifying the second root complex as a requestor, information from a header of the read packet is stored in the second queue.

3. The non-volatile memory switch of claim 2, wherein the program code to prevent an error condition of one of the different root complexes from affecting transactions of another of the different root complexes comprises program code to: Based on detecting an error condition from one of the first root complex and the second root complex, selecting a read completion packet in the nonvolatile memory switch associated with the root complex of the error condition; and The selected read completion packet is flushed from the non-volatile memory switch.

4. The non-volatile memory switch of claim 1 , wherein the program code to prevent an error condition of one of the different root complexes from affecting transactions of another of the different root complexes comprises program code executable by the processor to cause the non-volatile memory switch to: A root complex response to a read request for a non-volatile memory device from among the set of one or more single-port non-volatile memory devices indicates a determination of an error based on a completion, and sets a completion status of the root complex response to indicate a completion with a completer abort status, wherein the root complex response is received via a first interconnect interface among the plurality of interconnect interfaces; and Convey the root complex response having the completion status set to indicate the completion with the completer abort status to the one non-volatile memory device via the second interconnect interface.

5. The non-volatile memory switch according to claim 1, wherein the program code for preventing an error condition of one root complex among the different root complexes from affecting transactions of another root complex among the different root complexes includes program code executable by the processor to cause the non-volatile memory switch to: Based on detection of a data parity error in completion data of a root complex response for a non-volatile memory device from among the set of one or more single-port non-volatile memory devices, modify the root complex response to indicate a poisoned payload, wherein the root complex response is received via a first interconnect interface among the plurality of interconnect interfaces; and Convey the modified root complex response to the one non-volatile memory device via the second interconnect interface.

6. The non-volatile memory switch according to claim 1, wherein the program code for preventing an error condition of one root complex among the different root complexes from affecting transactions of another root complex among the different root complexes includes program code executable by the processor to cause the non-volatile memory switch to: Based on detection of a link disconnect event corresponding to a first root complex among the root complexes, trigger a flush of in-flight transactions on the non-volatile memory switch from the set of one or more single-port non-volatile memory devices targeted at the link corresponding to the link disconnect event; Convey an abort command for each command that has been issued on the non-volatile memory switch to the set of one or more single-port non-volatile memory devices but has not been completed; Reset a link component of the non-volatile memory switch corresponding to the first root complex; And Re-establish a link with the first root complex.

7. The non-volatile memory switch according to claim 1, wherein the program code for preventing an error condition of one root complex among the different root complexes from affecting transactions of another root complex among the different root complexes includes program code executable by the processor to cause the non-volatile memory switch to: Based on detection of a reset command from a first root complex to a first endpoint core, Discard read commands to be issued from the non-volatile memory switch from the first root complex; communicating an abort command for each command on the non-volatile memory switch that has been issued to the set of one or more single-port non-volatile memory devices but has not yet completed; triggering a transaction flush on the first endpoint core; deleting the command and submission queues of the first endpoint core; as well as A reset of the first endpoint core is initiated.

8. The nonvolatile memory switch of claim 1 , wherein the program code to present the different root complexes to a nonvolatile memory device as a single root complex comprises program code executable by the processor to cause the nonvolatile memory switch to modify an identifier in a read command received by the nonvolatile memory switch from the different root complexes to be the common value indicative of the single root complex.

9. The non-volatile memory switch of claim 1, wherein the plurality of interconnect interfaces comprises a Peripheral Component Interconnect Express interface.

10. The non-volatile memory switch of claim 1, wherein the memory transaction complies with the Non-Volatile Memory Express specification.

11. The nonvolatile memory switch of claim 1 , further comprising an internal timing source, wherein the program code to prevent an error condition of one of the different root complexes from affecting transactions of another of the different root complexes comprises program code executable by the processor to cause the nonvolatile memory switch to switch processing of transactions of a first one of the root complexes to use the internal timing source as a reference clock based on detection of a loss of a reference clock for the first root complex.

12. The non-volatile memory switch of claim 1, wherein the program code that presents the different root complexes as the single root complex to the set of one or more single-port non-volatile memory devices include: Program code executable by the processor to cause the nonvolatile memory switch to use the maintained association to determine a root complex of the different root complexes that corresponds to a memory transaction message received by the nonvolatile memory switch from the one or more single-port nonvolatile memory devices.

13. A non-volatile memory NVM package, include: Peripheral Component Interface Express PCIe interface; Multiple single-port NVM devices; as well as a NVM switch communicatively coupled between the PCIe interface and the plurality of single-port NVM devices, wherein the NVM switch comprises a processor and a machine-readable medium having program code stored thereon, the program code being executable by the processor to cause the NVM switch to: reserving different memory spaces on the plurality of single-port NVM devices for different root complexes; presenting the different root complexes as a single root complex to the plurality of single-port non-volatile memory devices, including modifying identifiers within memory transaction messages received by the NVM switch from the different root complexes to a common value indicative of the single root complex, each identifier indicating a corresponding root complex from which a memory transaction message having the identifier was received; as well as An error condition in one of the different root complexes is prevented from affecting transactions in another of the different root complexes.

14. The NVM package of claim 13, wherein the program code further comprises program code executable by the processor to cause the NVM switch to maintain an association of memory transactions with corresponding ones of the different root complexes without requesting an identifier.

15. The NVM package of claim 14, wherein the program code to prevent an error condition in one of the different root complexes from affecting transactions in another of the different root complexes comprises program code to: Based on detecting an error condition from one of the root complexes, selecting a read completion packet in the NVM switch associated with the root complex of the error condition; and The selected read completion packet is flushed from the NVM switch.

16. The NVM package of claim 14, wherein the program code presents the different root complexes as the single root complex to the plurality of single-port non-volatile memory devices include: Program code executable by the processor to cause the NVM switch to use the maintained associations to determine a root complex of the different root complexes that corresponds to a memory transaction message received by the nonvolatile memory switch from the plurality of single-port NVM devices.

17. The NVM package of claim 13, wherein the program code to prevent an error condition of one of the different root complexes from affecting transactions of another of the different root complexes comprises program code executable by the processor to cause the NVM switch to: Based on detecting a reset command from the first root complex to the first endpoint core, discarding, from the first root complex, a read command on the NVM switch to be issued from the NVM switch; communicating an abort command for each command on the NVM switch that has been issued to at least one of the plurality of NVM devices but has not yet completed; triggering a transaction flush to the first endpoint core; deleting the command and submission queues of the first endpoint core; as well as A reset of the first endpoint core is initiated.

18. The NVM package of claim 13, wherein the NVM switch further comprises an internal timing source, wherein the program code to prevent an error condition of one of the different root complexes from affecting transactions of another of the different root complexes comprises program code executable by the processor to cause the NVM switch to switch processing of transactions of a first of the root complexes to use the internal timing source as a reference clock based on detection of a loss of a reference clock for the first root complex.

19. A system, include: Backplane interconnects; a plurality of host devices, each of the host devices comprising a corresponding root complex and a corresponding interconnect interface, the interconnect interface connecting a corresponding one of the host devices to the backplane interconnect; as well as a non-volatile memory (NVM) package comprising a plurality of single-port non-volatile memory (NVM) devices, an interconnect interface connecting the NVM package to the backplane interconnect, and a NVM switch communicatively coupled between the plurality of single-port NVM devices and the interconnect interface of the NVM package, wherein the NVM switch comprises, A processor and a machine-readable medium having program code stored thereon, the program code being executable by the processor to cause the NVM switch to: presenting different root complexes of the plurality of host devices as a single root complex to the plurality of single-port NVM devices, including modifying identifiers within memory transaction messages received by the nonvolatile memory switch from the plurality of host devices to a common value indicative of the single root complex, each identifier indicating a corresponding root complex from which a memory transaction message having the identifier was received; as well as An error condition in one of the plurality of host devices is prevented from affecting transactions of another of the plurality of host devices.

20. The system of claim 19, wherein the program code to prevent an error condition in one of the plurality of host devices from affecting transactions in another of the plurality of host devices comprises program code to: based on detecting an error condition from one of the plurality of host devices, selecting a read completion packet in the NVM switch associated with the host device of the error condition; and The selected read completion packet is flushed from the NVM switch.

21. The system of claim 19, wherein the NVM switch further comprises an internal timing source, wherein the program code that prevents an error condition of one of the plurality of host devices from affecting transactions of another of the plurality of host devices comprises program code executable by the processor to cause the NVM switch to switch processing of transactions of a first of the plurality of host devices to use the internal timing source as a reference clock based on detection of a loss of a reference clock for the first host device.

22. The system of claim 19, wherein the program code comprises program code executable by the processor to cause the NVM switch to: maintaining an association of memory transactions with corresponding ones of the different root complexes; and A root complex of the different root complexes that corresponds to a memory transaction message received by the nonvolatile memory switch from the plurality of single-port NVM devices is determined using the maintained associations.

Citation Information

Patent Citations

  • Message transmission method, device, system and storage medium realizing pcie switching network

    CN103098428A

  • I / O system, downstream PCI express bridge, interface sharing method, and program

    US20120110233A1

  • Peripheral component interconnect express (PCIE) distributed non- transparent bridging designed for scalability,networking and IO sharing enabling the creation of complex architectures.

    US20150261709A1

  • Header parity error handling

    US20160147592A1