Systems, methods, and media for out-of-order input / output direct memory access operations

By leveraging the coordinated operation of the hardware processor's cache coherence control unit and the CXL.mem device in out-of-order input/output direct memory access operations between PCIe devices and CXL.mem devices, the data transfer process is optimized, performance bottlenecks are resolved, and the efficiency of the computing link is improved.

CN122397008APending Publication Date: 2026-07-14SK HYNIX NAND PRODUCT SOLUTIONS CORP

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SK HYNIX NAND PRODUCT SOLUTIONS CORP
Filing Date
2024-12-05
Publication Date
2026-07-14

Smart Images

  • Figure CN122397008A_ABST
    Figure CN122397008A_ABST
Patent Text Reader

Abstract

Mechanisms are provided for out-of-order input / output direct memory access operations, including: issuing, using a hardware processor, a return invalidate snoop request to a cache coherency control unit of a host processor; and issuing an out-of-order input / output direct memory access operation request to a compute express link memory device. In some of these mechanisms, the out-of-order input / output direct memory access operation request is for a read operation. In some of these mechanisms, the mechanism further includes receiving a response to the out-of-order input / output direct memory access operation request, including data from the compute express link memory device. In some of these mechanisms, the data is updated in response to the return invalidate snoop request. In some of these mechanisms, the out-of-order input / output direct memory access operation request is for a write operation.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications This application claims the benefit of U.S. Patent Application No. 18 / 536,878, filed December 12, 2023, which is hereby incorporated herein by reference in its entirety. Background Technology

[0002] Peripheral Component Interconnect Fast (PCIe) devices frequently request out-of-order input / output (UIO) direct memory access (DMA) operations at the Compute Fast Link Memory (CXL.mem) device. During these operations, the PCIe device can read data from or write data to the CXL.mem device.

[0003] Improving the performance of the PCIe and CXL.mem devices involved in such operations is desirable.

[0004] Therefore, new mechanisms (including systems, methods, and media) for out-of-order input / output direct memory access operations are desirable. Summary of the Invention

[0005] According to some embodiments, mechanisms (including systems, methods, and media) are provided for out-of-order input / output direct memory access operations.

[0006] In some embodiments, a system for out-of-order input / output direct memory access (OI / O) operations is provided. The system includes: a memory; and a hardware processor coupled to the memory and configured to at least: issue a return invalid listener request to a cache coherence control unit of a host processor; and issue an OI / O direct memory access operation request to a compute fast link memory device. In some of these embodiments, the OI / O direct memory access operation request is used for a read operation. In some of these embodiments, the hardware processor is also configured to receive a response to the OI / O direct memory access operation request, including data from the compute fast link memory device. In some of these embodiments, data is updated in response to the return invalid listener request. In some of these embodiments, the OI / O direct memory access operation request is used for a write operation. In some of these embodiments, the OI / O direct memory access operation request includes data to be written to the compute fast link memory device. In some of these embodiments, a cache line is invalidated in response to the return invalid listener request.

[0007] In some embodiments, a method is provided for an out-of-order input / output direct memory access (OI / O) operation, the method comprising: issuing a return invalid listener request to a cache coherence control unit of a host processor using a hardware processor; and issuing an OI / O direct memory access operation request to a compute fast link memory device. In some of these embodiments, the OI / O direct memory access operation request is for a read operation. In some of these embodiments, the method further comprises receiving a response to the OI / O direct memory access operation request, including data from the compute fast link memory device. In some of these embodiments, data is updated in response to the return invalid listener request. In some of these embodiments, the OI / O direct memory access operation request is for a write operation. In some of these embodiments, the OI / O direct memory access operation request includes data to be written to the compute fast link memory device. In some of these embodiments, a cache line is invalidated in response to the return invalid listener request.

[0008] In some embodiments, a non-transitory computer-readable medium containing computer-executable instructions that, when executed by a processor, cause the processor to perform a method for an out-of-order input / output direct memory access operation, the method comprising: issuing a return invalidation listener request to a cache coherence control unit of a host processor; and issuing an out-of-order input / output direct memory access operation request to a compute fast link memory device. In some of these embodiments, the out-of-order input / output direct memory access operation request is for a read operation. In some of these embodiments, the method further includes receiving a response to the out-of-order input / output direct memory access operation request, including data from the compute fast link memory device. In some of these embodiments, data is updated in response to the return invalidation listener request. In some of these embodiments, the out-of-order input / output direct memory access operation request is for a write operation. In some of these embodiments, the out-of-order input / output direct memory access operation request includes data to be written to the compute fast link memory device. In some of these embodiments, a cache line is invalidated in response to the return invalidation listener request. Attached Figure Description

[0009] Figure 1 It is a block diagram of an example of hardware used according to some embodiments.

[0010] Figure 2 This is a flowchart of an example process for an out-of-order input / output (UIO) direct memory access (DMA) read operation from a CXL.mem device by a PCIe device, according to some embodiments.

[0011] Figure 3 This illustrates some embodiments. Figure 2 A block diagram illustrating an exemplary UIO DMA read operation.

[0012] Figure 4 This is a flowchart of an exemplary process for performing a UIO DMA write operation on a CXL.mem device by a PCIe device, according to some embodiments.

[0013] Figure 5 This illustrates some embodiments. Figure 4 A block diagram illustrating an exemplary UIO DMA write operation. Detailed Implementation

[0014] According to some embodiments, mechanisms (including systems, methods, and media) are provided for out-of-order input / output direct memory access operations.

[0015] Go to Figure 1 Example 100 of hardware is shown, which can be used according to some embodiments. As illustrated, hardware 100 includes a host processor 102, a compute fast link (CXL) switch 104, a PCIe device 106, a CXL memory (CXL.mem) device 108, and memory 124. In some embodiments, hardware 100 may also include any other suitable devices. In some embodiments, hardware 100 may be included in, part of, and / or connected to any suitable general-purpose or special-purpose computing device, such as a desktop computer, laptop computer, tablet computer, server, embedded device, sensor, Internet of Things device, and / or any other suitable device.

[0016] like Figure 1As shown, in some embodiments, host processor 102 may include any suitable number (including only one) of processor cores 110 and 112, and each processor core may include one or more caches (e.g., caches 114 and 116 for processor cores 110 and 112, respectively). Host processor 102 may be implemented using hardware processors such as microprocessors. In some embodiments, host processor may be configured (e.g., via executable code) to provide a cache coherence control unit 117 (which may be any suitable hardware and / or software for managing the cache coherence state of memory) and a root complex 120. In some embodiments, host processor 102 may also be configured to provide a memory controller (e.g., implemented as hardware for implementing a bus scheduling algorithm or a combination of hardware and microcode) that can be used to interface host processor to memory 124 (which may be implemented using dynamic random access memory, static random access memory, and / or any suitable volatile and / or non-volatile byte-addressable memory technology).

[0017] In some embodiments, the CXL switch can be any suitable switch compatible with the CXL standard and can be implemented using any suitable hardware, executable code, or a combination thereof.

[0018] In some embodiments, PCIe device 106 may be any suitable PCIe device, such as a solid-state drive (SSD) (which may be implemented using random access memory, read-only memory, flash memory, NAND memory (e.g., multi-plane NAND memory, 3D NAND, memory having any of the following memory densities: single-level cell (SLC), multi-level cell (MLC), three-level cell (TLC), four-level cell (QLC), five-level cell (PLC) and any suitable memory density with more than five bits per memory cell, and / or any other suitable memory), hard disk drive, optical media drive, general-purpose or special-purpose GPU, computing or data mobility accelerator, network interface card, or data processing unit.

[0019] In some embodiments, CXL.mem device 126 may be any suitable memory device compatible with the CXL standard. As shown, in some embodiments, CXL.mem device 126 may include any suitable memory, such as random access memory, read-only memory, flash memory, NAND memory (e.g., multi-plane NAND memory, 3D NAND, memory having any of the following memory densities: single-level cell (SLC), multi-level cell (MLC), three-level cell (TLC), four-level cell (QLC), five-level cell (PLC), and any suitable memory density greater than five bits per memory cell, and / or any other suitable memory), hard disk storage device, optical media, and / or any other suitable memory.

[0020] In some embodiments, the host processor 102, CXL switch 104, PCIe device 106, and CXL.mem device 106 may communicate using any suitable interface(s). For example, in some embodiments, the host processor 102, CXL switch 104, PCIe device 106, and CXL.mem device 106 may communicate using a PCIe interface.

[0021] Figure 2 The illustration shows a flowchart of an example 200 of a process for an out-of-order input / output (UIO) direct memory access (DMA) read operation performed by a PCIe device 106 from a CXL.mem device 108, according to some embodiments. In some embodiments, the PCIe device can use this operation to read the memory of the CXL.mem device.

[0022] As shown, process 200 includes subprocesses 210, 220, and 230, executed by PCIe device 106, cache coherence control unit 117, and CXL.mem device 108, respectively. Subprocess 210 includes 212, 214, 216, and 218; subprocess 220 includes 222, 224, 226, and 228; and subprocess 230 includes 232, 234, and 236.

[0023] As illustrated, process 200 begins at 212 of subprocess 210, where the PCIe device issues a Proxy Return Invalid Listen Request (Read) (PBISRR) for a given address range in the CXL.mem device. In some embodiments, the PBISRR is considered proxy-based because it is issued by the PCIe device rather than the CXL.mem device. In some embodiments, the PBISRR can include any suitable content, can have any suitable format, and can be published in any suitable manner. For example, in some embodiments, the PBISRR may indicate that it is for reading and that it is for a given address range.

[0024] Next, at 222 of subprocess 220, the cache consistency control unit receives the PBISRR for a given address range. In some embodiments, the PBISRR can be received in any suitable manner.

[0025] refer to Figure 3 According to some embodiments, an example of the PBISRR process is illustrated at 302. As shown, in some embodiments, PBISRR originates from the PCIe device, is sent through the CXL switch and the root complex, and is received by the cache coherence control unit.

[0026] Return to reference Figure 2 At subprocess 220, at point 224, the cache consistency control unit then collects data from the dirty cache lines(s) corresponding to a given address range. A cache line is considered dirty if its content is the only up-to-date copy of that data. The cache consistency control unit collects data from the dirty cache lines(s) from the system cache hierarchy so that it can use the collected data from the dirty cache lines(s) to update the corresponding data in the CXL.mem device. In some embodiments, data can be collected from any suitable number of cache lines (including only one), and data from cache lines can be collected in any suitable manner.

[0027] Then, at 226 of subprocess 220, the cache coherence control unit writes the data collected from the system cache hierarchy to the CXL.mem device. In some embodiments, data from any suitable number of dirty cache lines(s), including only one, can be written, and data from dirty cache lines(s) can be written in any suitable manner. For example, in some embodiments, data from dirty cache lines(s) can be written using normal storage semantics as specified in the CXL specification. In some embodiments, the cache coherence control unit also updates the state of the cache lines(s) to reflect whether they are shared (if the host processor maintains a copy) or invalid (if the cache lines(s) are evicted from the host processor).

[0028] Following step 226, at step 232 of subprocess 230, the CXL.mem device receives data from dirty cache lines from the cache coherence control unit. In some embodiments, the CXL.mem device may receive data from dirty cache lines from the cache coherence control unit in any suitable manner.

[0029] refer to Figure 3 According to some embodiments, an example of collecting data from dirty cache lines(s) and writing data from dirty cache lines(s) to the CXL.mem device is illustrated at position 304. Figure 3 As illustrated in the example, data from dirty cache lines(s) is collected from cache 116 (or any other suitable cache) of processor core 112 by the cache coherence control unit, and then written to memory 126 of CXL.mem device 108 by the cache coherence control unit. In some embodiments, when data from dirty cache lines(s) is written to the CXL.mem device, in some embodiments, the data from dirty cache lines(s) passes through the root complex and the CXL switch.

[0030] Return to reference Figure 2 Similarly, following 226, at 228 of subprocess 220, the cache consistency control unit indicates to the PCIe device that PBISRR is complete. In some embodiments, this indication can be made in any suitable manner and can include any suitable content. For example, in some embodiments, this indication can be made by generating a completion indicator with a label for the PCIe device that helps that device associate completion with the original request.

[0031] Then, at 214 of subprocess 210, the PCIe device receives an indication of PBISRR completion from the cache coherence control unit. In some embodiments, the PCIe device may receive the indication of PBISRR completion from the cache coherence control unit in any suitable manner.

[0032] Then, at 216 of subprocess 210, the PCIe device issues a UIO DMA read operation request to the CXL.mem device. In some embodiments, the UIO DMA read operation request can be made in any suitable manner and can include any suitable content. For example, in some embodiments, the UIO DMA read operation request can be made as specified in the CXL specification and can indicate a given range of addresses to be read.

[0033] Then, at 234 of subprocess 230, the CXL.mem device receives a UIO DMA read operation request. In some embodiments, the CXL.mem device may receive the UIO DMA read operation request in any suitable manner.

[0034] refer to Figure 3 According to some embodiments, an example of a UIODMA read operation request being issued from a PCIe device to a CXL.mem device is illustrated at 306. As shown, in some embodiments, the UIODMA read operation request passes through the CXL switch, but not through the root complex or cache coherence control unit.

[0035] Return to reference Figure 2 At 236 of subprocess 230, the CXL.mem device then responds to the UIO DMA read operation request with the current data. In some embodiments, the response to the UIO DMA read operation request can be made in any suitable manner and can include any suitable content. For example, in some embodiments, the response to the UIO DMA read operation request can be made as specified in the CXL specification.

[0036] Finally, at 218 of subprocess 210, the PCIe device receives a response to the UIO DMA read operation request, containing current data from the CXL.mem device. In some embodiments, the PCIe device may receive the response to the UIO DMA read operation request, containing current data from the CXL.mem device, in any suitable manner. In some embodiments, once received by the PCIe device, the current data may be used in any suitable manner and for any suitable purpose.

[0037] refer to Figure 3 According to some embodiments, an example of a response to a UIODMA read operation request from a PCIe device to a CXL.mem device is illustrated at 308. As shown, in some embodiments, the response to the UIODMA read operation request passes through the CXL switch, but not through the root complex or cache coherence control unit.

[0038] Figure 4 The illustration shows a flowchart of an example 400 of a process for a PCIe device 106 to perform an out-of-order input / output (UIO) direct memory access (DMA) write operation to a CXL.mem device 108, according to some embodiments. In some embodiments, the PCIe device can use this operation to write to the memory of the CXL.mem device.

[0039] As shown, process 400 includes subprocesses 410, 420, and 430, executed by PCIe device 106, cache coherence control unit 117, and CXL.mem device 108, respectively. Subprocess 410 includes 412, 414, 416, and 418; subprocess 420 includes 422, 424, and 426; and subprocess 430 includes 432, 434, 436, and 438.

[0040] As illustrated, process 400 begins at 412 of subprocess 410, where the PCIe device issues a proxy return invalid listener request (write) (PBISRW) for a given address range in the CXL.mem device. Because the return invalid listener request (write) is issued by the PCIe device rather than the CXL.mem device, the PBISRW is considered proxy-based. In some embodiments, the PBISRW may include any suitable content, may have any suitable format, and may be published in any suitable manner. For example, in some embodiments, the PBISRW may indicate that it is for writing and that it is for a given address range.

[0041] Next, at 422 of subprocess 420, the cache consistency control unit receives the PBISRW for a given address range. In some embodiments, the PBISRW can be received in any suitable manner.

[0042] refer to Figure 5 According to some embodiments, an example of the PBISRW process is illustrated at 502. As shown, in some embodiments, the PBISRW originates from a PCIe device, is sent through a CXL switch and the root complex, and is received by the cache coherence control unit.

[0043] Return to reference Figure 4 At 424 of subprocess 420, the cache consistency control unit then invalidates all copies of cache lines(s) corresponding to a given address range. In some embodiments, any suitable number of cache lines (including only one) can be invalidated, and cache lines can be invalidated with or without being written back to the CXL.mem device in any suitable manner.

[0044] refer to Figure 5 According to some embodiments, an example of an invalid cache line(s) is illustrated at 504.

[0045] Return to reference Figure 4Following step 424, at step 426 of subprocess 420, the cache consistency control unit indicates to the PCIe device that PBISRW is complete. In some embodiments, this indication can be made in any suitable manner and can include any suitable content. For example, in some embodiments, this indication can be made by generating a completion indicator with a label for the PCIe device, which helps that device to complete the task associated with the original request.

[0046] Then, at 414 of subprocess 410, the PCIe device receives an indication that PBISRW is complete from the cache coherence control unit. In some embodiments, the PCIe device may receive the indication that PBISRW is complete from the cache coherence control unit in any suitable manner.

[0047] Then, at 416 of subprocess 410, the PCIe device issues a UIO DMA write operation request to the CXL.mem device. In some embodiments, the UIO DMA write operation request can be made in any suitable manner and can include any suitable content. For example, in some embodiments, the UIO DMA read operation request can be made as specified in the CXL specification and can indicate a given address range to be written to and the data to be written to that address range.

[0048] Then, at 434 of subprocess 430, the CXL.mem device receives a UIO DMA write operation request. In some embodiments, the CXL.mem device may receive the UIO DMA write operation request in any suitable manner.

[0049] refer to Figure 5 According to some embodiments, an example of issuing a UIODMA write operation request from a PCIe device to a CXL.mem device is illustrated at 506. As shown, in some embodiments, the UIODMA write operation request passes through the CXL switch, but not through the root complex or cache consistency control unit.

[0050] Return to reference Figure 4 At subprocess 430, at point 436, the CXL.mem device then updates the given address range of the CXL.mem device using the data specified in the UIO DMA write operation request. In some embodiments, updating the given address range with data can be performed in any suitable manner. For example, in some embodiments, updating the given address range can be performed as specified in the CXL specification.

[0051] Then, at 438 of subprocess 430, the CXL.mem device indicates to the PCIe device that the UIO DMA write operation request is complete. In some embodiments, this indication can be made in any suitable manner and can include any suitable content. For example, in some embodiments, this indication can be made by a UIO response.

[0052] Finally, at 418 of subprocess 410, the PCIe device receives an indication that the UIO DMA write operation request is complete from the CXL.mem device. In some embodiments, the PCIe device may receive the indication that the UIO DMA write operation request is complete from the CXL.mem device in any suitable manner.

[0053] refer to Figure 5 According to some embodiments, an example of an indication of a UIO DMA write operation request completion is illustrated at 508. As shown, in some embodiments, the indication of a UIO DMA write operation request completion is transmitted through the CXL switch, but not through the root complex or cache consistency control unit.

[0054] In some embodiments, because the PCIe device issues PBISRR and / or PBISRW early in processes 200 and / or 400, respectively, the PCIe device can perform certain actions that need to occur before making a UIO DMA request while the cache coherence control unit is a handline cache line as described herein (such as reading the local medium and calculating error correction codes / cyclic redundancy codes and / or encrypting data before UIO DMA write, fetching the flash translation table from the host processor memory, performing internal buffer allocation / writeback to create a "transfer" buffer space to absorb data when subsequent DMA is issued, and / or any other suitable operation). In some embodiments, parallelization of these activities reduces end-to-end transaction latency, thereby improving overall performance.

[0055] It should be understood that, Figure 2 and 4 At least some of the processes in the boxes above can be implemented or performed in any order or sequence, not limited to the order and sequence shown and described in the figures. Furthermore, Figure 2 and 4 Some of the processes outlined in the boxes above can be implemented or executed substantially simultaneously, either in appropriate circumstances or in parallel, to reduce latency and processing time. Alternatively or concurrently, certain steps can be omitted. Figure 2 and 4 Some of the processes described in the boxes above.

[0056] In some embodiments, any suitable computer-readable medium may be used to store instructions for performing the functions and / or processes described herein. For example, in some embodiments, the computer-readable medium may be transient or non-transitory. For example, a non-transitory computer-readable medium may include media such as non-transitory magnetic media (such as hard disks, floppy disks, and / or any other suitable magnetic media), non-transitory optical media (such as optical discs, digital video discs, Blu-ray discs, and / or any other suitable optical media), non-transitory semiconductor media (such as flash memory, electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), and / or any other suitable semiconductor media), any suitable medium that is not transient during transmission or has no persistent appearance, and / or any suitable non-transitory tangible medium. As another example, a transient computer-readable medium may include signals on a network, wires, conductors, optical fibers, circuits, any suitable medium that is transient during transmission and has no persistent appearance, and / or any suitable intangible medium.

[0057] While the invention has been described and illustrated in the foregoing illustrative embodiments, it is to be understood that this disclosure has been made by way of example only, and many changes may be made to the details of the implementation of the invention without departing from the spirit and scope of the invention, which is defined only by the following claims. The features of the disclosed embodiments can be combined and rearranged in various ways.

Claims

1. A system for out-of-order input / output direct memory access operations, comprising: Memory; as well as A hardware processor, coupled to the memory and configured to at least: Send an invalid listener request to the host processor's cache coherence control unit; as well as Issue an out-of-order input / output direct memory access operation request to the compute fast link memory device.

2. The system according to claim 1, wherein the out-of-order input / output direct memory access operation request is used for a read operation.

3. The system of claim 2, wherein the hardware processor is further configured to receive a response to the out-of-order input / output direct memory access operation request, including data from the compute fast link memory device.

4. The system of claim 3, wherein the data is updated in response to the return of the invalid listening request.

5. The system of claim 1, wherein the out-of-order input / output direct memory access operation request is used for a write operation.

6. The system of claim 5, wherein the out-of-order input / output direct memory access operation request includes data to be written to the compute fast link memory device.

7. The system of claim 6, wherein the cache line is invalidated in response to the return invalid listening request.

8. A method for out-of-order input / output direct memory access operations, comprising: The hardware processor sends an invalid listener request to the host processor's cache coherence control unit. as well as Issue an out-of-order input / output direct memory access request to the compute fast link memory device.

9. The method of claim 8, wherein the out-of-order input / output direct memory access request is used for a read operation.

10. The method of claim 9, further comprising receiving a response to the out-of-order input / output direct memory access operation request, including data from the compute fast link memory device.

11. The method of claim 10, wherein the data is updated in response to the return of the invalid listening request.

12. The method of claim 8, wherein the out-of-order input / output direct memory access request is used for a write operation.

13. The method of claim 12, wherein the out-of-order input / output direct memory access operation request includes data to be written to the compute fast link memory device.

14. The method of claim 13, wherein the cache line is invalidated in response to the return invalid listening request.

15. A non-transitory computer-readable medium comprising computer-executable instructions, which, when executed by a processor, cause the processor to perform a method for out-of-order input / output direct memory access operations, the method comprising: Send an invalid listen request to the host processor's cache coherence control unit; as well as Issue an out-of-order input / output direct memory access operation request to the compute fast link memory device.

16. The non-transitory computer-readable medium of claim 15, wherein the out-of-order input / output direct memory access operation request is for a read operation.

17. The non-transitory computer-readable medium of claim 16, wherein the method further comprises receiving a response to the out-of-order input / output direct memory access operation request, including data from the compute fast link memory device.

18. The non-transitory computer-readable medium of claim 17, wherein the data is updated in response to the return invalid listening request.

19. The non-transitory computer-readable medium of claim 15, wherein the out-of-order input / output direct memory access operation request is for a write operation.

20. The non-transitory computer-readable medium according to claim 19, wherein, The out-of-order input / output direct memory access operation request includes data to be written to the compute fast link memory device.

21. The non-transitory computer-readable medium of claim 20, wherein the cache line is invalidated in response to the return invalid listening request.