System and method for data replication offload for storage devices

By performing data copying operations directly between two storage devices in the storage system and utilizing the buffer of the destination device for data transfer, the problem of high host resource consumption in traditional data copying is solved, achieving more efficient data copying.

CN114637462BActive Publication Date: 2025-10-28KIOXIA CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202111538351.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-02-05
Filing Date
2021-12-15
Publication Date
2025-10-28
Estimated Expiration
2041-12-15

AI Technical Summary

Technical Problem

In traditional data replication, the host needs to consume a lot of resources to transfer data from the first solid-state device to the second solid-state device, resulting in high resource consumption, increased latency, and high bandwidth consumption.

Method used

By performing data copying operations directly between two storage devices in a storage system, and utilizing the buffer of the destination device for data transfer, the involvement of the host is reduced, enabling peer-to-peer (P2P) transfer and sharing the buffer address to achieve direct data transfer.

Benefits of technology

This reduces the use of host resources, lowers data replication latency and bandwidth consumption, improves data replication efficiency, and reduces costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114637462B_ABST
    Figure CN114637462B_ABST
Patent Text Reader

Abstract

This disclosure generally relates to systems and methods for data replication offloading of storage devices. Various embodiments described herein relate to systems and methods for transferring data from a source device to a destination device, comprising: receiving a replication request from a host via the destination device; performing a transfer with the source device via the destination device to transfer data from a buffer of the source device to a buffer of the destination device; and writing the data via the destination device to a non-volatile storage device of the destination device.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-referencing of related patent applications

[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 126,442, filed December 16, 2020, entitled "Systems and Methods for Data Copy Offload for Storage Devices," the contents of which are incorporated herein by reference in their entirety and for all purposes as if fully and completely set forth herein. Technical Field

[0003] This disclosure generally relates to systems, methods, and non-transitory processor-readable media for data replication in data storage devices. Background Technology

[0004] In a traditional data copying process, where data is copied from a first solid-state drive (SSD) to a second SSD, the host sends a read command to the first SSD, transferring the data from the first SSD to the host's local storage. Subsequently, the host sends a write command to the second SSD, transferring the data from the host's local storage to the second SSD. This process requires significant resources from the host to execute. Summary of the Invention

[0005] In some arrangements, a method for transferring data from a source device to a destination device includes: receiving a copy request from a host via the destination device; performing a transfer with the source device via the destination device to transfer the data from a buffer in the source device to a buffer in the destination device; and writing the data to a non-volatile storage device of the destination device via the destination device.

[0006] In some arrangements, a method for transferring data from a source device to a destination device includes: communicating with the destination device via the source device to set the filling of a buffer in the source device, and performing a transfer with the destination device via the source device to transfer the data from the buffer in the source device to the buffer in the destination device. Attached Figure Description

[0007] Figure 1A A block diagram illustrating an example of a system including a storage device and a host according to some implementation schemes is shown.

[0008] Figure 1B A block diagram illustrating examples of buffers for a source device and a buffer for a destination device according to some embodiments is shown.

[0009] Figure 2 This is a block diagram illustrating an example method for performing data replication, based on some implementation schemes.

[0010] Figure 3 This is a block diagram illustrating an example method for performing data replication, based on some implementation schemes.

[0011] Figure 4 This is a block diagram illustrating an example method for performing data replication, based on some implementation schemes.

[0012] Figure 5 This is a block diagram illustrating an example method for performing data replication, based on some implementation schemes.

[0013] Figure 6 This is a block diagram illustrating an example method for performing data replication, based on some implementation schemes.

[0014] Figure 7 This is a block diagram illustrating an example method for performing data replication, based on some implementation schemes.

[0015] Figure 8 This is a block diagram illustrating an example method for performing data replication, based on some implementation schemes.

[0016] Figure 9A This is a block diagram illustrating an example method for performing data replication, based on some implementation schemes.

[0017] Figure 9B This is a block diagram illustrating an example method for performing data replication, based on some implementation schemes.

[0018] Figure 10 This is a flowchart illustrating an example method for performing data replication according to some implementation schemes. Detailed Implementation

[0019] The arrangements disclosed herein relate to data replication schemes that are cost-effective solutions that do not compromise the need to meet business demands more quickly. This disclosure improves data replication while producing solutions consistent with current system architectures and evolving changes. In some arrangements, this disclosure relates to collaboratively performing data replication operations between two or more components of a storage system. The data replication operation occurs directly between the two or more components of the storage system, without involving any third party (e.g., but not limited to, a storage system controller or host) during data replication, thereby reducing or “offloading” command, control, or data buffering tasks from third parties. While non-volatile memory devices are presented as examples herein, the disclosed schemes can be implemented on any storage system or device that interfaces with a host and temporarily or permanently stores data for later retrieval by the host.

[0020] To help illustrate this implementation plan, Figure 1A A block diagram of a system comprising storage devices 100a, 100b, ..., 100n (collectively referred to as storage device 100) coupled to a host 101 is shown, according to some examples. The host 101 may be a user device operated by a user or an autonomous central controller of the storage device, wherein the host 101 and storage device 100 correspond to a storage subsystem or storage device. The host 101 may be connected to a communication network 109 (via a network interface 108) so that other host computers (not shown) can access the storage subsystem or storage device via the communication network 109. Examples of such storage subsystems or devices include all-flash array (AFA) or network-attached storage (NAS) devices. As shown, the host 101 includes a memory 102, a processor 104, and a bus 106. The processor 104 is operatively coupled to the memory 102 and the bus 106. The processor 104 is sometimes referred to as the central processing unit (CPU) of the host 101 and is configured to execute processes of the host 101.

[0021] Memory 102 is the local memory of host 101. In some instances, memory 102 is a buffer, sometimes referred to as a host buffer. In some instances, memory 102 is a volatile storage device. In other instances, memory 102 is a non-volatile permanent storage device. Examples of memory 102 include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static RAM (SRAM), magnetic RAM (MRAM), phase-change memory (PCM), and so on.

[0022] Bus 106 includes one or more of software, firmware, and hardware that provide an interface for components of host 101 to communicate. Examples of components include, but are not limited to, processor 104, network card, storage device, memory 102, graphics card, etc. Additionally, host 101 (e.g., processor 104) can use bus 106 to communicate with storage device 100. In some instances, storage device 100 is directly attached to or communicatively coupled to bus 106 via a suitable interface 140. Bus 106 is one or more of serial, PCIe bus or network, PCIe root complex, internal PCIe switch, etc.

[0023] Processor 104 can execute an operating system (OS) that provides a file system and applications using the file system. Processor 104 can communicate with storage devices 100 (e.g., controllers 110 of each storage device 100) via a communication link or network. To this end, processor 104 can use interface 140, which connects to the communication link or network, to send and receive data to one or more of the storage devices 100. Interface 140 allows software running on processor 104 (e.g., a file system) to communicate with storage devices 100 (e.g., their controllers 110) via bus 106. Storage devices 100 (e.g., their controllers 110) are operatively coupled directly to bus 106 via interface 140. Although interface 140 is conceptually shown as a dashed line between host 101 and storage device 100, interface 140 may include one or more controllers, one or more physical connectors, one or more data transfer protocols, including namespaces, ports, transport mechanisms, and their connections. Although the connection between host 101 and storage device 100 is shown as a direct link, in some implementations the link may include a network structure that may include networking components such as bridges and switches.

[0024] To send and receive data, processor 104 (and the software or file system running thereon) communicates with storage device 100 using a storage data transfer protocol running on interface 140. Examples of protocols include, but are not limited to, SAS, Serial ATA (SATA), and NVMe protocols. In some instances, interface 140 includes hardware (e.g., a controller) implemented on or operatively coupled to bus 106, storage device 100 (e.g., controller 110), or another device operatively coupled to bus 106 and / or storage device 100 via one or more suitable networks. Interface 140 and the storage protocols running thereon also include software and / or firmware executed on such hardware.

[0025] In some instances, processor 104 may communicate with communication network 109 via bus 106 and network interface 108. Other host systems (not shown) attached to or communicatively coupled to communication network 109 may communicate with host 101 using suitable network storage protocols, examples of which include, but are not limited to, NVMe (NVMeoF), iSCSI, Fibre Channel (FC), Network File System (NFS), Server Message Block (SMB), etc. Network interface 108 allows software running on processor 104 (e.g., storage protocols or file systems) to communicate with external hosts attached to communication network 109 via bus 106. In this manner, network storage commands can be issued by the external host and processed by processor 104, which can issue storage commands to storage device 100 as needed. Thus, data can be exchanged between the external host and storage device 100 via communication network 109. In this example, any exchanged data is buffered in memory 102 of host 101.

[0026] In some instances, storage device 100 is located in a data center (not shown for simplicity). The data center may contain one or more platform or rack units, each supporting one or more storage devices (e.g., but not limited to storage device 100). In some embodiments, host 101 and storage device 100 together form a storage node, where host 101 acts as a node controller. An example of a storage node is a Kioxia Kumoscale storage node. One or more storage nodes within a platform are connected to a top-of-rack (TOR) switch, each storage node connected to the TOR via one or more network connections, such as Ethernet, Fibre Channel, or wireless band, and can communicate with each other via the TOR switch or another suitable intra-platform communication mechanism. In some embodiments, storage device 100 may be a network-connected storage device (e.g., an Ethernet SSD) connected to a TOR switch, where host 101 is also connected to the TOR switch and can communicate with storage device 100 via the TOR switch. In some embodiments, at least one router facilitates communication between storage devices 100 in storage nodes across different platforms, racks, or cabinets via a suitable networking structure. Examples of storage device 100 include non-volatile devices, such as, but not limited to, solid-state drives (SSDs), Ethernet-connected SSDs, non-volatile dual in-line memory modules (NVDIMMs), universal flash memory (UFS), secure digital storage (SD) devices, and so on.

[0027] Each of the storage devices 100 includes at least a controller 110 and a memory array 120. For simplicity, other components of the storage device 100 are not shown. The memory array 120 includes NAND flash memory devices 130a-130n. Each of the NAND flash memory devices 130a-130n includes one or more individual NAND flash dies, which are NVMs capable of retaining data without power. Therefore, NAND flash memory devices 130a-130n refer to multiple NAND flash memory devices or dies within the flash memory device 100. Each of the NAND flash memory devices 130a-130n includes one or more dies, wherein each die has one or more planes. Each plane has multiple blocks, and each block has multiple pages.

[0028] Although NAND flash memory devices 130a-130n are shown as examples of memory array 120, other examples of non-volatile memory technologies used to implement memory array 120 include, but are not limited to, non-volatile (with backup battery) DRAM, magnetic random access memory (MRAM), phase-change memory (PCM), ferroelectric RAM (FeRAM), and the like. The arrangements described herein can also be implemented on memory systems using such memory technologies and other suitable memory technologies.

[0029] Instances of controller 110 include, but are not limited to, SSD controllers (e.g., client SSD controllers, data center SSD controllers, enterprise SSD controllers, etc.), UFS controllers, or SD controllers, etc.

[0030] Controller 110 can combine raw data storage from multiple NAND flash memory devices 130a-130n, enabling those NAND flash memory devices 130a-130n to logically operate as a single memory cell. Controller 110 may include a processor, microcontroller, buffers (e.g., buffers 112, 114, 116), error correction system, data encryption system, flash translation layer (FTL), and flash interface module. Such functionality may be implemented in hardware, software, and firmware, or any combination thereof. In some arrangements, the software / firmware of controller 110 may be stored in memory array 120 or any other suitable computer-readable storage medium.

[0031] Controller 110 includes suitable processing and memory capabilities for performing the functions described herein, as well as other functions. As described, controller 110 manages various features of NAND flash memory devices 130a-130n, including, but not limited to, I / O processing, reading, writing / programming, erasing, monitoring, logging, error handling, garbage collection, wear leveling, logical-to-physical address mapping, data protection (encryption / decryption, cyclic redundancy check (CRC)), error correction decoding (ECC), data scrambling, etc. Therefore, controller 110 provides visibility to NAND flash memory devices 130a-130n.

[0032] Buffer memory 111 is a memory device locally and operatively coupled to controller 110. For example, buffer memory 111 may be on-chip SRAM located on a chip of controller 110. In some embodiments, buffer memory 111 may be implemented using a memory device of storage device 110 external to controller 110. For example, buffer memory 111 may be on-chip DRAM located on a chip other than controller 110. In some embodiments, buffer memory 111 may be implemented using memory devices both internal and external to controller 110 (e.g., on-chip and off-chip of controller 110). For example, buffer memory 111 may be implemented using internal SRAM and external DRAM, which are transparent / exposed and accessible by other devices via interface 140, such as host 101 and other storage devices 100. In this example, controller 110 includes an internal processor using memory addresses within a single address space, and a memory controller that controls the internal SRAM and external DRAM, selecting whether to place data on the internal SRAM and external DRAM based on efficiency. In other words, the internal SRAM and external DRAM are addressed like a single memory. As shown in the figure, the buffer memory 111 includes a buffer 112, a write buffer 114, and a read buffer 116. In other words, buffer 112, write buffer 114, and read buffer 116 can be implemented using the buffer memory 111.

[0033] Controller 110 includes buffer 112, which is sometimes referred to as a drive buffer or controller memory buffer (CMB). In addition to being accessible by controller 110, buffer 112 is also accessible by other devices via interface 140, such as host 101 and others in storage device 100. In this manner, buffer 112 (e.g., the address of a memory location within buffer 112) is exposed across bus 106, and devices on bus 106 can issue read and write commands using the address or logical address (e.g., logical block address (LBA)) corresponding to the memory location within buffer 112 to read and write data within the buffer. In some instances, buffer 112 is a volatile memory device. In some instances, buffer 112 is a non-volatile permanent memory device. Examples of buffer 112 include, but are not limited to, RAM, DRAM, SRAM, MRAM, PCM, etc. Buffer 112 may refer to multiple buffers, each configured to store different types of data, as described herein.

[0034] In some implementation schemes, such as Figure 1A As shown, buffer 112 is local memory of controller 110. For example, buffer 112 may be on-chip SRAM memory located on the chip of controller 110. In some embodiments, buffer 112 may be implemented using memory devices of storage device 110 external to controller 110. For example, buffer 112 may be on-chip DRAM located on a chip other than controller 110. In some embodiments, buffer 112 may be implemented using memory devices both internal and external to controller 110 (e.g., on and off the chip of controller 110). For example, buffer 112 may be implemented using internal SRAM and external DRAM, which are transparent / exposed and accessible by other devices via interface 140, such as host 101 and other storage devices 100. In this example, controller 110 includes an internal processor using memory addresses within a single address space, and a memory controller that controls the internal SRAM and external DRAM, selecting whether to place data on the internal SRAM and external DRAM based on efficiency. In other words, the internal SRAM and external DRAM are addressed like a single memory.

[0035] In one example of a write operation, in response to receiving data from host 101 (via host interface 140), controller 110 acknowledges the write command from host 101 after writing the data to write buffer 114. In some embodiments, write buffer 114 may be implemented in a separate memory separate from buffer 112, or write buffer 114 may be a defined region or portion of the memory that includes buffer 112, wherein only the CMB portion of the memory is accessible by other devices but not by write buffer 114. Controller 110 may write the data stored in write buffer 114 to memory array 120 (e.g., NAND flash memory devices 130a-130n). Once the physical address of the data written to memory array 120 is complete, FTL updates the mapping between the logical address (e.g., LBA) used by host 101 to associate with the data and the physical address used by controller 110 to identify the physical location of the data. In another example relating to a read operation, controller 110 includes another buffer 116 (e.g., a read buffer), different from buffers 112 and 114, for storing data read from memory array 120. In some embodiments, read buffer 116 may be implemented in a separate memory different from buffer 112, or read buffer 116 may be a defined region or portion of the memory that includes buffer 112, wherein only the CMB portion of the memory is accessible by other devices but not by read buffer 116.

[0036] Although non-volatile memory devices (e.g., NAND flash memory devices 130a-130n) are presented as examples herein, the disclosed schemes can be implemented on any storage system or device connected to host 101 via an interface, wherein such system temporarily or permanently stores data for host 101 to retrieve later.

[0037] Although storage device 100 is shown and described as a separate physical device, the arrangements disclosed herein can also be applied to virtualized storage device 100. For example, the controller 110 and memory array 120 of each storage device 100 may be virtualized relative to hardware components such as processors and memory.

[0038] Traditionally, to copy data corresponding to a single logical address (e.g., an LBA) or more generally a range of logical addresses (e.g., an LBA range, specified using a starting LBA and the length B of the data bytes to be transferred) from a first storage device (e.g., storage device 100a) to a second storage device (e.g., storage device 100b), host 101 needs to allocate a buffer in memory 102 and send an I / O command (e.g., a read command) to the first storage device via interface 140 and bus 106 to read the data corresponding to the logical address, where the address of memory 102, the source LBA, and the data length B are provided as parameters in the read command. Storage device 100a translates the logical address into a physical address, at which B data bytes are read from memory array 120 into read buffer 116. After reading all the data corresponding to the B data bytes, controller 110 transfers the B bytes of data from read buffer 116 in the first storage device (typically using a direct memory transfer operation on bus 106) to the allocated buffer in memory 102 via interface 140 and bus 106. Finally, controller 110 indicates to host 101 that the read command has been completed without errors. Host 101 sends another I / O command (e.g., a write command) to the second device via interface 140 and bus 106, providing the address of the buffer in memory 102, the destination LBA, and the length B as command parameters. In response to the second I / O command, controller 110 transfers data from the allocated buffer in memory 102 (typically using a direct memory transfer operation on bus 106) to write buffer 114. Controller 110 then performs a destination LBA-to-physical address translation and transfers B bytes of data from write buffer 114 to the physical address in the memory array 120 of the second storage device via interface 140 and bus 106. Consecutive blocks of data can be transferred by specifying an LBA range using a starting LBA and a length value indicating the size of the data block. Alternatively, using a storage protocol such as NVMe, a list of LBA and length pairs (e.g., a hash list (SGL)) can be used to specify a set of non-consecutive data blocks. In this case, the aforementioned process is repeated for each LBA and length pair in the list until all data blocks have been copied.

[0039] All these operations require CPU cycles, context switching, etc., on the processor 104 of host 101. Furthermore, the transfers performed between processor 104 and memory 102 consume memory space (data buffers, commit queues (SQ) / complete queues (CQ)) and memory bus bandwidth between processor 104 and memory 102. Additionally, data communication between processor 104 and bus 106 consumes the bandwidth of bus 106, which is considered a valuable resource because bus 106 acts as the interface between different components of host 101 and between the storage devices 100a, 100b…100n themselves. Therefore, conventional data replication schemes consume significant resources on host 101 (e.g., bandwidth, CPU cycles, and buffer space).

[0040] One arrangement disclosed herein relates to data replication based on peer-to-peer (P2P) transfer within storage device 100. In a data replication scheme using P2P transfer, a local memory buffer (e.g., buffer 112) of storage device 100 is used to perform data transfer from one storage device (e.g., storage device 100a) to another storage device (e.g., storage device 100b). Data replication involving multiple storage devices 100 can be performed within storage device 100, without the participation of host 101, which triggers the data replication operation by sending an I / O command to the first storage device (e.g., storage device 100a) among more than two storage devices 100. Therefore, data no longer needs to be copied to the memory 102 of host 101, thereby reducing the latency and bandwidth required to transfer data in and out of memory 102. The number of I / O operations can be reduced to a single I / O for replication operations involving logical address ranges and a considerable number of storage devices 100. The efficiency benefits not only improve performance but also reduce cost, power consumption, and network usage.

[0041] To achieve this improved efficiency, the address of the buffer 112 of each storage device 100 is shared with all storage devices 100, so that each storage device 100 knows the address of the buffer 112 of the other storage devices 100. For example, storage device 100a knows the address of the buffer 112 of each storage device 100b-100n, storage device 100b knows the address of the buffer 112 of each storage device 100a, 100c-100n, storage device 100n knows the address of the buffer 112 of each storage device 100a-100n-1, and so on.

[0042] The address of buffer 112 of storage device 100 is shared when using various mechanisms. In some embodiments, each storage device 100 may obtain the address of buffer 112 of other storage devices 100 from a designated entity, or the address of buffer 112 may be stored in a shared address register. For example, the address of buffer 112 of storage device 100 may be exposed to host 101 across bus 106 (e.g., via a base address register) to be shared with storage device 100. An example of a base address register is a shared PCIe base address register, which is a shared address register and a designated entity. The address of buffer 112 for CMB may be, for example, the NVMe controller register CMBLOC, which contains the PCI address location of the start of buffer 112 (the beginning of buffer 112) and the controller register CMBSZ (the size of buffer 112).

[0043] In another instance, the address of buffer 112 may be managed and provided by another suitable entity with appropriate processing and memory capabilities (e.g., processing circuitry), and communicatively coupled to storage device 100 using interface 140 or another suitable network (e.g., local or Ethernet connection). Considering that storage devices 100 within a group may change from time to time due to one or more storage devices coming online, being added to the group, being removed from the group, being shut down, etc., the designated entity may maintain an up-to-date list of addresses for buffer 112.

[0044] In another instance, the address of buffer 112 does not need to be shared separately beforehand, but is specified separately in the command used in the copy operation. In one embodiment, host 101 can issue a copy command to the source device, specifying the destination device's buffer address 112 as a parameter. Thus, the source device can read data from the source device's memory array 120 and transfer the data to buffer 112 in the destination device. The source device can issue a write command to the destination device, specifying the destination device's own buffer 112 as the address of the data to be written, to complete the copy operation. In another embodiment, the source device receives a copy command from host 101, specifying its own buffer 112 as the address where the data being read is to be placed. Then, source device 100a specifies its own buffer 112 as the address of the data to be written to the destination device in the write command to complete the copy operation.

[0045] In some embodiments, host 101 may issue a copy command to the destination device, designating the storage device as the source device and specifying the address of the source device's buffer 112 as the location where the source data will be located. The destination device then issues a read command to the source device, specifying its own buffer 112 for the read data. The destination device then uses the source device's buffer 112 as the location for the data to be written to perform the data write operation. In another embodiment, host 101 may issue a copy command to the destination device, designating the storage device as the source device and specifying the address of the destination device's buffer 112 as the location where the source data is to be transferred. The destination device then issues a read command to the source device, specifying its own buffer 112 for the read data. The destination device then uses the destination device's buffer 112 as the location for the data to be written to perform the data write operation.

[0046] The sharing of the address of buffer 112 allows storage devices 100 to communicate directly with each other and transfer data in a P2P manner as described herein. In one instance, buffer 112 is a CMB separate from write buffer 114 and read buffer 116. In another instance, buffer 112 may be the same as write buffer 114 or read buffer 116, which shares its address with storage device 100 in the manner described herein. In yet another instance, the address of the CMB of one of the storage devices 100 shared with the other storage devices 100 may be mapped to write buffer 114 or read buffer 116. For this purpose, buffer 112 is used herein to refer to a buffer of a storage device whose address is shared among storage devices 100 for data copying operations. Examples of addresses include, but are not limited to, CMB addresses, identifiers, pointers, or other suitable indicators that identify buffer 112 of a storage device.

[0047] As used herein, a source device refers to a storage device containing the data to be copied. For example, a source device stores the data to be copied in memory array 120. Examples of source devices include, but are not limited to, PCIe NVMe devices, NVMeoF devices (e.g., Ethernet SSDs), and so on.

[0048] As used herein, a destination device refers to a storage device to which data is transferred (copied). For example, data may be transferred from a source device to a destination device, and in some cases, existing data stored in the memory array 120 of the destination device may be updated to the transferred data. Examples of destination devices include, but are not limited to, PCIe NVMe devices, fabricated NVMe (NVMeoF) devices (e.g., Ethernet SSDs), and the like.

[0049] As used herein, the target device is either the source device or the destination device.

[0050] As used in this article, namespace copying refers to the copying of the entire range of LBAs within a namespace. For namespace copying, specify the source namespace and destination namespace, as well as the starting LBA and the length of the LBA range.

[0051] As used in this document, LBA copying refers to the copying of a specified LBA using a starting LBA and an LBA range length definition. For LBA copying, the source namespace and destination namespace are specified.

[0052] As used herein, intra-network replication refers to the replication of data from a source device to a destination device, wherein the source device and the destination device are PCIe NVMe devices within the same PCIe network or NVMeoF devices within the same NVMeoF network.

[0053] As used herein, off-network replication refers to the replication of data from a source device to a destination device, where the source and destination devices are NVMeoF devices on two different networks connected via NVMeoF (e.g., NVMeoF devices connected via an Ethernet network structure, or Ethernet SSDs implementing the NVMeoF protocol).

[0054] As used herein, one-to-one replication refers to a use case where only one source device stores the data to be replicated (instead of multiple source devices), and the data is replicated only to one destination device (instead of multiple destination devices). Instances of one-to-one replication include, but are not limited to, namespace replication or LBA replication. One-to-one replication can be intra-network replication or inter-network replication.

[0055] As used herein, many-to-one replication refers to a use case where multiple source devices store data to be replicated, and the data is replicated to a single destination device. In some instances, the order specified in the P2P replication descriptor indicates the order in which the destination device replicates data from the multiple source devices. Examples of many-to-one replication include, but are not limited to, namespace replication or LBA replication. Many-to-one replication can be intra-network replication or inter-network replication.

[0056] Figure 1B A block diagram illustrating examples of buffers for a source device and a buffer for a destination device according to some embodiments is shown. (Reference) Figure 1A and 1B The buffer 112 of the source device (e.g., storage device 100a) and the buffer 112 of the destination device (e.g., storage device 100b) are in Figure 1B As shown in the image.

[0057] The source device's buffer 112 includes a P2P reserved area 150 and a data area 152. The P2P reserved area 150 and data area 152 are different predefined areas or partitions within the memory device of the source device's buffer 112. The P2P reserved area 150 can be identified by its own address range exposed on interface 140, defined by a starting address and a size. The P2P reserved area 150 is a designated area or partition within the source device's buffer 112 reserved for P2P messages. The source device and the destination device can directly transmit P2P messages to each other via interface 140 or another suitable communication protocol or network, as the reserved area 150 can be read from or written to by the storage device 100 using the exposed address range on interface 140. Another device (e.g., the destination device) can send the P2P message to the source device by writing the P2P message to an address within the P2P reserved area 150. In response to the source device detecting that a P2P message has been written to the P2P reserved area 150, the source device may begin processing the P2P message. In some instances, the current operation of the source device may be paused or interrupted to process the P2P message.

[0058] As described herein, a P2P message to the source device can trigger the source device to read data stored in the source device's memory array 120 into a dynamically allocated buffer in the source device's data area 152 or into the source device's read buffer 116. For this purpose, the source device's controller 110 can dynamically allocate one or more buffers 154, 156, ..., 158 within the data area 152. Each buffer 154, 156, ..., 158 has its own address. Data can be read from the source device's memory array 120 into buffers 154, 156, ..., 158 using a read operation. Each of these buffers can have a buffer size suitable for temporarily storing or temporarily holding blocks of the entire data to be transferred to the destination device. In some instances, the source device can provide addresses to the destination device by writing the addresses of buffers 154, 156, ..., 158 into the P2P reserved area 160 of the destination device. This allows the destination device to perform read operations on those addresses to read data stored in the source device's buffers 154, 156, ..., 158 into the destination device's buffers (e.g., the write buffer of write buffer 114). Buffers 154, 156, ..., 158 can be allocated in a sliding window, meaning that some of buffers 154, 156, ..., 158 are allocated in a first time window in response to a P2P message to temporarily store data blocks for transfer. Once the transfer of these data blocks to the destination device is complete, those buffers are deallocated or recycled, and their associated memory capacity is freely allocated as additional buffers for buffers 154, 156, ..., 158. Corresponding write buffers in the destination device's write buffer 114 can be allocated similarly.

[0059] The destination device's buffer 112 includes a P2P reserved area 160 and a data area 162. The P2P reserved area 160 and the data area 162 are different predefined areas or partitions within the memory device of the destination device's buffer 112. The P2P reserved area 160 can be identified by its own address range exposed on interface 140, defined by a starting address and a size. The P2P reserved area 160 is a designated area or partition within the address space of the destination device's buffer 112 reserved for P2P messages. Another device (e.g., a source device) can send a P2P message to the destination device by writing it to an address within the address range of the P2P reserved area 160. In response to the destination device detecting that a P2P message has been written to the P2P reserved area 160, the destination device can begin processing the P2P message.

[0060] The controller 110 of the destination device can dynamically allocate one or more buffers 164, 166, ..., 168 within the data area 162. Each buffer 164, 166, ..., 168 has its own address. Data can be written from buffers 164, 166, ..., 168 to the memory array 120 of the destination device using write operations. Each of these buffers may have a buffer size suitable for temporarily storing or pausing blocks of data received from the source device. In some instances, the destination device can provide these addresses to the source device by writing the addresses of buffers 164, 166, ..., 168 to the P2P reserved area 150 of the source device, so that the source device can perform write operations on those addresses to write data stored in buffers of the source device (e.g., read buffers in read buffer 116) to buffers 164, 166, ..., 168 of the destination device. Buffers 164, 166, ..., 168 can be allocated in a sliding window, meaning that some of buffers 164, 166, ..., 168 are allocated in a first time window to receive the first data chunk from the source device. When writing these data chunks to the memory array 120 of the destination device is complete, those buffers are deallocated or recycled, and their associated memory capacity is freely allocated as additional buffers for buffers 164, 166, ..., 168. Corresponding read buffers in the read buffer 116 of the source device can be allocated similarly.

[0061] Figure 2 This is a block diagram illustrating an example method 200 for performing data replication, based on some implementation schemes. (Reference) Figure 1A-2 In method 200, data is copied from a source device (e.g., storage device 100a) to a destination device (e.g., storage device 100b). Figure 2 Within this context, storage devices 100a and 100n are mirror drives, and storage device 100n has failed. Therefore, a backup drive (storage device 100b) is brought online to mirror the data stored on storage device 100a, making both mirror devices online simultaneously. For this purpose, the source and destination devices will have the same namespace size, and LBA replication will maintain the same LBA offset within the namespace. Therefore, Figure 2 The one-to-one LBA copy operation at the same offset is described.

[0062] like Figure 2As shown, the memory array 120 of storage device 100a has a storage capacity 201 that stores data containing data corresponding to namespace 210. The memory array 120 of storage device 100b has a storage capacity 202 that stores data containing data corresponding to namespace 220. Data corresponding to LBA range 211 of namespace 210 is copied to storage capacity 202 of storage device 100b as data corresponding to LBA range 221 of namespace 220. Namespaces 210 and 220 may have the same identifier (e.g., the same number / index, such as "namespace 1") and the same size. LBA range 211 begins with a start LBA 212 and ends with a stop LBA 213, and is generally defined in commands by the length of the start LBA 212 and the LBA range 211. LBA range 221 begins with start LBA 222 and ends with end LBA 223, and is typically defined in a command by the length of the start LBA 222 and the LBA range 221. Start LBAs 212 and 222 may have the same identifier (e.g., the same number / index) provided the offsets are the same.

[0063] In some arrangements, such as Figure 2 In the one-to-one LBA copy operation shown, host 101 may send a command containing descriptors to the destination device (e.g., storage device 100b), such descriptors being one or more of the following: (1) the LBA range 211 to be copied (defined by the starting LBA 212 and the length of the LBA range 211); (2) the address of the buffer 112 of the source device (e.g., storage device 100a); and (3) other information, such as the data copy type (e.g., one-to-one LBA copy operation).

[0064] In some arrangements, the address of buffer 112 may be shared beforehand, rather than being sent directly by host 101 in a command to the buffer 112 of the source device. For example, a table may contain the starting address of the buffer for each storage device 100 mapped to an ID. This table may be generated and updated by host 101 or as a central address agent (not shown) of host 101, or the central address agent may access each storage device 100. The central address agent is communicatively coupled to host 101 and storage device 100 via a suitable network. In some instances, the table may be stored by the central address agent. In instances where the table is generated by host 101, host 101 may send the table (and its updates) to the central address agent for storage. In some instances, host 101 or the central address agent may send the table (and its updates) to each storage device 100 for storage. The command contains a descriptor of the ID of the buffer 112 of the source device, rather than the address of the buffer 112 of the source device. In response to receiving an ID, the destination device looks up a table stored and managed in the central address agent or the destination device itself to determine the starting address of the buffer 112 corresponding to the ID.

[0065] In some instances, a device (e.g., a destination device) uses the source device's buffer 112. In some instances, the source device submits the buffer 112 for use by another device, which interrupts or stops the other use during the span of the copy operation.

[0066] In response to a received command, the destination device initiates communication with the source device to transfer data corresponding to LBA range 211 from the source device to the destination device. For example, the destination device may send a request to the source device via interface 140, whereby the request includes the LBA range 211 to be copied (defined by the starting LBA 212 and the length of LBA range 211). Upon receiving the request, the source device reads data stored in array 120 into its buffer 112, which is exposed to the storage device 100 containing the destination device. The destination device can then transfer the data stored in the source device's buffer 112 to the destination device's write buffer via interface 140 based on LBA range 221, and then to array 120. This can be a write operation to the destination device. For example, a write command specifying the source device's buffer 112 as the location to be written is sent from the source device to the destination device.

[0067] In some instances, the command sent from the destination device to the source device includes the address or ID of the destination device's buffer 112. In response to receiving the command, the source device reads data from the source device's array 120 and transfers the data directly to the destination device's buffer 112.

[0068] Therefore, when the destination device issues a command to the source device depends on the address or ID provided by the destination device. In the case where the address ID of the source device's own buffer 112 is provided in the command, the copy operation does not require the source device's buffer 112.

[0069] In some instances, the source device notifies the destination device that it has finished reading a data chunk from array 112 and that the data chunk has been placed. In some instances, chunk and descriptor transfers are always in units of LBAs. Therefore, when the source device notifies the destination that the transfer of a data chunk is complete, the source device writes an LBA count indicating that the transfer is just completed into the destination device's buffer 112 (e.g., P2P reservation area 160). In some instances, the destination device acknowledges receipt of the transfer, allowing the source device to reuse the same buffer to place the next data chunk from array 112. The destination device acknowledges receipt by writing a count of the remaining LBAs that still need to be transferred through the source device to satisfy the descriptor transfer requirements into the source device's buffer 112 (e.g., P2P reservation area 150). In response to the destination device writing zero for the remaining LBAs, the source device determines that the descriptor transfer request has been fully processed. In some arrangements, such as... Figure 2 In the illustrated one-to-one LBA copy operation, host 101 may send a command containing descriptors to a source device (e.g., storage device 100a), such descriptors include one or more of the following: (1) the LBA range 211 to be copied (defined by a starting LBA 212 and the length of the LBA range 211); (2) the address or ID of the buffer 112 of the destination device (e.g., storage device 100b); and (3) other information, such as the data copy type (e.g., one-to-one LBA copy operation). In response to receiving the command, the source device reads data stored in array 120 and transfers it (e.g., via DMA or non-DMA write transfer) to the buffer 112 of the destination device, which is exposed to the storage device 100 containing the source device. The source device then initiates communication with the destination device to write the data in the destination device's buffer 112 corresponding to the LBA range 211 read from the source device to the destination device array 120. For example, a source device may send a write request to a destination device via interface 140, wherein the request includes data already containing data corresponding to LBA range 211 (defined by the starting LBA 212 and the length of LBA range 211), the destination starting LBA, and the address of a destination buffer 112 containing the length of the data to be written (which may be the same as the source starting LBA and length). The destination device writes the data stored in the destination buffer 112 to the array 120 of the destination device.

[0070] In some instances, the source device reads data from array 120 into its buffer 112 and sends a write request to the destination device specifying the address or ID of the source device's buffer 112. The destination device then performs a transfer, moving the data from the source device's buffer 112 into the destination device.

[0071] Figure 3 This is a block diagram illustrating an example method 300 for performing data replication, based on some implementation schemes. (Reference) Figure 1A , 1B In method 300, data is copied from a source device (e.g., storage device 100a) to a destination device (e.g., storage device 100b). Figure 3 Within this context, storage devices 100a and 100n are virtualization mirroring drivers, and storage device 100n fails. Therefore, a backup virtualization driver (storage device 100b) is brought online to mirror the data stored on storage device 100a, resulting in both mirroring virtualization devices being online simultaneously. For this reason, although the source and destination devices have the same namespace size, LBA replication corresponds to a different offset in the destination namespace due to virtualization. Therefore, Figure 3 The one-to-one LBA copy operation at different offsets is described.

[0072] like Figure 3 As shown, the memory array 120 of storage device 100a has a storage capacity 301, which stores data containing data corresponding to namespace 310. The memory array 120 of storage device 100b has a storage capacity 302, which stores data containing data corresponding to namespace 320. Data corresponding to LBA range 311 of namespace 310 is copied to storage capacity 302 of storage device 100b as data corresponding to LBA range 321 of namespace 320. Namespaces 310 and 320 may have the same identifier (e.g., the same number / index, such as "namespace 1") and the same size. LBA range 311 begins with a start LBA 312 and ends with a stop LBA 313, and is generally defined in commands by the LBA lengths of the start LBA 312 and LBA range 311. LBA range 321 begins with start LBA 322 and ends with end LBA 323, and is typically defined in a command by the lengths of start LBA 322 and LBA range 321. Start LBAs 312 and 322 may have different identifiers (e.g., different numbers / indices) depending on the offset.

[0073] Figure 4 This is a block diagram illustrating an example method 400 for performing data replication, based on some implementation schemes. (Reference) Figure 1A ,1B In method 400, data is copied from a source device (e.g., storage device 100a) to a destination device (e.g., storage device 100b). Figure 4 Within the context of [the context], physical drive copying is performed, causing all information corresponding to the namespace (including partition information, file system, etc.) to be copied from the source device to the destination device. For this purpose, the source and destination devices have the same namespace size, and the entire namespace is copied. Therefore, Figure 4 A one-to-one LBA copy for the same namespace on storage devices 100a and 100b is described.

[0074] like Figure 4 As shown, the memory array 120 of storage device 100a has a storage capacity 401, which stores data containing data corresponding to namespace 410. The memory array 120 of storage device 100b has a storage capacity 402, which stores data containing data corresponding to namespace 420. Data corresponding to the entire namespace 410 is copied as data corresponding to namespace 420 to the storage capacity 402 of storage device 100b. Namespaces 210 and 220 may have the same identifier (e.g., the same number / index, such as "namespace 1") and the same size. Namespace 410 begins with a start LBA 412 and ends with a stop LBA 413, and is generally defined in commands by the start LBA 412 and the LBA length of namespace 410. Namespace 420 begins with a start LBA 422 and ends with a stop LBA 423, and is generally defined in commands by the start LBA 422 and the LBA length of namespace 420. Starting LBAs 412 and 422 can have the same identifier (e.g., the same number / index).

[0075] Figure 5 This is a block diagram illustrating an example method 500 for performing data replication, based on some implementation schemes. (Reference) Figure 1A , 1B In method 500, data is copied from a source device (e.g., storage device 100a) to a destination device (e.g., storage device 100b). Figure 5 Within this context, storage devices 100a and 100n are virtualization drivers. Physical driver replication is performed, causing all information corresponding to the namespace (including partition information, file systems, etc.) to be copied from the source device to the destination device. For this purpose, the source and destination devices have different namespaces due to virtualization, and the entire namespace 511 is copied to a different namespace 522. Therefore, Figure 5 The text describes a one-to-one namespace copying process on storage devices 100a and 100b, from one namespace to another.

[0076] like Figure 5 As shown, the memory array 120 of storage device 100a has a storage capacity 501, which stores data containing data corresponding to namespaces 511 and 512. The memory array 120 of storage device 100b has a storage capacity 502, which stores data containing data corresponding to namespaces 521 and 522. Data corresponding to the entire namespace 511 is copied as data corresponding to namespace 522 to the storage capacity 502 of storage device 100b. Namespaces 511 and 522 may have different identifiers (e.g., different numbers / indices, such as "namespace 1" and "namespace 2" respectively) and the same size. Namespace 511 begins with a start LBA and ends with a stop LBA, and is generally defined in commands by the start LBA and the length of the LBA of namespace 511. Namespace 522 begins with a start LBA and ends with a stop LBA, and is generally defined in commands by the start LBA and the length of the LBA of namespace 522.

[0077] Figure 6 This is a block diagram illustrating an example method 600 for performing data replication, based on some implementation schemes. (See reference) Figure 1A , 1B In method 600, data is copied from multiple source devices (e.g., storage devices 100a and 100n) to a destination device (e.g., storage device 100b). Figure 6 Within this context, the destination device can be a mass drive. As shown, LBA range 611 in namespace 610 and LBA range 621 in namespace 620 of storage device 100a are copied to LBA range 631 in namespace 630 and LBA range 641 in namespace 640 of storage device 100b, respectively. Therefore, Figure 6 It describes many-to-one LBA copying to different destination namespaces and the same offset as the source namespace.

[0078] like Figure 6As shown, the memory array 120 of storage device 100a has a storage capacity of 601, which stores data corresponding to namespace 610. The memory array 120 of storage device 100n has a storage capacity of 602, which stores data corresponding to namespace 620. The memory array 120 of storage device 100b has a storage capacity of 603, which stores data corresponding to namespaces 630, 640, and 650. Data corresponding to LBA range 611 of namespace 610 is copied to the storage capacity 603 of storage device 100b as data corresponding to LBA range 631 of namespace 630. Additionally, data corresponding to LBA range 621 of namespace 620 is copied to the storage capacity 603 of storage device 100b as data corresponding to LBA range 641 of namespace 640.

[0079] In some instances, namespaces 610 and 630 may have the same identifier (e.g., the same number / index, such as "namespace 1") and the same size. In other instances, namespaces 610 and 630 may have different identifiers (e.g., different numbers / indexes, such as "namespace 1" and "namespace 3" respectively) and the same size. LBA range 611 begins with start LBA 612 and ends with end LBA 613, and is typically defined in commands by the length of start LBA 612 and LBA range 611. LBA range 631 begins with start LBA 632 and ends with end LBA 633, and is typically defined in commands by the length of start LBA 632 and LBA range 631. Start LBAs 612 and 632 may have the same identifier (e.g., the same number / index) provided the offset is the same.

[0080] Namespaces 620 and 640 may have different identifiers (e.g., different numbers / indices, such as "Namespace 1" and "Namespace 2" respectively) and the same size. LBA range 621 begins with a starting LBA 622 and ends with a last LBA 623, and is typically defined in commands by the length of the starting LBA 622 and LBA range 621. LBA range 641 begins with a starting LBA 642 and ends with a last LBA 643, and is typically defined in commands by the length of the starting LBA 642 and LBA range 641. Under the condition of the same offset, starting LBAs 622 and 642 may have the same identifier (e.g., the same number / indication).

[0081] In some arrangements, host 101 sends a command to a destination device (e.g., storage device 100b) containing descriptors indicating the order in which LBA ranges 611 and 621 are copied to storage capacity 603. Specifically, the descriptor indicates that LBA range 611 will be copied to a namespace (e.g., namespace 630) whose number / index is lower than that of the namespace (e.g., namespace 640) to which LBA range 621 will be copied. In some instances, the command includes other descriptors, such as a descriptor indicating the address of buffer 112 of storage device 100a from which storage device 100b can transfer data corresponding to LBA range 611, a descriptor indicating the address of buffer 112 of storage device 100n from which storage device 100b can transfer data corresponding to LBA range 621, a descriptor indicating the transfer type (e.g., many-to-one LBA copy to a different destination namespace with the same offset as the source namespace), and other descriptors.

[0082] Figure 7 This is a block diagram illustrating an example method 700 for performing data replication, based on some implementation schemes. (Reference) Figure 1A , 1B In method 700, data is copied from multiple source devices (e.g., storage devices 100a and 100n) to a destination device (e.g., storage device 100b). As shown, LBA range 711 in namespace 710 and LBA range 721 in namespace 720 of storage device 100a are copied to LBA range 731 and LBA range 741 in namespace 730 of storage device 100b, respectively. Therefore, Figure 7 It describes many-to-one LBA replication to the same destination namespace with different offsets, where different LBA ranges are obtained from different source namespaces and written to a single destination namespace.

[0083] like Figure 7As shown, the memory array 120 of storage device 100a has a storage capacity 701, which stores data corresponding to namespace 710. The memory array 120 of storage device 100n has a storage capacity 702, which stores data corresponding to namespace 720. The memory array 120 of storage device 100b has a storage capacity 703, which stores data corresponding to namespace 730. Data corresponding to LBA range 711 of namespace 710 is copied to storage capacity 703 of storage device 100b as data corresponding to LBA range 741 of namespace 730. Additionally, data corresponding to LBA range 721 of namespace 720 is copied to storage capacity 703 of storage device 100b as data corresponding to LBA range 731 of namespace 730. In other words, data corresponding to LBA ranges 711 and 721 are copied to the same namespace 730.

[0084] In some instances, namespaces 710, 720, and 730 may have the same identifier (e.g., the same number / index, such as "namespace 1") and the same size. LBA range 711 begins with a starting LBA 712 and ends with a last LBA 713, and is typically defined in commands by the length of the starting LBA 712 and LBA range 711. LBA range 721 begins with a starting LBA 722 and ends with a last LBA 723, and is typically defined in commands by the length of the starting LBA 722 and LBA range 721. LBA range 731 begins with a starting LBA 732 and ends with a last LBA 733, and is typically defined in commands by the length of the starting LBA 732 and LBA range 731. LBA range 741 begins at start LBA 742 and ends at end LBA 743, and is typically defined in a command by the length of the LBA between start LBA 742 and LBA range 741.

[0085] As shown in the figure, after the copy operation, the starting LBA 742 of LBA range 741 immediately follows the ending LBA 733 of LBA range 731. Therefore, within the same namespace 730, data from LBA range 711 is concatenated with data from LBA range 721. Consequently, the offset of the starting LBA 722 is different from the offset of the starting LBA 742.

[0086] In some arrangements, host 101 sends a command to a destination device (e.g., storage device 100b) containing a descriptor indicating the order in which LBA ranges 671 and 721 are copied to storage capacity 703. Specifically, the descriptor indicates that LBA range 711 will be copied to an LBA range (e.g., LBA range 741) immediately following its starting LBA 742 after the last LBA 733 of the LBA range to which LBA range 721 will be copied (e.g., LBA range 731). In some instances, the command includes other descriptors, such as a descriptor indicating the address of buffer 112 of storage device 100a from which storage device 100b can transfer data corresponding to LBA range 711, a descriptor indicating the address of buffer 112 of storage device 100n from which storage device 100b can transfer data corresponding to LBA range 721, a descriptor indicating the transfer type (e.g., many-to-one LBA copy with the same namespace and different offsets, concatenation), and other descriptors.

[0087] Figure 8 This is a block diagram illustrating an example method 800 for performing data replication, based on some implementation schemes. (Reference) Figure 1A , 1B In method 800, data is copied from multiple source devices (e.g., storage devices 100a and 100n) to a destination device (e.g., storage device 100b). As shown, the entire namespace 810 of storage device 100a and the LBA range 821 in namespace 820 of storage device 100n are copied to namespace 840 and namespace 830 of storage device 100b, respectively. Therefore, Figure 8 It describes many-to-one LBA copying and namespace copying with different offsets.

[0088] like Figure 8 As shown, the memory array 120 of storage device 100a has a storage capacity 801, which stores data corresponding to namespace 810. The memory array 120 of storage device 100n has a storage capacity 802, which stores data corresponding to namespace 820. The memory array 120 of storage device 100b has a storage capacity 803, which stores data corresponding to namespaces 830 and 840. Data corresponding to the entire namespace 810 is copied as data corresponding to namespace 840 to the storage capacity 803 of storage device 100b. Additionally, the LBA range 821 corresponding to namespace 820 is copied as data corresponding to LBA range 831 of namespace 830 to the storage capacity 803 of storage device 100b.

[0089] In some instances, namespaces 810, 820, and 830 may have the same identifier (e.g., the same number / index, such as "namespace 1"). In some instances, namespaces 810 and 840 may have different identifiers (e.g., different numbers / indexes, such as "namespace 1" and "namespace 2" respectively) and the same size. Namespace 810 begins with a starting LBA 812 and ends with a ending LBA 813, and is typically defined in commands by the length of the starting LBA 812 and the LBA of namespace 810. LBA range 821 begins with a starting LBA 822 and ends with a ending LBA 823, and is typically defined in commands by the length of the starting LBA 822 and the LBA of range 821. LBA range 831 begins with a starting LBA 832 and ends with a ending LBA 833, and is typically defined in commands by the length of the starting LBA 832 and the LBA of range 831. Namespace 840 begins with start LBA 842 and ends with end LBA 843, and is typically defined in commands by the length of the LBAs described in start LBA 842 and namespace 840. Start LBAs 822 and 832 may have different identifiers (e.g., different numbers / indexes) under different offset conditions.

[0090] In some arrangements, host 101 sends a command to a destination device (e.g., storage device 100b) containing a descriptor indicating the order in which namespace 810 and LBA range 821 are copied to storage capacity 803. Specifically, the descriptor indicates that LBA range 821 will be copied to a namespace (e.g., namespace 830) whose number / index is lower than that of the namespace to which namespace 810 will be copied (e.g., namespace 840). In some instances, the command includes other descriptors, such as a descriptor indicating the address of buffer 112 of storage device 100a from which storage device 100b can transfer data corresponding to namespace 810, a descriptor indicating the address of buffer 112 of storage device 100n from which storage device 100b can transfer data corresponding to LBA range 821, a descriptor indicating the transfer type (e.g., many-to-one LBA copy with different namespaces and the same offset), and other descriptors.

[0091] In the arrangement disclosed herein, the copy offloading mechanism uses buffer 112, which may be a persistent (i.e., non-volatile) CMB buffer for transferring data. In some embodiments, the persistent CMB buffer may be allocated in a persistent memory region (PMR), which may be implemented using non-volatile byte-addressable memory, such as, but not limited to, phase-change memory (PCM), magnetic random access memory (MRAM), or resistive random access memory (ReRAM). Buffer 112 may be shared by host 101 and may be accessed by peer devices (e.g., other storage devices 100) for the copy operations disclosed herein. In addition to accessibility, the copy offloading mechanism disclosed herein also uses buffer 112 to copy data, much like a sliding window.

[0092] to this end, Figure 9A This is a block diagram illustrating an example method 900a for performing data copying using a sliding window mechanism, according to some implementation schemes. In method 900a, a buffer 112 of the source device and a write buffer 114 of the destination device are used.

[0093] like Figure 9A As shown, data corresponding to LBA range 911 stored in namespace 910 on the source device (in this context, storage device 100a) is copied via method 900a to the destination device (in this context, storage device 100b) as LBA range 921 in namespace 920. Due to bandwidth and processing power limitations of the source device, destination device, and interface 140, data corresponding to LBA range 911 is transferred through multiple (n) different transfers (from the first transfer to the nth transfer). The first transfer is performed before the second transfer, and so on, with the nth transfer being the last transfer. In each transfer, data blocks corresponding to the buffer size of each of buffers 915, 916, 917, 925, 926, and 927 are transferred from the source device to the destination device.

[0094] For example, in the first transfer, the source device reads data corresponding to the initial LBA 912 from a physical location on memory array 120 into a persistent memory region (PMR) buffer 915. The PMR buffer 915 is an instance of a buffer 154 dynamically allocated in the data region 152 of the source device. Data transfers (e.g., direct memory access (DMA) transfers, non-DMA transfers, etc.) are performed by the destination device to transfer data from the PMR buffer 915 to the write buffer 925 of the destination device via interface 140 in a P2P read operation. The write buffer 925 is an instance of a buffer dynamically allocated in the write buffer 114 of the destination device. The destination device then writes the data from the write buffer 925 to the physical location on the memory array 120 of the destination device corresponding to the initial LBA 922.

[0095] In each of the second through (n-1)th transfers, the source device reads the data corresponding to the recipient in intermediate LBA 913 from a physical location on memory array 120 into the recipient in PMR buffer 916. Each PMR buffer 916 is an instance of buffer 156 dynamically allocated in data region 152 of the source device. Data transfers (e.g., DMA transfers, non-DMA transfers, etc.) are performed by the destination device to transfer data from each PMR buffer 916 to the recipient in write buffer 926 of the destination device via interface 140 during a P2P read operation. Each write buffer 926 is an instance of buffer dynamically allocated in write buffer 114 of the destination device. The destination device then writes the data from each write buffer 926 to the physical location on memory array 120 of the destination device corresponding to the recipient in intermediate LBA 923.

[0096] In the final transfer, the source device reads data corresponding to the end LBA 914 from a physical location on memory array 120 into PMR buffer 917. PMR buffer 917 is an instance of buffer 158 dynamically allocated in data region 152 of the source device. Data transfer (e.g., DMA transfer, non-DMA transfer, etc.) is performed by the destination device to transfer data from PMR buffer 917 to write buffer 927 of the destination device via interface 140 in a P2P read operation. Write buffer 927 is an instance of buffer dynamically allocated in write buffer 114 of the destination device. The destination device then writes data from write buffer 927 to the physical location on memory array 120 of the destination device corresponding to the end LBA 924.

[0097] Figure 9BThis is a block diagram illustrating an example method 900b for performing data copying using a sliding window mechanism, according to some implementation schemes. In method 900b, a buffer 112 for the destination device and a read buffer 116 for the source device are used.

[0098] like Figure 9B As shown, data corresponding to LBA range 911 stored in namespace 910 on the source device (in this context, storage device 100a) is copied to the destination device (in this context, storage device 100b) as LBA range 921 in namespace 920 via method 900b. Due to bandwidth and processing power limitations of the source device, destination device, and interface 140, data corresponding to LBA range 911 is transferred through multiple (n) different transfers (the first transfer is to the nth transfer). The first transfer is performed before the second transfer, and so on, with the nth transfer being the last transfer. In each transfer, data blocks corresponding to the buffer size of each of buffers 935, 936, 937, 945, 956, and 957 are transferred from the source device to the destination device.

[0099] For example, in the first transfer, the source device reads data corresponding to the starting LBA 912 from a physical location on the memory array 120 into a read buffer 935. The read buffer 935 is an example of a dynamically allocated buffer within the source device's read buffer 116. Data transfers (e.g., DMA transfers, non-DMA transfers, etc.) are performed by the source device 100a to transfer data from the read buffer 935 to the destination device's PMR buffer 945 via interface 140, for example, in a P2P write operation. The PMR buffer 945 is an example of a dynamically allocated buffer 164 within the destination device's data region 162. The destination device then writes data from the PMR buffer 945 to a physical location on the destination device's memory array 120 corresponding to the starting LBA 922.

[0100] In each of the second to (n-1)th transfers, the source device reads the data corresponding to the recipient in intermediate LBA 913 from a physical location on memory array 120 into the recipient in read buffer 936. Each read buffer 936 is an instance of a dynamically allocated buffer in read buffer 116 of the source device. Data transfers (e.g., DMA transfers, non-DMA transfers, etc.) are performed by source device 100a to transfer data from each read buffer 936 to a recipient in PMR buffer 946 of the destination device via interface 140, for example, in a P2P write operation. Each PMR buffer 946 is an instance of a dynamically allocated buffer 166 in data area 162 of the destination device. The destination device then writes the data from each PMR buffer 946 to the physical location on memory array 120 of the destination device corresponding to the recipient in intermediate LBA 923.

[0101] In the final transfer, the source device reads data corresponding to the end LBA 914 from a physical location on memory array 120 into read buffer 937. Read buffer 937 is an instance of a dynamically allocated buffer in read buffer 116 of the source device. Data transfer (e.g., DMA transfer, non-DMA transfer, etc.) is performed by the source device to transfer data from read buffer 937 to PMR buffer 947 of the destination device via interface 140, for example, in a P2P write operation. PMR buffer 947 is an instance of a dynamically allocated buffer 168 in data area 162 of the destination device. The destination device then writes data from PMR buffer 947 to a physical location on memory array 120 of the destination device corresponding to the end LBA 924. Buffers 915-917 can be allocated in a sliding window, meaning that some of buffers 915-917 are allocated in a first time window in response to a P2P message that temporarily stores data blocks for transfer. When the transfer of these data blocks to the destination device is complete, those buffers are deallocated or recycled, and the associated memory capacity can be freely allocated as additional buffers to buffers 915-917. Buffers 935-937 can be allocated similarly.

[0102] Similarly, buffers 925-927 can be allocated in a sliding window, meaning that some of buffers 925-927 are allocated in a first time window to receive the first data chunks from the source device. Once the writing of these data chunks to the memory array 120 of the destination device is complete, those buffers are deallocated or recycled, and their associated memory capacity is freely allocated as additional buffers for buffers 925-927. Buffers 945-947 can be allocated similarly.

[0103] In this instance copy operation, data (e.g., 1GB) is copied from a source device (e.g., storage device 100a) to a destination device (e.g., storage device 100b). In this instance, both the source and destination devices use a buffering mechanism (e.g., buffer 112) with a specific buffer size (e.g., 128KB) for each dynamically allocated buffer (e.g., each of buffers 154, 156, 158, 164, 166, and 168 and each of PMR buffers 915, 916, 917, 925, 926, and 927). Therefore, the total data to be copied is divided into 8192 chunks and will be used 8192 times (e.g., in...). Figure 9A and Figure 9B In the case of n being 8192), a transfer is made from the source device to the destination device, with each block or transfer corresponding to 128KB of data.

[0104] Traditionally, to perform such a copy operation, host 101 submits a read request (e.g., an NVMe read request) for a data block (128KB) to the source device (e.g., for a "Sq Rd" operation). In response to the read request, the source device performs a read operation to read the requested data from memory array 120 into read buffer 116 (e.g., in an "NVM Rd" operation). The data is then transferred from read buffer 116 via interface 140 (e.g., an NVMe interface) (e.g., in a "PCIe Rd" operation). Next, the data is transferred via bus 106. For example, the data is transferred via the root complex (e.g., in an "RC Rd" operation) and via the memory bus (e.g., in a "Mem Rd" operation) to be stored in the memory buffer of memory 102. Next, host 101 submits a write request (e.g., an NVMe write request) to the destination device (e.g., in a "Sq Wt" operation). In response to the write request, the destination device causes data to be read from the memory buffer of memory 102 and transferred via bus 106. For example, data is transferred via the memory bus (e.g., in the "Mem Wt" operation) and via the root compound (e.g., in the "RC Wt" operation). The data is then transferred to the write buffer 114 of the destination device via interface 140 (e.g., an NVMe interface) (e.g., in the "PCIe Wt" operation). Next, the destination device performs a write operation to write the data from the write buffer 114 to the memory array 120 (e.g., in the "NVM Wt" operation). This process (comprising 12 operations) is repeated for each block until all 8192 blocks have been transferred to the memory array 120 of the destination device. Therefore, to transfer 1GB of data, a total of 98,304 operations are typically required. In other words, to process each block, a read operation has three separate transactions (including a command phase / transaction in which a read command is sent from host 101 to the source device, a read phase / transaction in which data is read from the memory array 120 of the source device, and a data transfer phase / transaction in which data is transferred from the source device to the memory 102 of host 101), and a write operation has three separate transactions (including a command phase / transaction in which a write command is sent from host 101 to the destination device, a data transfer phase / transaction in which data is transferred from the memory 102 of host 101 to the destination device, and a write phase / transaction in which data is written to the memory array 120 of the destination device). Each operation / transaction consumes bandwidth at the system level or CPU cycles of host 101.

[0105] Figure 10 This is a block diagram illustrating an example method 1000 for performing data replication, based on some implementation schemes. (Reference) Figure 1A-10Method 1000 is performed by the destination device (e.g., storage device 100b) and the source device 100a. Blocks 1010, 1020, 1030, 1040, and 1050 are performed by the destination device. Blocks 1025, 1035, 1045, and 1055 are performed by the source device. For example, communication between the source and destination devices at blocks 1020 / 1025 and 1030 / 1045 is performed via interface 140 or another suitable communication channel between the source and destination devices without routing to host 101.

[0106] At 1010, the destination device (e.g., controller 110) receives a replication request from host 101 via interface 140. For example, host 101 submits a replication request (e.g., a command, such as a P2P NVMe replication request or another suitable request / command) to the destination device. In some instances, this replication request is not a standard NVMe write request, but a new type of NVMe request to the destination device, instructing it to initiate a P2P data transfer and directly acquire the data chunks to be copied from the source device. In other instances, the behavior of the destination device triggered by the replication request is similar to that triggered by a write command pointing to the buffer address of buffer 112 on the source device rather than the memory 102 of host 101. In such instances, the destination device may not even be able to distinguish whether the buffer address is on another storage device or on host 101.

[0107] In some arrangements, the copy request includes descriptors or parameters that define or identify the data copying method. For example, the copy request may include one or more descriptors identifying the address of buffer 112 on the source device, at least one namespace of the data on the source device (referred to as the “source namespace”), at least one logical base address (LBA) of the data on the source device (referred to as the “source start LBA”), and the data byte length B of the data. In instances where host 101 instructs the destination device to copy data from multiple source devices (e.g., via many-to-one copy operations of methods 600, 700, and 800), the copy request may include such descriptors for each source device from which data will be transferred (e.g., each of storage devices 100a and 100n). The address of buffer 112 refers to the address of the P2P reserved area 150 of buffer 112 on the source device.

[0108] Instances of source namespaces include, but are not limited to, namespaces 210, 310, 410, 511 and 512, 610, 620, 710, 720, 810, 820, etc. Instances of source start LBAs include, but are not limited to, start LBA 212, start LBA 312, start LBA 412, start LBA of namespace 511, start LBA of namespace 512, start LBA 612, start LBA 622, start LBA 712, start LBA 722, start LBA 812, start LBA 822, etc.

[0109] In some arrangements, the replication request further includes one or more descriptors identifying at least one namespace (referred to as the "destination namespace") and at least one starting logical address (LBA) of the data on the destination device (referred to as the "destination start LBA"). The data byte length B of the data on the source device and the destination device is the same. Instances of the destination namespace include, but are not limited to, namespaces 220, 320, 420, 521 and 522, 630, 640, 730, 740, 830, 840, etc. Instances of the destination start LBA include, but are not limited to, start LBA 222, start LBA 322, start LBA 422, the start LBA of namespace 521, the start LBA of namespace 522, start LBA 632, start LBA 642, start LBA 732, start LBA 742, start LBA 832, start LBA 842, etc.

[0110] Therefore, a replication request contains logical address information for both the source and destination devices. For each replication operation, the descriptor included in the replication request is called a source-destination pair descriptor. As described herein, for each replication operation, there is one destination device and one or more source devices. Including information for multiple source devices allows multiple replication operations to be performed for the same destination device as part of a single replication request.

[0111] At 1020, the destination device communicates with the source device to set the data identified in the copy request to fill the buffer 112 of the source device. At 1025, the source device communicates with the destination device to set the data identified in the copy request to fill the buffer 112 of the source device. Specifically, the destination device can identify the data destined for the source device by sending the source namespace, source start LBA, and data length (their descriptors are contained in the copy request received at 1010) to the source device. In instances where multiple source devices are identified in the copy request, the destination device communicates with each source device in the manner described above to set the data identified in the copy request to fill the buffer 112 of each source device.

[0112] For example, in response to receiving a copy request at 1010, the destination device sends a P2P message to the source device. Therefore, box 1020 includes sending a message from the destination device to the source device, and box 1025 includes receiving a message from the source device to the destination device. The message is sent directly to the source device via interface 140 without being routed through host 101. The message includes at least the source namespace, the source starting LBA, and the remaining data length (e.g., the remaining number of data chunks to be transferred to the destination device).

[0113] In some implementations, the destination device sends the message to the source device by performing a write operation that writes the P2P message to the address of the P2P reserved area 150 of the buffer 112 of the source device. As described, the message contains at least the source namespace, the source start LBA, and the remaining data length (e.g., the remaining number of data blocks to be transferred to the destination device). In some instances, the message further contains the address of the P2P reserved area 160 of the buffer 112 of the destination device, such that the destination device can send a message, a response, and an acknowledgment to the source device by performing a write operation that writes the message to the P2P reserved area 160 of the buffer 112 of the destination device.

[0114] In other implementations, the P2P message is a regular NVMe read command sent from the destination device to the source device. In this case, the destination device temporarily switches to operate as the NVMe initiator, and the source device acts as the NVMe target device until the NVMe read command is completed.

[0115] In response to receiving this message, the source device temporarily stores one or more data blocks corresponding to the source namespace, source start LBA, and remaining data length into the source device's buffers. For example, at 1035, the source device reads one or more data blocks from the source device's memory array 120 into one or more buffers. If no data blocks are transferred to the destination device, then the remaining data length will be the entire data length.

[0116] In some instances, as described with respect to method 900a, the source device, for example, uses a sliding window mechanism (e.g., in the current window) to read one or more blocks of data from the source device's memory array 120 into one or more of the source device's PMR buffers 915-917 (e.g., buffers 154-158). Within each window, one or more of the PMR buffers 915-917 are filled.

[0117] In other instances, as described with respect to method 900b, the source device uses a sliding window mechanism (e.g., in the current window) to read one or more blocks of data from the source device's memory array 120 into one or more read buffers 935-937 (e.g., read buffer 116) of the source device. Within each window, one or more of the read buffers 935-937 are filled.

[0118] At 1030, the destination device performs a transfer with the source device to transfer each data block from the corresponding buffer of the source device to the buffer of the destination device. At 1045, the source device performs a transfer with the destination device to transfer each data block from the corresponding buffer of the source device to the buffer of the destination device.

[0119] In some instances, as described with respect to method 900a, the destination device may perform a P2P read operation to read each data block from its counterpart in PMR buffers 915-917 (which are already filled in the current window) into its counterpart in the destination device's write buffers 925-927. For example, in response to determining that one or more of the PMR buffers 915-917 are filled in the current window, the source device sends a P2P response message to the destination device (e.g., by writing a P2P message to the P2P reservation area 160 or by sending a regular NVMe command). The P2P response message contains the address of the one or more of the PMR buffers 915-917 that are already filled in the current window. The P2P response message further indicates the number of blocks ready to be transferred in the current window (also the number of the one or more of the PMR buffers 915-917). The destination device can then perform a P2P read operation using the address of each of the one or more (currently filled) PMR buffers 915-917 to read each data block from the corresponding one of the one or more (currently filled) PMR buffers 915-917 into the corresponding one of the write buffers 925-927 of the destination device. In some instances, the destination device sends another P2P message to the source device (e.g., by writing a P2P message to P2P reserved area 160 or by sending a regular NVMe command), where such a message indicates that the destination device has retrieved the data block and updated the remaining data blocks to be transferred.

[0120] In some instances, as described with respect to method 900b, the source device may perform a P2P write operation to write each data block from its corresponding read buffer 935-937 (which is already filled in the current window) to its corresponding PMR buffer 945-947 in the destination device. For example, in response to determining that one or more of the read buffers 935-937 are filled in the current window, the source device sends a P2P response message to the destination device (e.g., by writing a P2P message to the P2P reservation area 160 or by sending a regular NVMe command). The P2P response message contains the number of blocks ready to be transferred in the current window (which is also the number of said one or more read buffers 935-937). In response, the destination device allocates the same number of PMR buffers 945-947 as the number of blocks ready to be transferred. Furthermore, the destination device sends another P2P response message to the source device (e.g., by writing a P2P message to P2P reserved area 150 or by sending a regular NVMe command), wherein such a P2P response message contains the address of the allocated recipient in PMR buffers 945-947. The source device can then perform a P2P write operation using the address of each of the one or more of the PMR buffers 945-947 to write each data block from the corresponding recipient in the one or more read buffers 935-937 (which are already filled in the current window) to the corresponding recipient in the allocated recipient of the PMR buffers 945-947 of the destination device.

[0121] At 1040, in response to the completion of the transfer of each data block in the determination window to the destination device's buffer (e.g., write buffer or PMR buffer), the destination device performs an NVM write operation to write data from the destination device's buffer to the destination device's memory array 120.

[0122] Boxes 1035, 1045, and 1030 are operations performed by the source and destination devices for a window. Therefore, at 1050, in response to the destination device determining that all data blocks have been transferred to the destination device's memory array 120 (1050: Yes), method 1000 terminates for the destination device, and any write buffer or PMR buffer is released by the destination device. Conversely, in response to the destination device determining that not all data blocks have been transferred to the destination device's memory array 120 (1050: No), method 1000 returns for the destination device to box 1030, where one or more data blocks for subsequent windows are transferred in the manner described herein.

[0123] Similarly, at 1055, in response to the source device determining that all data blocks have been transferred to the destination device (1055: Yes), method 1000 terminates for the source device, and any read buffer or PMR buffer is released by the source device. On the other hand, in response to the source device determining that not all data blocks have been transferred to the destination device (1055: No), method 1000 returns to box 1035 for the destination device, where one or more data blocks in subsequent windows are transferred in the manner described herein. In one instance, in response to receiving a P2P message from the source device indicating that 0 data blocks are pending transfer, the source device may determine that all data blocks have been transferred to the destination device.

[0124] In some instances, the destination device polls the source device for the status, indicating whether data is ready to be transferred or a read request has been completed. In other instances, polling can be replaced by an interrupt-based architecture that simulates an SQ / CQ pair model, where the destination device interrupts once data is ready. In some instances, the destination device sends a message request to the source device for the data chunks to be transferred from the source device. After processing the message request, the source device reports the status to the destination device via a message response. If at least one chunk remains, message requests and reports continue to be sent until the destination device determines that all chunks have been transferred from the source device. The final message from the destination device contains a transfer length of zero, informing the source device that no other chunks need to be transferred.

[0125] Method 1000 can be performed between the destination device and each source device in each source-destination pair descriptor of a replication request received by the destination device from the host 101. After completing the processing of all transfers identified in all source-destination pair descriptors, the destination device completes the replication request.

[0126] In method 1000, assuming each data block or buffer is 128KB in size and the total data size is 1GB, the source and destination devices perform three operations for each data block. For a total of 8,192 blocks, 24,576 operations plus two additional operations are performed (e.g., a total of 24,578 operations). In other words, method 1000 reduces the number of operations by 75%. Given that each operation consumes system bandwidth, CPU cycles, and other resources, method 1000 improves the efficiency of data copying without increasing hardware costs.

[0127] In some instances where a power outage occurs during a copy operation, in the destination device driving method disclosed herein, the destination device stores an indicator showing how far the copy operation has progressed for resumption after the power outage and subsequent power-on. Therefore, the destination device treats the P2P reserved area 160 as a power-loss protected (PLP) area and refreshes or saves its contents to memory 120 in response to a power outage. In response to a determination of power-on, the contents are restored to the P2P reserved area 160, and the copy operation resumes from the last stopped position. In some instances, the source device is stateless because it may lose all previous information related to the copy operation (e.g., how much data has been transferred). Therefore, if the destination device does not acknowledge transferred chunks temporarily stored by the source device due to a power outage or even interruption, the destination device will request the data chunks again after power-on, and the transfer operation can resume from that point. The source device does not need to remember that the destination device has not acknowledged them, nor does it need to temporarily store unacknowledged data chunks itself.

[0128] Although method 1000 corresponds to a destination device-driven method, other methods may be driven by the source device. In other words, other methods involve the source receiving a copy request from host 101 and initiating communication with the destination device to transfer data from the source device to the destination device.

[0129] In one implementation, host 101 sends a copy request to the source device via interface 140. The copy request specifies the source namespace, source start LBA, and data byte length B, the address of the destination device (e.g., the address of the P2P reservation area 160 of buffer 112 of the destination device), the destination namespace, and the destination start LBA. For example, in the sliding window described herein, the source device reads data corresponding to the source namespace, source start LBA, and data byte length B from memory array 120 into dynamically allocated buffers 154-158 of buffer 152. The source device then sends a write command to the destination device or sends a P2P message to the destination device (e.g., by writing a P2P message to the address of the P2P reservation area 160 of buffer 112 of the destination device), the P2P message specifying the destination namespace, destination start LBA, and the address of the allocated device in buffers 154-158. The destination device then performs a P2P read operation to read the data stored in the allocated device in buffers 154-158 into the write buffer of the destination device. The destination device then writes the data from its write buffer to its memory array 120. This operation can be performed for subsequent sliding windows until all data blocks have been transferred.

[0130] In one implementation, host 101 sends a copy request to the source device via interface 140. The copy request specifies the source namespace, source start LBA, and data byte length B, the address of the destination device (e.g., the address of the P2P reservation area 160 of buffer 112 of the destination device), the destination namespace, and the destination start LBA. For example, in the sliding window described herein, the source device reads data corresponding to the source namespace, source start LBA, and data byte length B from memory array 120 into a dynamically allocated read buffer of read buffer 116. Next, the source device performs a P2P write operation to transfer data from each allocated buffer in the source device's read buffer to buffers 164-168 of the destination device. For example, the source device may send a P2P message to the destination device indicating the number of buffers 164-168 to be allocated for the sliding window (e.g., by writing a P2P message to the address of the P2P reservation area 160 of buffer 112 of the destination device). The destination device may send a P2P message response to the source device indicating the address of the allocator in buffers 164-168 (e.g., by writing a P2P message to the address of the P2P reserved area 150 in the source device's buffer 112). The source device may perform a P2P write operation to transfer data from each allocator in the source device's read buffer to an allocator in the destination device's buffers 164-168. The destination device then writes the data from its write buffer to its memory array 120. In some instances, the source device may send a write command to the destination device specifying the allocator in buffers 164-168 as the location where data is to be written to memory array 120. Such operations may be performed against a subsequent sliding window until all data blocks have been transferred. The preceding description is provided to enable those skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects. Therefore, the appended claims are not intended to be limited to the aspects shown herein, but should be given the full scope consistent with the language of the appended claims, wherein the reference to an element in the singular does not mean "one and only one" (unless specifically stated so), but rather "one or more". Unless specifically stated so, the term "some" refers to one or more. All structural and functional equivalents of elements known or to be known hereafter by those skilled in the art throughout the various aspects described herein are expressly incorporated herein by reference and are intended to be covered by the appended claims. Furthermore, nothing disclosed herein is intended to be exclusive to the public, whether or not such disclosure is expressly stated in the claims. No claim element will be construed as a component plus a function unless the element is expressly stated using the phrase "component for..."

[0131] It should be understood that the specific order or hierarchy of steps in the disclosed process is an example of an illustrative method. It should be understood that the specific order or hierarchy of steps in the process may be rearranged based on design preferences, while remaining within the scope of the previously described method. The accompanying method claims present the elements of the various steps in an exemplary order, but are not intended to limit them to the specific order or hierarchy presented.

[0132] The prior description of the disclosed embodiments is provided to enable those skilled in the art to make or use the disclosed subject matter. Those skilled in the art will readily understand various modifications to these embodiments, and the general principles defined herein may be applied to other embodiments without departing from the spirit or scope of the prior description. Therefore, the prior description is not intended to limit it to the embodiments shown herein, but is to be endowed with the widest scope consistent with the principles and novel features disclosed herein.

[0133] The various examples illustrated and described are provided by way of example only to illustrate the various features of the appended claims. However, the features shown and described with respect to any given example are not necessarily limited to the associated example and may be used or combined with other examples shown and described. Furthermore, the claims are not intended to be limited to any single example.

[0134] The foregoing method descriptions and process flowcharts are provided merely as illustrative examples and are not intended to require or imply that the steps in the various examples must be performed in the presented order. As those skilled in the art will understand, the order of the steps in the foregoing examples can be performed in any order. For example, words such as "after," "following," and "next" are not intended to limit the order of steps; these words are merely used to guide the reader through the description of the method. Furthermore, any reference to singular claim elements, for example, using the articles "a," "an," or "the," should not be construed as limiting said element to the singular.

[0135] The various illustrative logic blocks, modules, circuits, and algorithmic steps described in conjunction with the examples disclosed herein can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability between hardware and software, the functionality of the various illustrative components, blocks, modules, circuits, and steps has been described above in general terms. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the system as a whole. Those skilled in the art can implement the described functionality in different ways for each specific application, but such implementation decisions should not be construed as causing a deviation from the scope of this disclosure.

[0136] The hardware used to implement the various illustrative logics, logic blocks, modules, and circuits described in conjunction with the examples disclosed herein can be implemented or executed using a general-purpose processor, DSP, ASIC, FPGA, or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. The general-purpose processor may be a microprocessor, but alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors incorporating a DSP core, or any other such configuration. Alternatively, some steps or methods may be performed by a circuit system specific to a given function.

[0137] In some exemplary instances, the functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored as one or more instructions or codes on a non-transitory computer-readable storage medium or a non-transitory processor-readable storage medium. The steps of the methods or algorithms disclosed herein may be embodied in a processor-executable software module that may reside on a non-transitory computer-readable or processor-readable storage medium. A non-transitory computer-readable or processor-readable storage medium may be any storage medium accessible by a computer or processor. For example, but not limitingly, such non-transitory computer-readable or processor-readable storage media may include RAM, ROM, EEPROM, flash memory, CD-ROM or other optical drive storage devices, magnetic drive storage devices or other magnetic storage devices, or any other medium that can be used to store desired program code in the form of instructions or data structures and is accessible by a computer. As used herein, drives and optical discs include compact optical discs (CDs), laser optical discs, optical discs, digital versatile optical discs (DVDs), floppy drives, and Blu-ray discs, wherein drives typically reproduce data magnetically, while optical discs reproduce data optically using lasers. Combinations of the foregoing are also included within the scope of non-transitory computer-readable and processor-readable media. Furthermore, the operation of a method or algorithm may reside, as a whole or in any combination or as a set of code and / or instructions, on non-transitory processor-readable and / or computer-readable storage media, which may be incorporated into a computer program product.

[0138] The prior description of the disclosed examples is provided to enable those skilled in the art to make or use this disclosure. Various modifications to these examples will readily be apparent to those skilled in the art, and the general principles defined herein may be applied to some examples without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not intended to be limited to the examples shown herein, but should be given the widest scope consistent with the appended claims and the principles and novel features disclosed herein.

Claims

1. A method for transferring data from a source device to a destination device, comprising: The destination device receives a copy request from the host. The destination device sends a read request to the source device. The destination device requests the status of the source device, wherein the request includes a message from the destination device indicating to the source device the amount of data transferred or the completion of the read request. The data is transferred from the buffer of the source device to the buffer of the destination device by performing a transfer with the source device. as well as The data is written to the non-volatile storage device of the destination device via the destination device.

2. The method of claim 1, wherein the copy request is defined as one or more of the following: At least one namespace for the data on the source device; At least one starting logical address of the data on the source device; The length of the data; At least one namespace for the data on the destination device; and At least one starting logical address of the data on the destination device.

3. The method of claim 2, further comprising communicating with the source device via the destination device to set the filling of the buffer of the source device, wherein the source device fills the buffer with data blocks, and one or more of the data blocks are transferred simultaneously.

4. The method according to claim 3, wherein Communicating with the source device to set the filling of the buffer of the source device includes sending a message to the source device via the destination device; and The message includes at least one namespace of the data on the source device, at least one starting logical address of the data on the source device, and the length of the data.

5. The method according to claim 4, wherein The copy request further defines the address of the reserved area of ​​the buffer in the destination device; and Sending the message to the source device includes writing the message to the reserved area of ​​the buffer of the destination device using the address of the reserved area.

6. The method according to claim 1, wherein The buffer of the source device includes a buffer that the destination device can access; The buffer of the destination device includes a write buffer that is inaccessible to the source device; and Performing the transfer with the source device to transfer the data from the buffer of the source device to the buffer of the destination device includes performing a read operation through the destination device to read the data in chunks from each of the buffers accessible by the destination device into corresponding entries in the write buffers inaccessible by the source device.

7. The method of claim 6, wherein The buffer that the destination device can access is either a controller memory buffer (CMB) or a persistent memory region (PMR) buffer; and The destination device includes the non-volatile storage device and a controller operatively coupled to the non-volatile storage device.

8. The method according to claim 1, wherein The buffer of the source device includes a read buffer that is inaccessible to the destination device; The buffer of the destination device includes a buffer that the source device can access; and Performing the transfer with the source device to transfer the data from the buffer of the source device to the buffer of the destination device includes receiving, via a write operation performed by the source device, the data chunks in each of the buffers accessible by the source device from the corresponding read buffers that are inaccessible to the destination device.

9. A destination device for data replication, comprising: buffer; Non-volatile storage devices; as well as The controller is configured as follows: Receive a replication request from the host; Send a read request to the source device; The status of the source device is requested, wherein the request includes a message from the destination device indicating to the source device the amount of data transferred or the completion of the read request; Perform a transfer with the source device to transfer data from the buffer of the source device to the buffer of the destination device; as well as The data is written to the non-volatile storage device.

10. A non-transitory processor-readable medium comprising processor-readable instructions that, when executed by at least one processor of a controller of a destination device, cause the processor to perform the following operations: Receive replication requests from the host; Send a read request to the source device; The request requests the status of the source device, wherein the request includes a message from the destination device indicating to the source device the amount of data transferred or the completion of the read request. Perform a transfer with the source device to transfer data from the buffer of the source device to the buffer of the destination device; and The data is written to a non-volatile storage device.

11. A method for transferring data from a source device to a destination device, comprising: The source device communicates with the destination device to set the filling of the buffer of the source device; as well as The data is transferred from the buffer of the source device to the buffer of the destination device by performing a transfer between the source device and the destination device. The buffer of the source device includes a buffer accessible by the destination device, and the buffer of the destination device includes a write buffer inaccessible by the source device. The transfer performed with the source device to transfer the data from the buffer of the source device to the buffer of the destination device includes transferring blocks of the data in each of the buffers accessible by the destination device to the corresponding write buffers inaccessible by the source device via a read operation performed by the destination device.

12. The method of claim 11, wherein Communicating with the destination device to set the filling of the buffer of the source device includes receiving messages from the destination device via the source device; and The message includes at least one namespace of the data on the source device, at least one starting logical address of the data on the source device, and the length of the data.

13. The method of claim 12, wherein the message is written to the reserved area using the address of a reserved area of ​​the buffer of the destination device.

14. The method of claim 11, wherein the copy request from the host includes a source descriptor and a destination descriptor specifying the source device and the destination device.

15. The method of claim 14, wherein the buffer accessible by the destination device is a controller memory buffer (CMB) or a persistent memory region (PMR) buffer.

16. The method of claim 11, further comprising reading each data block from the non-volatile storage device of the source device into one of the buffers of the source device via the source device.

17. The method of claim 16, wherein The buffer of the source device includes a buffer that the destination device can access; or The buffer of the source device includes a read buffer that is inaccessible to the destination device.

18. A source device for data replication, comprising: buffer; Non-volatile storage devices for storing data; as well as The controller is configured as follows: Communicate with the destination device to set the filling of the buffer of the source device; as well as Perform a transfer to the destination device to transfer the data from the buffer of the source device to the buffer of the destination device. The buffer of the source device includes a buffer accessible by the destination device, and the buffer of the destination device includes a write buffer inaccessible by the source device. The transfer performed with the source device to transfer the data from the buffer of the source device to the buffer of the destination device includes transferring blocks of the data in each of the buffers accessible by the destination device to the corresponding write buffers inaccessible by the source device via a read operation performed by the destination device.

19. A non-transitory processor-readable medium comprising processor-readable instructions that, when executed by at least one processor of a controller of a source device, cause the processor to perform the following operations: Communicating with the destination device to set the filling of the buffer of the source device; and Perform a transfer to the destination device to transfer data from the buffer of the source device to the buffer of the destination device. The buffer of the source device includes a read buffer that is inaccessible to the destination device, and the buffer of the destination device includes a buffer that is accessible to the source device. The transfer to the source device to transfer the data from the buffer of the source device to the buffer of the destination device includes performing a write operation by the source device to write blocks of the data in each of the read buffers that are inaccessible to the destination device to the corresponding buffers that are accessible to the source device.

Citation Information

Patent Citations

  • Transmit buffers in connection-oriented interface

    US20070005833A1

  • Supplemental Communication Interface

    US20080147923A1

  • Storage device initiated copy back operation

    US20190065382A1

  • Multi-mode processor bus bridge

    US6757762B1