Memory device, host system, and method of operating a memory device

By introducing a DMA engine into the memory device, internal data copying between memory regions is achieved, solving the problems of data transfer latency and bandwidth efficiency in computing systems and improving system computing performance.

CN114442916BActive Publication Date: 2026-05-29SAMSUNG ELECTRONICS CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SAMSUNG ELECTRONICS CO LTD
Filing Date
2021-07-22
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

In computing systems, increased latency and reduced interface bandwidth efficiency during data transmission negatively impact the efficiency of computational operations.

Method used

Memory devices employing direct memory access (DMA) engines communicate with multiple host devices via interconnects to enable data copying operations between memory regions, avoiding data output through external interconnects and utilizing the DMA engine to complete data copying and transfer within the memory device.

Benefits of technology

It reduces data processing waiting time, improves interface bandwidth efficiency, and optimizes computational performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114442916B_ABST
    Figure CN114442916B_ABST
Patent Text Reader

Abstract

A memory device, a host system, and a method of operating a memory device are provided. The memory device is configured to communicate with a plurality of host devices over an interconnect and includes a memory having a plurality of memory regions, including a first memory region allocated to a first host device and a second memory region allocated to a second host device. The memory device further includes a direct memory access (DMA) engine configured to, based on a request from the first host device including a copy command to copy data stored in the first memory region to the second memory region, read the stored data from the first memory region and write the read data to the second memory region without outputting the read data to the interconnect.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims priority to Korean Patent Application No. 10-2020-0148133, filed with the Korean Intellectual Property Office on November 6, 2020, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This disclosure relates to memory devices, and more specifically, to a memory device including a direct memory access (DMA) engine, a system including the memory device, and a method of operating the memory device. Background Technology

[0004] A system configured to process data (e.g., a computing system) may include a central processing unit (CPU), memory devices, input / output (I / O) devices, and a root complex configured to transfer information between devices included in the system. As an example, devices included in the system may send and receive requests and responses based on various types of protocols such as Fast Peripheral Component Interconnect (PCIe).

[0005] The system may include a memory device that can be shared between at least two devices different from the memory device. During data processing, various types of computational operations can be performed, and data accessed by the memory device can be frequently migrated. In this case, latency during data transfer may increase, or interface bandwidth efficiency may decrease. As a result, the time spent on computational operations may increase. Summary of the Invention

[0006] A memory device is provided that can reduce latency and improve interface bandwidth efficiency during data transmission for processing computational operations of a computing system, a system including the memory device, and a method for operating the memory device.

[0007] Other aspects will be set forth in part in the description which follows, and other aspects will be partially obvious from the description, or may be partially learned by practice of the embodiments presented.

[0008] According to an embodiment, a memory device is provided, configured to communicate with a plurality of host devices via an interconnect. The memory device includes a memory comprising a plurality of storage regions, including a first storage region allocated to a first host device and a second storage region allocated to a second host device. The memory device further includes a direct memory access (DMA) engine configured to perform an operation based on a request from the first host device, wherein the request includes a copy command to copy data stored in the first storage region to the second storage region, the operation including: reading the stored data from the first storage region; and writing the read data to the second storage region without outputting the read data to the interconnect.

[0009] According to an embodiment, a method for operating a memory device is provided, the memory device being configured to communicate with a plurality of host devices via an interconnect, the memory device including a plurality of storage regions, the plurality of storage regions including a first storage region allocated to a first host device and a second storage region allocated to a second host device, the method including: receiving a request from the first host device, the request including a copy command to copy data stored in the first storage region to the second storage region. The method further includes, based on receiving the request, reading the stored data from the first storage region; and writing the read data to the second storage region without outputting the read data to the interconnect.

[0010] According to an embodiment, a host system is provided, the host system comprising: a root association including a first root port and a second root port, the root association being configured to provide interconnects based on a predetermined protocol; a first host device configured to communicate with a memory device via the first root port, the first memory region of the memory device corresponding to a first logical device being allocated to the first host device. The host system further comprises a second host device configured to communicate with the memory device via the second root port, the second memory region of the memory device corresponding to a second logical device being allocated to the second host device. The host system is configured to: receive a response from the memory device indicating that the copying of the data stored in the first memory region to the second memory region has been completed, based on a request sent by the first host device to the memory device to copy data stored in the first memory region to the second memory region, without receiving the data from the memory device through the root association. Attached Figure Description

[0011] The above and other aspects, features and advantages of the embodiments of this disclosure will become more apparent from the following description taken in conjunction with the accompanying drawings, wherein:

[0012] Figure 1 This is a block diagram of a system including a memory device according to an embodiment;

[0013] Figure 2A and Figure 2B This is a block diagram of the host devices included in the system according to the embodiment;

[0014] Figure 3 This is a block diagram of the system according to an embodiment;

[0015] Figure 4 yes Figure 3 A block diagram of the first memory device;

[0016] Figure 5A and Figure 5B This is a diagram illustrating an example of a data transmission path in a system according to an embodiment;

[0017] Figure 6A and Figure 6B These are diagrams illustrating implementation examples and operation examples of requests not applied in the embodiments;

[0018] Figure 7A and Figure 7B These are diagrams illustrating implementation examples and operation examples of the request according to the embodiments;

[0019] Figure 8 This is a flowchart of a method for operating a host system according to an embodiment;

[0020] Figure 9 This is a block diagram of a system for applying a memory device according to an embodiment to a CXL-based Type 3 device;

[0021] Figure 10 This is a diagram illustrating the application of a memory device according to an embodiment to a CXL-based Type 2 device.

[0022] Figure 11A and Figure 11B This is a block diagram of the system according to an embodiment; and

[0023] Figure 12 This is a block diagram of a data center including a system, according to an embodiment. Detailed Implementation

[0024] The embodiments will now be described more fully with reference to the accompanying drawings.

[0025] Figure 1 This is a block diagram of a system 10 including a memory device according to an embodiment.

[0026] Reference Figure 1 System 10 may be referred to differently, such as a data processing system or a computing system, and includes at least one device. As an example, system 10 may include at least one memory device and a device configured to request data access to the at least one memory device. Because the device configured to request data access operates against the at least one memory device, this device may be referred to as a host device.

[0027] According to an embodiment, system 10 may include a first host device 11, a second host device 12, and a first memory device 13 to a third memory device 15. Although Figure 1 The system 10 shown includes two host devices and three memory devices as an example; however, the embodiments are not limited thereto, and the system 10 may include a variety of numbers of devices.

[0028] The first memory device 13 through the third memory device 15 may each include various types of memory. As an example, the first memory device 13 through the third memory device 15 may each include non-volatile memory such as solid-state drives (SSDs), flash memory, magnetic random access memory (MRAM), ferroelectric RAM (FRAM), phase-change RAM (PRAM), and resistive RAM (RRAM). However, the embodiments are not limited thereto, and the first memory device 13 through the third memory device 15 may each include dynamic RAM (DRAM) such as double data rate synchronous DRAM (DDR SDRAM), low-power DDR (LPDDR) SDRAM, graphics DDR (GDDR) SDRAM, and Rambus DRAM (RDRAM).

[0029] Devices included in System 10 can communicate with each other via interconnects (or links) configured to support at least one protocol. Each device may include an internal component configured to perform protocol-based communication supported by the interconnect. For example, at least one protocol selected from the following technologies may be applied to the interconnect: Fast Peripheral Component Interconnect (PCIe) protocol, Compute Express Link (CXL) protocol, XBus protocol, NVLink protocol, Infinite Structure protocol, Cache Coherent Interconnect for Accelerators (CCIX) protocol, and Coherent Accelerator Processor Interface (CAPI) protocol. In the following embodiments, communication based on the CXL protocol will be described primarily. However, the embodiments are not limited thereto, and various protocols other than the CXL protocol may be applied.

[0030] Although interconnections between the first host device 11, the second host device 12, and the first memory devices 13 through 15 are shown briefly for simplicity, system 10 may include a root union connected to multiple devices via a root port, through which the first host device 11, the second host device 12, and the first memory devices 13 through 15 can communicate with each other. For example, the root union may manage transactions between the first host device 11 and the second host device 12 and the first memory devices 13 through 15. In embodiments, mutual communication may be performed based on various other components and functions according to the CXL standard. As an example, communication using various protocols may be enabled based on components disclosed in the CXL standard (e.g., flexible buses and switches). At least some of the first memory devices 13 through 15 may be connected to the first host device 11 and / or the second host device 12 via a bridge (e.g., a PCI bridge) configured to control communication paths based on a predetermined protocol.

[0031] According to embodiments, both the first host device 11 and the second host device 12 may include various types of devices. For example, both the first host device 11 and the second host device 12 may include any one or any combination of programmable components, components configured to provide fixed functions (e.g., intellectual property (IP) cores), reconfigurable components, and peripheral devices. The programmable component may be a central processing unit (CPU), graphics processing unit (GPU), or neural processing unit (NPU) configured as the main processor for all operations of the control system 10. The reconfigurable component may be a field-programmable gate array (FPGA). The peripheral device may be a network interface card (NIC).

[0032] According to an embodiment, any one or any combination of the first memory device 13 to the third memory device 15 may be shared between the first host device 11 and the second host device 12. For example, the first memory device 13 may be pooled memory shared between the first host device 11 and the second host device 12. The first memory device 13 may store instructions executed by the first host device 11 and the second host device 12, or store input data for computational operations and / or the results of computational operations. The first memory device 13 may include a direct memory access (DMA) engine 13_1 and memory 13_2.

[0033] According to an embodiment, memory 13_2 may include multiple storage regions allocated to multiple host devices that are different from each other. For example, storage regions may correspond to logically separate logical devices, and a first memory device 13 as a physical device may be recognized by system 10 as multiple devices (e.g., multiple memory devices). Memory 13_2 may include first storage region LD0 to nth storage region LD(n-1). Figure 1 In the illustrated embodiment, it is assumed that a first storage region LD0 is allocated to a first host device 11, and an nth storage region LD(n-1) is allocated to a second host device 12. Both the first host device 11 and the second host device 12 may include a request generator. The first storage region LD0 and the nth storage region LD(n-1) can be accessed independently by different host devices. As described above, system 10 may include various CXL-based devices. When the first host device 11 corresponds to a CPU and the first memory device 13 is a CXL-based Type 3 device, the first host device 11 can execute layered software including applications, generate a data packet including data, and send the data packet to the first memory device 13. Figure 1The request generator of the first host device 11 may include hardware and / or software components related to the generation of data packets. Additionally, the first memory device 13 may include components configured to process data packets containing data. As described below, the first memory device 13 may include a memory controller configured to process data packets. The memory controller may be implemented as a separate device from memory 13_2, or it may be included in the same device as memory 13_2. Furthermore, the first storage regions LD0 to nth storage regions LD(n-1) may be allocated differently to the host device, and as an example, at least two of the first storage regions LD0 to nth storage regions LD(n-1) may be allocated to one host device.

[0034] Furthermore, DMA engine 13_1 can control the transfer paths of data stored in memory 13_2 and data read from memory 13_2. As an example, DMA engine 13_1 can send data read from any storage area of ​​memory 13_2 to another storage area. In various data transfer operations performed by system 10, data stored in areas allocated to a host device can be read and stored (or copied) in areas allocated to another host device. For example, in memory 13_2, data stored in the first storage area LD0 can be read, and the read data can be stored in the nth storage area LD(n-1).

[0035] According to an embodiment, when data is transferred between multiple logical devices in a memory device corresponding to a physical device, based on the control of DMA engine 13_1, data read from the first memory device 13 can be transferred through the path of the first memory device 13 without being output to the interconnect. As an example, a request including a copy command can be provided to the memory device, and the copy command can be defined differently from normal read and write commands. As an example, the first host device 11 can provide the first memory device 13 with a first access request Req_1 including a copy command. The first access request Req_1 may include an address indicating the location of a first storage region LD0 (e.g., a source address) and an address indicating the location of the nth storage region LD(n-1) (e.g., a destination address).

[0036] According to an embodiment, in response to a first access request Req_1, DMA engine 13_1 can perform path control operations to receive data Data read from a first storage region LD0 and send the data Data to the nth storage region LD(n-1). For example, the first memory device 13 may include a processor configured to process the first access request Req_1. The first memory device 13 may send the data Data read from the first storage region LD0 to the nth storage region LD(n-1) based on the control of the DMA engine 13_1 included in the first memory device 13, without outputting the data Data read from the first storage region LD0 to an external interconnect based on the information included in the first access request Req_1.

[0037] In contrast, when a request from the host device corresponds to a normal read request or a request to migrate data to a storage device physically different from the first memory device 13, the first memory device 13 can output the data Data read from memory 13_2 via an external interconnect. In the operational example, in response to a first access request Req_1 from the first host device 11, based on the control of the DMA engine 13_1, the data Data read from the first storage region LD0 can be transferred via a path through the external interconnect of the first memory device 13.

[0038] According to the example embodiments described above, data transfer efficiency between logical devices identified as different devices within the same memory device can be improved. For example, a copy function can be performed using the DMA engine in the memory device to send read data to the memory area without processing data read from the storage area, based on a predetermined protocol. Therefore, the latency of data processing operations can be reduced, and the data traffic of interconnects can be decreased.

[0039] Figure 2A and Figure 2B This is a block diagram of the host device included in system 100 according to an embodiment. Figure 2A and Figure 2B An example of a link based on the CXL protocol as an interconnect between devices is shown.

[0040] Reference Figure 2A System 100 may include various types of host devices. Although in Figure 2AThe embodiments are illustrated using host processor 110 and accelerator 120 (e.g., GPU and FPGA) as examples; however, the embodiments are not limited thereto. Various other types of devices configured to send access requests can be applied to system 100. Host processor 110 and accelerator 120 can send and receive messages and / or data to each other via link 150 configured to support the CXL protocol. In an embodiment, host processor 110, as the main processor, can be a CPU configured to control all operations of system 100.

[0041] Additionally, system 100 may include host memory 130 connected to host processor 110 and device memory 140 mounted at accelerator 120. Host memory 130 connected to host processor 110 may support cache coherence. Accelerator 120 may manage device memory 140 independently of host memory 130. Host memory 130 and device memory 140 may be accessed by multiple host devices. As an example, accelerator 120 and devices such as NICs may access host memory 130 via PCIe DMA.

[0042] In some embodiments, link 150 may support multiple protocols (e.g., sub-protocols) defined in the CXL protocol, and messages and / or data may be transmitted through multiple protocols. For example, protocols may include a non-consistent protocol (or I / O protocol CXL.io), a consistent protocol (or caching protocol CXL.cache), and a memory access protocol (or memory protocol CXL.memory).

[0043] The I / O protocol CXL.io can be a PCIe-like I / O protocol. Shared memory (e.g., pooled memory) included in system 100 can communicate with host devices based on PCIe or the CXL.io I / O protocol. The main processor 110 and accelerator 120 can access [the system] via PCIe or CXL.io-based interconnects. Figure 1 The embodiment shown is a memory device 13. Additionally, the caching protocol CXL.cache can provide a protocol through which the accelerator 120 can access the host memory 130, and the memory protocol CXL.memory can provide a protocol through which the host processor 110 can access the device memory 140.

[0044] Accelerator 120 can refer to any device configured to provide functionality to host processor 110. For example, at least some of the computational and I / O operations performed on host processor 110 can be offloaded to accelerator 120. In some embodiments, accelerator 120 may include any one or any combination of programmable components (e.g., GPUs and NPUs), components configured to provide fixed functionality (e.g., IP cores), and reconfigurable components (e.g., FPGAs).

[0045] Accelerator 120 may include physical layer 121, multiprotocol multiplexer (MUX) 122, interface circuitry 123, and accelerator logic 124, and communicates with device memory 140. Accelerator logic 124 may communicate with host processor 110 using the aforementioned protocols via multiprotocol MUX 122 and physical layer 121.

[0046] Interface circuitry 123 can determine one of several protocols based on messages and / or data for communication between accelerator logic 124 and host processor 110. Interface circuitry 123 can connect to at least one protocol queue included in multi-protocol MUX 122 and send messages and / or data to and receive messages and / or data from host processor 110 via at least one protocol queue.

[0047] The multi-protocol MUX 122 may include at least one protocol queue and send and receive messages and / or data to and from the host processor 110 through the at least one protocol queue. In some embodiments, the multi-protocol MUX 122 may include multiple protocol queues corresponding to multiple protocols supported by the link 150, respectively. In some embodiments, the multi-protocol MUX 122 may arbitrate communication between different protocols and communicate based on the selected protocol.

[0048] Device memory 140 can be connected to accelerator 120 and is referred to as device-attached memory. Accelerator logic 124 can communicate with device memory 140 based on a protocol independent of link 150 (i.e., a device-specific protocol). In some embodiments, accelerator 120 may include a controller as a component for accessing device memory 140, and accelerator logic 124 can access device memory 140 through the controller. The controller can access device memory 140 of accelerator 120 and also enables host processor 110 to access device memory 140 via link 150. In some embodiments, device memory 140 may correspond to CXL-based device-attached memory.

[0049] The host processor 110 may be the main processor (e.g., CPU) of the system 100. In some embodiments, the host processor 110 may be a CXL-based host processor or host. Figure 2A As shown, the host processor 110 can be connected to the host memory 130 and includes a physical layer 111, a multi-protocol MUX 112, an interface circuit 113, a coherence / cache circuit 114, a bus circuit 115, at least one core 116, and I / O devices 117.

[0050] At least one core 116 can execute instructions and is connected to a coherence / cache circuit 114. The coherence / cache circuit 114 may include a cache hierarchy and is referred to as coherence / cache logic. For example... Figure 2A As shown, the coherence / caching circuit 114 can communicate with at least one core 116 and interface circuit 113. For example, the coherence / caching circuit 114 can communicate via at least two protocols, including a coherence protocol and a memory access protocol. In some embodiments, the coherence / caching circuit 114 may include DMA circuitry. I / O device 117 can be used to communicate with bus circuit 115. For example, bus circuit 115 may be PCIe logic, and I / O device 117 may be a PCIe I / O device.

[0051] Interface circuitry 113 enables communication between components of host processor 110 (e.g., coherence / caching circuitry 114 and bus circuitry 115) and accelerator 120. In some embodiments, interface circuitry 113 can enable communication between components of host processor 110 and accelerator 120 according to multiple protocols (e.g., non-coherence protocols, coherence protocols, and memory protocols). For example, interface circuitry 113 can determine one of multiple protocols based on messages and / or data for communication between components of host processor 110 and accelerator 120.

[0052] The multi-protocol MUX 112 may include at least one protocol queue. Interface circuitry 113 may be connected to at least one protocol queue and send messages and / or data to and receive messages and / or data from accelerator 120 via the at least one protocol queue. In some embodiments, interface circuitry 113 and multi-protocol MUX 112 may be integrally formed as a single component. In some embodiments, multi-protocol MUX 112 may include multiple protocol queues corresponding to multiple protocols supported by link 150. In some embodiments, multi-protocol MUX 112 may arbitrate communication between different protocols and provide selected communication to physical layer 111.

[0053] Furthermore, according to the embodiments, Figure 1The request generator of the host device shown can correspond to Figure 2A The various components shown may be included within components. For example, the functionality of a request generator may be provided by... Figure 2A The various components shown (e.g., accelerator logic 124, at least one core 116, and I / O device 117) are used to perform this.

[0054] The host processor 110, accelerator 120, and various peripheral devices can access Figure 1 The memory device 13 shown is an example. Figure 1 The illustrated memory device may include a first memory region allocated to the host processor 110 and a second memory region allocated to the accelerator 120, and data may be directly transferred between the first and second memory regions based on the control of a DMA engine included therein. As an example, data may be read from the first memory region in response to a request from the host processor 110, and the read data may be copied to the second memory region without being output to the outside of the memory device based on the control of the DMA engine.

[0055] Figure 2B It shows the use of Figure 2A An example of multi-protocol communication in system 100. Furthermore, Figure 2B An example is shown where both the host processor 110 and the accelerator 120 include a memory controller. Figure 2B The memory controller shown may include being included in Figure 2A Some components in each of the devices shown, or implementations of these components separately.

[0056] The host processor 110 can communicate with the accelerator 120 based on multiple protocols. According to the CXL example above, these protocols may include the memory protocol CXL.memory (or MEM), the consistency protocol CXL.cache (or COH), and the non-consistency protocol CXL.io (or IO). The memory protocol MEM can define transactions between the master and slave devices. For example, the memory protocol MEM can define transactions from the master to the slave and from the slave to the master. The consistency protocol COH can define the interaction between the accelerator 120 and the host processor 110. For example, the interface of the consistency protocol COH may include three channels for transmitting requests, responses, and data. The non-consistency protocol IO can provide a non-consistent load / store interface for I / O devices.

[0057] Accelerator 120 may include a memory controller 125 configured to communicate with and access device memory 140. In some embodiments, memory controller 125 may be external to accelerator 120 and integrated with device memory 140. Additionally, host processor 110 may include a memory controller 118 configured to communicate with and access host memory 130. In some embodiments, memory controller 118 may be external to host processor 110 and integrated with host memory 130.

[0058] Figure 3 This is a block diagram of system 200 according to an embodiment. Figure 3 The multiple host devices shown (e.g., 201 to 203) may include various types of devices (e.g., CPU, GPU, NPU, FPGA and peripheral devices) according to the above embodiments.

[0059] System 200 may include a root union 210 and host devices 201 to 203. Root union 210 may include a DMA engine 211 and at least one root port (e.g., a first root port (RP) 213 and a second root port (RP) 214) connected to a memory device. According to embodiments, root union 210 may also include a structure manager 212 configured to transmit data or requests via a fabric such as Ethernet. Root union 210 may be connected to an endpoint 223 via a fabric. As an example, endpoint 223 may include flash memory (e.g., SSDs and Universal Flash Memory (UFS)), volatile memory (e.g., SDRAM), and non-volatile memory (e.g., PRAM, MRAM, RRAM, and FRAM). Figure 3 The root assembly 210 is shown as a component separate from the host device, but the root assembly 210 can be integrated into each of the host devices 201 to 203.

[0060] Root union 210 can provide data communication between host devices 201 to 203 and first memory devices 230 to third memory devices 250 based on various types of protocols, and supports the CXL protocol described in the above embodiments. In the embodiments, root union 210 and first memory devices 230 to third memory devices 250 can perform interface functions including various protocols defined in CXL (e.g., the I / O protocol CXL.io).

[0061] Furthermore, the first memory device 230 to the third memory device 250 can all correspond to Type 3 devices as defined in the CXL protocol. Therefore, the first memory device 230 to the third memory device 250 can all include a memory expander. According to an embodiment, the memory expander can include a controller. Although in Figure 3 The image shows a device that includes memory, but memory and memory expanders that include multiple memory regions can be implemented as separate devices.

[0062] According to an embodiment, system 200 can support multiple virtual channels. A virtual channel can provide multiple transmission paths that are logically separated from each other within a single physical interface. Although Figure 3 The illustrated embodiment shows a first virtual channel (VCS) 221 corresponding to the first root port 213 and a second virtual channel 222 corresponding to the second root port 214. The virtual channels included in the system 200 can also be implemented in various other forms. Furthermore, the host devices 201 to 203, the root union 210, and the virtual channels can be described as components constituting the host system.

[0063] Through virtual channels, a single root port can connect to multiple devices that are different from each other, or at least two root ports can connect to a single device. The PCI / PCIe example will now be used. A first root port 213 can be connected to a second memory device 240 via a path including at least one PCI-PCI bridge (PPB) and at least one virtual PCI-PCI bridge (vPPB), and to a first memory device 230 via another path including at least one PPB and at least one vPPB. Similarly, a second root port 214 can be connected to the first memory device 230 via a path including at least one PPB and at least one vPPB, and to a third memory device 250 via another path including at least one PPB and at least one vPPB.

[0064] System 200 can provide multiple logical devices (MLDs) supported by the CXL protocol. In an embodiment, in Figure 3 In the structure of system 200 shown, the first memory device 230 can communicate with at least two host devices via a first root port 213 and a second root port 214. The memory 232 of the first memory device 230 may include first memory regions LD0 to nth memory regions LD(n-1) allocated to different host devices. For example, the first memory region LD0 can be allocated to the first host device 201 and communicate with it via the first root port 213. Similarly, the second memory region LD1 can be allocated to the second host device 202 and communicate with it via the second root port 214. Furthermore, both the second memory device 240 and the third memory device 250 can communicate via either root port and correspond to a single logic device (SLD).

[0065] The first memory device 230 may include a DMA engine 231 and is logically identified as multiple devices. In an operational example, a data processing operation can be performed to copy data stored in any one logical device (e.g., a storage area) to another logical device. According to the above embodiment, based on the control of the DMA engine 231, data read from any storage area can be copied to another storage area without being output to the outside. That is, data transfer between multiple logical devices in the first memory device 230 can be performed by the DMA engine 231 in the first memory device 230 without migrating data through virtual channels and root ports. Therefore, latency can be reduced and interface bandwidth efficiency can be improved. According to the above embodiment, the structure of the system 200 can remain unchanged or the changes can be minimized, and unnecessary data movement can be minimized. For example, when a memory device including a DMA engine is added to the system 200, data copying as described above can be controlled by adding resources for driving the DMA engine in the memory device to the system 200.

[0066] Figure 4 yes Figure 3 Block diagram of the first memory device 230.

[0067] Reference Figure 3 and Figure 4 The first memory device 230 may include a DMA engine 231, a memory 232, and a command (CMD) executor 233. The memory 232 may include multiple memory regions, namely, a first memory region LD0 to an nth memory region LD(n-1). Although in Figure 4 The first memory device 230 is exemplarily shown as a single device, but the memory 232 and the memory expander including the DMA engine 231 and the command executor 233 can be implemented as separate devices as described above.

[0068] The first memory device 230 may include interface circuitry and communicate with an external device based on a predetermined protocol. The command executor 233 may receive a request including a command CMD from an external host device, execute the command CMD, and determine whether the request instructs the copying of data Data between memory regions of memory 232. The DMA engine 231 may control the transfer path of data Data based on the control of the command executor 233. As an example, the DMA engine 231 may receive data Data read from a first memory region LD0 allocated to either host device and send the data Data to a second memory region LD1 allocated to another host device. The DMA engine 231 may provide information indicating that the copying of data Data is complete to the command executor 233, and the command executor 233 may output a response indicating that the copying of data Data is complete to an external device.

[0069] Figure 5A and Figure 5B This is a diagram illustrating an example of a data transmission path in a system according to an embodiment. As described in the above embodiments... Figure 5A and Figure 5B The components shown are omitted here, so a detailed description of them can be omitted. Figure 5A The data transmission path in the example not being applied is shown. Figure 5B The data transmission path according to an embodiment is shown.

[0070] Reference Figure 5A The system can provide a memory device including an MLD (e.g., a first memory device (Type 3 (Pooled Memory)) 230), and can provide requests from a host device to the first memory device 230. Data transfer paths can be controlled in response to requests from the host device. For example, assuming data from a first memory region LD0 is copied to a second memory region LD1 corresponding to another logical device, the data read from the first memory region LD0 can be provided to a DMA engine 211 included in the root union 210 via a virtual channel and a root port. The DMA engine 211 can then transfer data to the second memory region LD1 of the first memory device 230 via another root port and another virtual channel (path (1)). Additionally, when data from the first memory region LD0 is provided to a memory device different from the first memory device 230 (e.g., a third memory device 250), the data can be transferred to the third memory device 250 via the DMA engine 211 of the root union 210 (path (2)).

[0071] In contrast, when Figure 5B When the embodiment shown is applied, the command executor (reference) Figure 4 233) can determine whether data should be transferred between memory regions and can control the data transfer path based on the determination result. For example, based on the control of DMA Engine 231, data read from the first memory region LD0 can be sent to the second memory region LD1 instead of being output to the outside (path (1)). In contrast, when the request from the host device is a normal read request or corresponds to the transfer of data to a memory device that is physically different from the first memory device (Type 3 (Pooled Memory)) 230, based on the determination result of command executor 233, data read from the first memory region LD0 can be output to the outside and then transferred to the third memory device 250 through DMA Engine 211 of root association 210 (path (2)).

[0072] Figure 6A and Figure 6B These are diagrams illustrating implementation examples and operation examples of requests not applied in the embodiments. Figure 7A and Figure 7B These are diagrams illustrating implementation examples and operation examples of the request according to the embodiments. Figure 6A and Figure 6B as well as Figure 7A and Figure 7B In the diagram, memory device A includes multiple logical devices, and the operation of copying data from a source address (or read address ReadAddr) to a destination address (or write address WriteAddr) is illustrated.

[0073] Reference Figure 6A and Figure 6B The host device can sequentially provide at least two requests to copy data. Each request may include a command code OPCODE indicating the type of memory operation and an address and / or data indicating the access location.

[0074] For example, to read data stored in the first memory region LD0 of memory MEM, the host device can provide the memory device A with the following request: this request includes a command code corresponding to the read command RD and a read address (Read Addr) indicating the location of the first memory region LD0, and provides the data to the host device based on the control of the DMA engine. Alternatively, the host device can provide the memory device A with the following request: this request includes a command code corresponding to the write command WR, a write address (Write Addr) indicating the location of the second memory region LD1, and data, and writes the data to the second memory region LD1 corresponding to the location indicated by the write address (Write Addr) based on the control of the DMA engine. Thus, a copy operation can be completed.

[0075] In addition, refer to Figure 7A and Figure 7B According to an embodiment, a copy command CP can be defined to instruct data copying between multiple logical devices in the same memory device. A request including the copy command CP, a read address (ReadAddr) as the source address, and a write address (WriteAddr) as the destination address can be provided to memory device A. Memory device A can read data from a first memory region LD0 corresponding to the read address ReadAddr in memory MEM without outputting the data externally, based on the control of the DMA engine, and store the read data in a second memory region LD1 corresponding to the write address WriteAddr. Therefore, as... Figure 6BData migration in the host system shown does not occur during data copying. Additionally, storage device Device A can send a request to the host device (Host) including a command code and a message MSG indicating that data copying is complete (Comp).

[0076] Figure 8 This is a flowchart of a method for operating a host system according to an embodiment.

[0077] Reference Figure 8 The host system may include at least one host device and communicate with at least one memory device. According to a specific example, the host system may be defined as including, in addition to at least one host device, a root union and a virtual channel according to the above embodiments. According to the above embodiments, at least two host devices may share a memory device comprising multiple logical devices, and a first storage area (or a first logical device) of the first memory device may be allocated to the first host device.

[0078] In operation S11, the first host device can perform a computational operation based on access to a memory device. As an example of a computational operation, data can be copied from a first storage area of ​​a first memory device allocated to the first host device to a memory device allocated to a second host device. In operation S12, the first host device can determine the memory devices allocated to other host devices in the host system and determine the location of the memory device allocated to the second host device. In operation S13, the first host device can determine whether the location of the memory device allocated to the second host device to which data is to be copied corresponds to another logical device of the same first memory device.

[0079] When it is determined that the location to which data is to be copied corresponds to a device physically different from the first memory device, data read and write operations can be performed sequentially according to the above embodiments. Therefore, in operation S14, a normal data access request can be sent to the memory device, thereby enabling the data copy operation to be performed. Otherwise, when the location to which data is to be copied corresponds to another logical device in the first memory device (e.g., a second storage area corresponding to a second logical device), in operation S15, the first host device can send a request to the first memory device including a copy command between logical devices in the first memory device. According to an embodiment, this request may include an address indicating the location from which data is read and an address indicating the location to which data is copied. In response to this request, the first memory device can internally perform the data copy operation based on the control of the DMA engine and output a response indicating that the data copy operation is complete, and in operation S16, the first host device can receive the response indicating that the data copy operation is complete.

[0080] Figure 9 This is a block diagram of a system 300 in which a memory device according to an embodiment is applied to a CXL-based Type 3 device.

[0081] Reference Figure 9 System 300 may include a root union 310 and CXL memory extenders 320 and memories 330 connected to the root union 310. The root union 310 may include a home agent and an I / O bridge. The home agent may communicate with the CXL memory extender 320 based on the memory protocol CXL.memory, and the I / O bridge may communicate with the CXL memory extender 320 based on the I / O protocol CXL.io. Based on the CXL protocol, the home agent may correspond to an agent located on the host side, which is positioned to fully resolve the consistency of system 300 for a given address.

[0082] The CXL memory expander 320 may include a memory controller (MC). According to the example embodiment described above, Figure 9 An example is shown below: the memory controller includes a DMA engine 321 configured to control data transfer paths between multiple logical devices. However, this embodiment is not limited to this, and the DMA engine 321 may be external to the memory controller. In this embodiment, the CXL memory expander 320 may output data to the root union 310 via an I / O bridge based on the I / O protocol CXL.io or a similar PCIe protocol.

[0083] Furthermore, according to the above embodiments, the memory 330 may include multiple memory regions (e.g., first memory region LD0 to nth memory region LD(n-1)), and each memory region may be implemented as various cells of the memory. As an example, when the memory 330 includes multiple volatile or non-volatile memory chips, the cell of each memory region may be a memory chip. Alternatively, the memory 330 may be implemented such that the cell of each memory region corresponds to one of various sizes defined in the memory (e.g., semiconductor die, block, bank, and rank).

[0084] Figure 10 This diagram illustrates the application of a memory device according to an embodiment to a CXL-based Type 2 device. The CXL-based Type 2 device can include various types of devices. According to the above embodiment, the Type 2 device can include an accelerator, which includes accelerator logic such as GPUs and FPGAs.

[0085] Reference Figure 10Device 400 (i.e., memory device 400) may include a DMA engine 410, memory 420, a command executor 430, a CXL-based Data Coherence Engine (DCOH) 440, and accelerator logic 450. As an example, because device 400 corresponds to a CXL-based Type 2 device, device 400 may include memory 420. In a specific example, as described in the above embodiments, device 400 may include device-attached memory or be connectable to device-attached memory, and memory 420 may correspond to device-attached memory. Although... Figure 10 The illustration shows memory 420 within device 400, but memory 420 may also be located externally to device 400. Device 400 may access host memory based on the caching protocol CXL.cache, or enable the host processor to access memory 420 based on the memory protocol CXL.memory. Additionally, device 400 may include a memory controller. The memory controller may include a DMA engine 410 and a command executor 430.

[0086] According to the above embodiments, the memory 420 may include multiple storage regions identified as different logical devices by the host device, and the command executor 430 can control data access operations on the memory 420 by executing commands included in requests from the host device. Furthermore, the DMA engine 410 can control the transfer path of data read from or written to the memory 420. According to embodiments, the data transfer path can be controlled such that data read from any storage region is copied to another storage region without being output through interconnects located outside the memory device 400.

[0087] Furthermore, DCOH 440 can correspond to an agent configured to resolve consistency related to the device cache on the device. DCOH 440 can perform consistency-related functions (such as updating metadata fields) based on the processing results of command executor 430 and DMA engine 410, and provide the update results to command executor 430. Accelerator logic 450 can request access to memory 420 through the memory controller, and also requests access to memory in the system located outside of device 400.

[0088] Figure 11A and Figure 11B This is a block diagram of the system according to an embodiment. In detail, Figure 11A and Figure 11B These are block diagrams of System 1A and System 1B, each containing multiple CPUs.

[0089] Reference Figure 11ASystem 1A may include a first CPU 21, a second CPU 31, a first double data rate (DDR) memory 22 and a second DDR memory 32 connected to the first CPU 21 and the second CPU 31, respectively. An interconnect system based on processor interconnect technology can be used to connect the first CPU 21 and the second CPU 31. Figure 11A The interconnect system shown can provide at least one CPU-to-CPU coherent link.

[0090] System 1A may include a first I / O device 23 and a first accelerator 24 communicating with a first CPU 21, and a first device memory 25 connected to the first accelerator 24. The first CPU 21 may communicate with each of the first I / O device 23 and the first accelerator 24 via a bus. Additionally, System 1A may include a second I / O device 33 and a second accelerator 34 communicating with a second CPU 31, and a second device memory 35 connected to the second accelerator 34. The second CPU 31 may communicate with each of the second I / O device 33 and the second accelerator 34 via a bus. In some embodiments, either or any combination of the first device memory 25 and the second device memory 35 in System 1A may be omitted.

[0091] Furthermore, system 1A may include remote far memory 40. The first CPU 21 and the second CPU 31 may be connected to remote far memory 40 via a bus, respectively. Remote far memory 40 can be used to expand the memory in system 1A. In some embodiments, remote far memory 40 in system 1A may be omitted.

[0092] System 1A can perform communication via a bus based on at least some of a plurality of protocols. As an example, the CXL protocol will be described. Information such as initialization information can be sent based on the I / O protocol CXL.io, or messages and / or data can be sent based on the cache protocol CXL.cache and / or the memory protocol CXL.memory.

[0093] exist Figure 11A In the system 1A shown, at least two host devices can share the remote remote memory 40. Although Figure 11A The illustration shows a scenario where the remote remote memory 40 is shared between the first CPU 21 and the second CPU 31, but the remote remote memory 40 can also be shared between various other host devices. When the memory device according to the embodiment is applied to the remote remote memory 40, the remote remote memory 40 can include multiple logical devices identified as different devices, and data can be transferred between the multiple logical devices based on the control of the DMA engine 41 without outputting data to the outside.

[0094] In addition, Figure 11B System 1B does not include remote remote memory 40. Figure 11B System 1B can be with Figure 11A The system 1A is essentially the same. The operation of the memory device according to the embodiment can be applied to… Figure 11B The various memory devices shown. As an example, any one or any combination of the first DDR memory 22, the second DDR memory 32, the first device memory 25, and the second device memory 35 may include a DMA engine according to the above embodiments. Therefore, data copying can be performed between the plurality of logical devices included in each of the first DDR memory 22, the second DDR memory 32, the first device memory 25, and the second device memory 35.

[0095] Figure 12 This is a block diagram of a data center 2 including a system according to an embodiment. In some embodiments, the system described above with reference to the accompanying drawings can be used as an application server and / or a storage server and can be included in data center 2. Furthermore, according to an embodiment, data replication between the memory device and the DMA engine-based controlled logic device according to the embodiment can be applied to each of the application server and / or storage server.

[0096] Reference Figure 12 Data Center 2 can collect various types of data and provide services, and is also known as a data storage center. For example, Data Center 2 could be a system configured to run search engines and databases or computing systems used by companies such as banks or government agencies. Figure 12 As shown, data center 2 may include application servers 50_1 to 50_n and storage servers 60_1 to 60_m (where m and n are both integers greater than 1). According to an embodiment, the number of application servers 50_1 to 50_n (n) and the number of storage servers 60_1 to 60_m (m) may be selected differently. The number of application servers 50_1 to 50_n (n) may differ from the number of storage servers 60_1 to 60_m (m).

[0097] Application servers 50_1 to 50_n may include any one or any combination of processors 51_1 to 51_n, memory 52_1 to 52_n, switches 53_1 to 53_n, NICs 54_1 to 54_n, and storage devices 55_1 to 55_n. Processors 51_1 to 51_n may control all operations of application servers 50_1 to 50_n, access memory 52_1 to 52_n, and execute instructions and / or data loaded in memory 52_1 to 52_n. Non-limiting examples of memory 52_1 to 52_n may include DDR SDRAM, high-bandwidth memory (HBM), hybrid memory cube (HMC), dual in-line memory modules (DIMM), Optane DIMM, or non-volatile DIMM (NVDIMM).

[0098] According to embodiments, the number of processors and the number of memories included in application servers 50_1 to 50_n can be selected differently depending on the embodiment. In some embodiments, processors 51_1 to 51_n and memories 52_1 to 52_n can provide processor-memory pairs. In some embodiments, the number of processors 51_1 to 51_n may be different from the number of memories 52_1 to 52_n. Processors 51_1 to 51_n may include single-core processors or multi-core processors. In some embodiments, such as Figure 12 As shown by the dashed lines, storage devices 55_1 to 55_n in application servers 501 to 50n can be omitted. The number of storage devices 55_1 to 55_n included in storage servers 50_1 to 50_n can be selected differently depending on the embodiment. Processors 51_1 to 51_n, memories 52_1 to 52_n, switches 53_1 to 53_n, NICs 54_1 to 54_n and / or storage devices 55_1 to 55_n can communicate with each other through the links described above with reference to the accompanying drawings.

[0099] Storage servers 60_1 to 60_m may include any one or any combination of processors 61_1 to 61_m, memory 62_1 to 62_m, switches 63_1 to 63_m, NICs 64_1 to 64_n, and storage devices 65_1 to 65_m. Processors 61_1 to 61_m and memory 62_1 to 62_m may operate similarly to processors 51_1 to 51_n and memory 52_1 to 52_n of application servers 50_1 to 50_n described above.

[0100] Application servers 50_1 to 50_n can communicate with storage servers 60_1 to 60_m via network 70. In some embodiments, network 70 can be implemented using Fibre Channel (FC) or Ethernet. FC can be a medium for relatively high-speed data transmission. Optical switches providing high performance and high availability can be used as FC. Depending on the access method of network 70, storage servers 60_1 to 60_m can be provided as file storage, block storage, or object storage.

[0101] In some embodiments, network 70 may be a storage-only network such as a Storage Area Network (SAN). For example, the SAN may be an FC-SAN that can use an FC network and can be implemented using the FC protocol (FCP). In another case, the SAN may be an Internet Protocol (IP)-SAN, which uses a Transmission Control Protocol / Internet Protocol (TCP / IP) network and is implemented according to the SCSI over TCP / IP or Internet SCSI (iSCSI) protocol. In some embodiments, network 70 may be a general-purpose network such as a TCP / IP network. For example, network 70 may be implemented according to protocols such as FC over Ethernet (FCoE), network attached storage (NAS), or fabric-based high-speed non-volatile memory (NVMe) (NVMe over fabrics, NVMe-oF).

[0102] The following descriptions will primarily focus on application server 50_1 and storage server 60_1. However, it should be noted that the description of application server 50_1 can also be applied to another application server (e.g., 50_n), and the description of storage server 60_1 can also be applied to another storage server (e.g., 60_m).

[0103] Application server 50_1 can store data requested by users or clients in one of storage servers 60_1 to 60_m via network 70. Additionally, application server 50_1 can retrieve data requested by users or clients from one of storage servers 60_1 to 60_m via network 70. For example, application server 50_1 can be implemented as a web server or a database management system (DBMS).

[0104] Application server 50 can access the memory 52_n and / or storage device 55_n included in another application server 50 via network 70, and / or access the memory 62_1 to 62_m and / or storage device 65_1 to 65_m included in storage servers 60_1 to 60_m via network 70. Therefore, application server 50_1 can perform various operations on data stored in application servers 50_1 to 50_n and / or storage servers 60_1 to 60_m. For example, application server 50_1 can execute instructions to migrate or copy data between application servers 50_1 to 50_n and / or storage servers 60_1 to 60_m. In this case, data can be migrated from the memory 62_1 to 62_m of storage servers 60_1 to 60_m to the memory 52_1 to 52_n of application servers 50_1 to 50_n, or directly, via the memory 62_1 to 62_m of storage servers 60_1 to 60_m. In some embodiments, for security or privacy reasons, the data migrated over network 70 may be encrypted data.

[0105] In storage server 60_1, the interface IF can provide the physical connection between processor 61_1 and controller CTRL, and between NIC 64_1 and controller CTRL. For example, the interface IF can be implemented using a Direct Attached Storage (DAS) method where storage device 65_1 is directly connected to a dedicated cable. For example, the interface IF can be implemented using various interface methods such as: Advanced Technology Attachment (ATA), Serial ATA (SATA), External SATA (e-SATA), Small Computer System Interface (SCSI), Serial Attached SCSI (SAS), PCI, PCIe, NVMe, IEEE 1394, Universal Serial Bus (USB), Secure Digital (SD) card, Multimedia Card (MMC), Embedded MMC (eMMC), UFS, Embedded UFS (eUFS), and Compact Flash (CF) card interfaces.

[0106] In storage server 60_1, switch 63_1 can selectively connect processor 61_1 to storage device 65_1 or selectively connect NIC 64_1 to storage device 65_1 based on the control of processor 61_1.

[0107] In some embodiments, NIC 64_1 may include a network interface card (NIC) and a network adapter. NIC 64_1 can connect to network 70 via a wired interface, a wireless interface, a Bluetooth interface, or an optical interface. NIC 64_1 may include internal memory, a digital signal processor (DSP), and a host bus interface, and is connected to processor 61_1 and / or switch 63_1 via the host bus interface. In some embodiments, NIC 64_1 may be integrated with any one or any combination of processor 61_1, switch 63_1, and storage device 65_1.

[0108] In application servers 50_1 to 50_n or storage servers 60_1 to 60_m, processors 51_1 to 51_m and 61_1 to 61_n can send commands to storage devices 55_1 to 55_n and 65_1 to 65_m or memories 52_1 to 52_n and 62_1 to 62_m to program or read data. In this case, the data may be data that has been corrected for errors by an error-correcting code (ECC) engine. The data may be data processed by data bus inversion (DBI) or data masking (DM) and may contain cyclic redundancy check (CRC) information. For security or privacy, the data may be encrypted.

[0109] In response to read commands received from processors 51_1 to 51_m and 61_1 to 61_n, storage devices 55_1 to 55_n and 65_1 to 65_m can send control signals and command / address signals to the non-volatile memory device (e.g., a NAND flash memory device) NVM. Therefore, when reading data from the non-volatile memory device NVM, a read enable signal can be input as a data output control signal to output data to the DQ bus. The read enable signal can be used to generate a data strobe signal. Command and address signals can be latched based on the rising or falling edge of the write enable signal.

[0110] The controller CTRL can control all operations of storage device 65_1. In embodiments, the controller CTRL may include static RAM (SRAM). The controller CTRL can write data to the non-volatile storage device NVM in response to a write command, or read data from the non-volatile storage device NVM in response to a read command. For example, write commands and / or read commands can be generated based on requests provided from a host (e.g., processor 61_1 of storage server 60_1, processor 61_m of another storage server 60_m, or processors 51_1 to 51_n of application servers 50_1 to 50_n). The buffer BUF can temporarily store (or buffer) data to be written to or read from the non-volatile storage device NVM. In some embodiments, the buffer BUF may include DRAM. The buffer BUF may store metadata. Metadata may refer to user data or data generated by the controller CTRL for managing the non-volatile storage device NVM. For security or privacy, storage device 65_1 may include a secure element (SE).

[0111] Although the inventive concept has been shown and described with reference to embodiments thereof, it should be understood that various changes in form and detail may be made without departing from the spirit and scope of the appended claims.

Claims

1. A memory device configured to communicate with a plurality of host devices via interconnects, the memory device comprising: The memory includes multiple storage areas, the multiple storage areas including a first storage area allocated to a first host device among the multiple host devices and a second storage area allocated to a second host device among the multiple host devices; as well as A direct memory access (DMA) engine is configured to perform operations based on a request from the first host device, wherein the request includes a copy command to copy data stored in the first storage area to a second storage area, and the operation includes: Read the stored data from the first storage area; and The read data is written to the second storage area, instead of being output to the interconnect.

2. The memory device according to claim 1, wherein, The plurality of storage regions are logical devices that are identified as different devices by the plurality of host devices.

3. The memory device according to claim 1, wherein, The DMA engine is also configured to: read the stored data from the first storage area based on a request from the first host device corresponding to an operation that copies the stored data to another memory device that is physically different from the memory device, and output the read data to the interconnect.

4. The memory device according to claim 1, wherein, The copy command is defined differently from the normal read and write commands of the memory.

5. The memory device according to claim 4, wherein, The request, including the copy command, includes an indication of the read address of the first storage region and an indication of the write address of the second storage region.

6. The memory device according to claim 1, wherein, The DMA engine is also configured to send a response including a message to the first host device upon completion of writing the read data to the second storage area, the message indicating completion of copying the data stored in the first storage area to the second storage area.

7. The memory device according to claim 1, wherein, The DMA engine is also configured to communicate with the first host device and the second host device via a root union including a first root port and a second root port, and The DMA engine is connected to the first root port via a first virtual channel and to the second root port via a second virtual channel.

8. The memory device of claim 1, further comprising a command executor configured to: Based on the request from the first host device, execute the command; and Based on the command being executed, the DMA engine is controlled to copy the data stored in the first storage area to the second storage area, without outputting the data to the interconnect.

9. The memory device according to claim 1, wherein, The memory includes a Type 3 pooled memory as defined in the computation fast link protocol.

10. The memory device according to claim 1, wherein, The multiple storage regions include multiple non-volatile memory chips. The first storage region includes any one of the plurality of non-volatile memory chips, and The second storage region includes any other non-volatile memory chip among the plurality of non-volatile memory chips.

11. A method of operating a memory device, the memory device being configured to communicate with a plurality of host devices via interconnects, the memory device including a plurality of storage regions, the plurality of storage regions including a first storage region allocated to a first host device among the plurality of host devices and a second storage region allocated to a second host device among the plurality of host devices, the method comprising: Receive a request from the first host device, the request including a copy command to copy data stored in the first storage area to the second storage area; Based on the received request, the stored data is read from the first storage area; as well as The read data is written to the second storage area, instead of being output to the interconnect.

12. The method of claim 11, further comprising: Based on a request from the first host device corresponding to an operation to copy the stored data to another memory device that is physically different from the memory device, the following operations are performed: Read the stored data from the first storage area; and The read data is output to the interconnect.

13. The method of claim 11, further comprising: Based on the completion of writing the read data into the second storage area, a response including a message is sent to the first host device, the message indicating that the copying of the data stored in the first storage area to the second storage area is complete.

14. The method according to claim 11, wherein, The memory device includes a Type 3 pooled memory as defined in the compute fast link protocol.

15. The method according to claim 11, wherein, It also includes communication with the first host device and the second host device based on the Peripheral Component Interconnect Protocol or the Fast Peripheral Component Interconnect Protocol.

16. The method of claim 11, further comprising: Communicating with the first host device and the second host device through a root union including the first root port and the second root port. The memory device is connected to the first root port via a first virtual channel and to the second root port via a second virtual channel.

17. A host system, the host system comprising: A root association, comprising a first root port and a second root port, the root association being configured to provide interconnects based on a predetermined protocol; A first host device is configured to communicate with a memory device via a first root port, wherein a first storage area of ​​the memory device corresponding to a first logical device is allocated to the first host device. as well as A second host device is configured to communicate with the memory device via the second root port, wherein a second storage area of ​​the memory device corresponding to the second logical device is allocated to the second host device. The host system is configured to: receive a response from the memory device indicating that the copying of the data stored in the first memory area to the second memory area has been completed, based on a request sent by the first host device to the memory device to copy data stored in the first memory area to the second memory area, without receiving the data from the memory device through the root union.

18. The host system according to claim 17, wherein, Both the first host device and the second host device include any one or any combination of a central processing unit, a graphics processing unit, a neural processing unit, a field-programmable gate array, a network interface card, and peripheral devices configured to perform communication based on a compute fast link protocol.

19. The host system according to claim 17, wherein, The first host device is also configured to communicate with the memory device via the first root port and the first virtual channel, and The second host device is further configured to communicate with the memory device via the second root port and the second virtual channel.

20. The host system according to claim 17, wherein, The request to copy the stored data includes a copy command, an instruction for the read address of the first storage area, and an instruction for the write address of the second storage area.