Data copying method, CXL equipment, computer system and storage medium
By completing data copying between different memories within the CXL device, the problems of low data copying efficiency, large delay and CPU computing power occupation in the prior art are solved, and a more efficient data copying process and better system performance are achieved.
Patent Information
- Application Number
- CN202311713697.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-13
- Publication Date
- 2025-06-13
AI Technical Summary
When CXL devices copy data between different memories, they are less efficient and have long paths, resulting in large delays and CPU computing power occupied.
Data copy is implemented within the CXL device, the host's data copy request is received through the processor, and the data copy from the source memory to the destination memory is completed within the device, and finally the host's copy is notified to be completed.
Shorten the data copy path, reduce latency, and offload the CPU's workload, freeing up the CPU's computing power.
Smart Images

Figure CN120144330A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to, but is not limited to, storage technologies, and particularly to a data copy method, a CXL device, a computer system, and a storage medium. Background Art
[0002] CXL (Compute Express Link) technology is a new type of high-speed interconnection technology that supports high-bandwidth, low-latency data transmission, has better flexibility and scalability, and can achieve the mixed use of different types of hardware devices. The application scenarios of CXL technology are very extensive, including data centers, artificial intelligence, and processor interconnection. In the field of data centers, CXL technology can interconnect different computing and storage resources to improve system performance and efficiency; in the field of artificial intelligence, CXL technology can enable accelerators such as GPUs and FPGAs to better cooperate with the main processor to improve the speed of AI model training and inference; in terms of processor interconnection, CXL technology can achieve the interconnection between processors of different manufacturers to improve the overall performance and flexibility of the system.
[0003] A CXL device can expose device memory to the host based on the CXL.mem protocol, that is, as host managed device memory (HDM). The host can copy data between device memories, but the efficiency is not high. Summary of the Invention
[0004] Embodiments of the present disclosure provide a CXL device, including a processor, a device cache, a first memory, and a second memory, characterized in that the processor is configured to perform the following processing:
[0005] Receiving a data copy request sent by the host, the data copy request carrying information of a source address, a destination address, and a data length, the source address being the address of the first memory, and the destination address being the address of the second memory;
[0006] Copying the data stored at the source address to the destination address through operations inside the device, and notifying the host that the data copy is completed.
[0007] Embodiments of the present disclosure also provide a computer system, including a host and the CXL device according to any embodiment of the present disclosure.
[0008] Embodiments of the present disclosure also provide a data copy method, applied to a CXL device including a first memory and a second memory, the data copy method including:
[0009] Receive a data copy request sent by the host, where the data copy request carries information about the source address, destination address, and data length. The source address is the address of the first memory, and the destination address is the address of the second memory;
[0010] Through operations within this device, copy the data stored at the source address to the destination address, and notify the host that the data copy is complete.
[0011] The embodiments of the present disclosure also provide a non-transitory computer storage medium storing a computer program, which when executed by a processor, can implement the data copy method described in any embodiment of the present disclosure.
[0012] After the CXL device in the above embodiments of the present disclosure receives a data copy request from the host, through operations within this device, copy the data stored at the source address to the destination address, and return a response indicating that the data copy is complete to the host. During this data copy process, there is no need to transfer data between the CXL device and the host, which can shorten the data copy path, reduce latency, and at the same time unload the CPU workload and release the CPU computing power.
[0013] Other features and advantages of the present disclosure will be described in the subsequent specification, and some of them will become obvious from the specification or be understood by implementing the present disclosure. Other advantages of the present disclosure can be achieved and obtained through the solutions described in the specification and the accompanying drawings. Description of the Drawings
[0014] The drawings are used to provide an understanding of the technical solutions of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the technical solutions of the present disclosure and do not constitute a limitation to the technical solutions of the present disclosure.
[0015] Figure 1 It is a schematic diagram for an embodiment to implement data copy between different memories of a CXL device;
[0016] Figure 2 It is a structural block diagram of a CXL device according to an embodiment of the present disclosure;
[0017] Figure 3 It is a schematic diagram for an embodiment to implement data copy between different memories of a CXL device;
[0018] Figure 4 It is a flowchart of a data copy method according to an embodiment of the present disclosure. Detailed Embodiments
[0019] The present disclosure describes multiple embodiments, but the description is exemplary rather than restrictive, and it will be apparent to those of ordinary skill in the art that there can be more embodiments and implementation solutions within the scope of the embodiments described in the present disclosure. Although many possible feature combinations are shown in the drawings and discussed in the detailed description, many other combinations of the disclosed features are also possible. Unless specifically restricted, any feature or element of any embodiment can be used in combination with any other feature or element in any other embodiment, or can replace any other feature or element in any other embodiment.
[0020] The present disclosure includes and contemplates combinations with features and elements known to those of ordinary skill in the art. The embodiments, features, and elements already disclosed in the present disclosure can also be combined with any conventional features or elements to form an inventive solution protected by the present disclosure. Any feature or element of any embodiment can also be combined with features or elements from other inventive solutions to form another inventive solution protected by the present disclosure. Therefore, it should be understood that any feature shown and / or discussed in the present disclosure can be implemented alone or in any suitable combination. Therefore, the embodiments are not subject to other limitations except those made in accordance with the appended claims and their equivalents. In addition, various modifications and changes can be made within the scope of the appended claims.
[0021] In addition, when describing representative embodiments, the specification may have presented the method and / or process as a particular sequence of steps. However, to the extent that the method or process does not depend on the particular order of the steps described herein, the method or process should not be limited to the particular order of steps described. As will be understood by those of ordinary skill in the art, other step sequences are possible. Therefore, the particular order of steps set forth in the specification should not be construed as a limitation on the claims. In addition, the claims directed to the method and / or process should not be limited to performing their steps in the order written, and those skilled in the art can readily understand that these orders can vary and still remain within the spirit and scope of the embodiments of the present disclosure.
[0022] In CXL devices, there are Type 1 CXL devices (abbreviated as Type 1 devices), Type 2 CXL devices (abbreviated as Type 2 devices), and Type 3 CXL devices (abbreviated as Type 3 devices). Among them, Type 2 devices and Type 3 devices can be configured with multiple device memories, which can be of types such as DDR, HBM (High Bandwidth Memory), etc. Taking Type 2 devices as an example, Type 2 devices can be acceleration cards (such as GPUs) in actual application scenarios. The Host provides data, and the acceleration card is responsible for computing. The HDM model allows the host to directly manage and control the storage space on Type 2 devices. Dynamically allocate and configure device memory according to application requirements, and directly perform read and write operations.
[0023] When the host determines to perform data copying between multiple device memories inside a Type 2 device, as Figure 1 shown, the host 2 sends a memory data copy request (Memcpy) to the CXL device 1. The CXL device 1 includes a control chip and multiple DARM chips, and this data copy is to be performed between different DARM chips of this CXL device 1. The CXL device needs to send the data (Data) in one DRAM chip to the host according to the source address, store it in the host cache, and then the host writes the data from the host cache to another DRAM chip according to the destination address. The data copy path is long and the CPU will be blocked during the copy process, resulting in low efficiency. Especially in the case of large - block data copying (such as data of one or more memory pages, greater than or equal to 4KB), limited by the memory bus bandwidth (such as the maximum bus bandwidth is 16 bytes) and the size of the CPU cache, the data is still copied byte by byte, which requires a large number of CPU clock cycles to complete the data copy, and the copy latency is large. The CXL interface and memory interface of the control chip are omitted in the figure.
[0024] To solve the problem of low efficiency in the data copy process, the embodiments of the present disclosure complete data copying between different memories inside the CXL device to shorten the data copy path, reduce latency, and save the resources of the host CPU. The embodiments of the present disclosure can be applied to large - block data copying, but are not limited thereto.
[0025] An embodiment of the present disclosure provides a CXL device, including a processor, a device cache, a first memory, and a second memory. The processor is configured to perform the following processing:
[0026] Upon receiving a data copy request sent by the host, the data copy request carries information about the source address, destination address, and data length. The source address is the address of the first memory, and the destination address is the address of the second memory;
[0027] Through operations within this device, copy the data stored at the source address to the destination address, and notify the host that the data copy is complete.
[0028] In an example of this embodiment, the CXL device is a Type 2 device (such as an acceleration card), and both the first memory and the second memory are device memories managed by the host. As Figure 1 shown, the CXL device includes a control chip 11 and multiple DRAM chips 12. The control chip 11 is connected to the DRAM chips (such as DDR, HBM) 12 through a memory interface 113 (such as a DRAM controller), and realizes interaction with the host based on the CXL.io protocol, CXL.mem protocol, and CXL.cache protocol through the CXL interface 114. A processor 111 and a device cache 112 are also integrated in the control chip 11. In an example, the processor 111 is a CPU core, and the device cache 112 can be implemented with SDRAM. The memory in the CXL device can include one or more of types of memories such as DRAM, FLASH, and SSD. The Bias management unit 115 and the DMA engine 116 in the control chip 11 are optional and will be described below.
[0029] In an example of this embodiment, the information of the source address can include the starting address and data length of the data carried in the data copy request in the first memory. Based on this starting address and data length, all the addresses of the data to be copied this time in the first memory, that is, the source address of the data, can be determined. Similarly, the destination address of the data can be determined by adding the data length to the starting address of the second memory where the data carried in the data copy request is to be written. The addresses carried in the data copy request can be the physical addresses of the device memory or the host-side bus domain addresses. The CXL device can convert the host-side bus domain addresses into the physical addresses of the device memory to access the first memory and the second memory. After CXL completes the data copy, it can send a copy completion message to the host in an interrupt manner to notify that the data copy is complete and the host can normally access the data at the source address and destination address in the CXL device.
[0030] After the CXL device of the embodiment of the present disclosure receives a data copy request (Memcpy) from the host, through operations within this device, copy the data stored at the source address to the destination address, and return a response (Copy Complete Response) indicating that the data copy is complete to the host. As Figure 3As shown, during the data copy process, there is no need to transfer data between the CXL device 1 and the host 2, which can shorten the data copy path and reduce latency. At the same time, it can offload the workload of the CPU and release the computing power of the CPU.
[0031] Data copy within the CXL device can be solved in two ways. The first is the DMA method, that is, data transfer is completed through specific hardware such as a DMA controller, without the participation of the processor (such as a CPU) inside the CXL device and without involving the device cache. The other method requires the participation of the processor inside the CXL device. The data stream is read from the source address of the first memory to the device cache and then written from the device cache to the target address of the second memory.
[0032] In an exemplary embodiment of the present disclosure, the CXL device further includes a DMA controller. The processor copies the data stored at the source address to the destination address through operations inside the device, including: sending a data copy instruction to the DMA controller, carrying the source address, destination address, and data length, and copying the data in the first memory stored at the source address to the destination address of the second memory in a DMA manner through the DMA controller. DMA (Direct Memory Access) means direct memory access. When it is necessary to copy data from one memory to another, the processor can initialize this transfer operation, such as setting the address and registers of the DMA channel, the direction of data transfer, etc., and then instruct the DMA controller (such as Figure 2 the DMA engine 116 in [reference] to start the transfer. The transfer itself is implemented by the DMA controller, and the DMA controller notifies the processor after the transfer is completed.
[0033] In an exemplary embodiment of the present disclosure, the processor copies the data stored at the source address to the destination address through operations inside the device, including: reading the data stored at the source address and writing it to the device cache, and then writing the data from the device cache to the storage space corresponding to the destination address. This embodiment is a method of implementing data copy between different memories inside the device based on the device cache.
[0034] In the above two implementation methods, the data stream during data copy does not pass through the cache at the host end, so both can shorten the data copy path and reduce latency.
[0035] In an exemplary embodiment of the present disclosure, large - block data copying between different memories inside a device is implemented based on device caching, and data pre - fetching is used to improve the efficiency of data copying. In this embodiment, both the first memory and the second memory are device memories, such as DRAM memories like DDR and HBM. The processor reads the data stored at the source address and writes it into the device cache, and then writes the data from the device cache into the storage space corresponding to the destination address, including: when the data volume of the data is greater than or equal to the capacity of a set number of memory pages, after starting the copy, the copy operation and the pre - fetch operation are executed in parallel. The set number can be 1, 2, 3 or more, and the data volume of the number and the capacity of the memory page can be in bytes. Among them:
[0036] The copy operation includes: sequentially searching for data blocks included in the data in the device cache according to the source address. If found, the found data block is read, and the data block is written into the second memory according to the destination address; if not found, the data block is read from the first memory according to the source address and written into the device cache, and then the data block is written from the device cache into the second memory according to the destination address;
[0037] The pre - fetch operation includes: starting from a specified address in the first memory, reading data blocks included in the data and writing them into the device cache until the data reading is completed. The specified address is determined according to the starting address of the source address and a set offset, and the size of the data block is an integer multiple of the memory page size, such as 4KB, 8KB, etc.
[0038] For the copying of large - block data, data pre - fetching is utilized inside the device in this embodiment, which can improve the memory copying efficiency. When the device cache is divided into multiple layers, such as L1 layer and L2 layer, the pre - fetch operation can also occur between different layers of the device cache. For example, data to be copied in the L2 layer can be pre - fetched and written into the L1 layer simultaneously. When the processor executes the copy operation, it will first search for the data to be copied in the L1 layer, and if not found, it will search in the L2 layer. Therefore, this pre - fetch operation can also improve the memory copying efficiency. If the device cache also includes an L3 layer, data in the L3 layer can also be pre - fetched and written into the L2 layer simultaneously, that is, the pre - fetch of each layer is parallelized to maximize the system effect.
[0039] In an exemplary embodiment of the present disclosure, the CXL device further includes a Bias management unit; after receiving a data copy request sent by the host, the processor switches the Bias mode of the memory pages corresponding to the source address and the target address from host Bias to device Bias through the Bias management unit and notifies the host, maintains the data consistency between the device cache and the device memory, and after completing the data copy, switches the Bias mode of the memory pages corresponding to the source address and the target address from device Bias to host Bias through the Bias management unit and notifies the host.
[0040] The cache coherence between the host and the CXL device (i.e., the coherence of the data in the host cache and the device cache with the corresponding data in the device memory) can be maintained through the Bias mode. The Bias mode is divided into host bias (i.e., the host bias mode) and device Bias (i.e., the device bias mode) according to different usage scenarios. The Bias mode maintains cache coherence in units of memory pages. A memory page is either in the Host bias mode or the Device Bias mode at a certain moment. In the Device Bias mode, it is preferred that the host side should not cache any content for this memory page, or the Bias hardware inside the device can deprive the host of the cache for this page (i.e., notify the host to give up the cache for the memory page to be accessed currently). In the Host Bias mode, the host accesses the device memory of the CXL device, and there should be no access or change to this memory inside the CXL device; if there is, an instruction needs to be notified to apply to the host for access to this memory page, and the host maintains the memory coherence; and the host returns the content of this section of memory to the device instead of directly accessing this section of the memory page at the device side.
[0041] In this embodiment, the conversion of the Bias mode is performed using the hardware in the CXL device (such as Figure 2It is implemented by the Bias management unit 115). After the processor receives a data copy request sent by the host, it notifies the Bias management unit 115 to switch the Bias mode of the memory pages to be accessed for this data copy (including the memory pages storing the data in the first memory and the memory pages of the second memory to which the data is to be written) from the host Bias to the device Bias and notifies the host. After receiving this notification, the host abandons the caching of the data in these memory pages; after the data copy is completed, the Bias mode of the memory pages accessed for this data copy is switched from the device Bias to the host Bias and the host is notified. After receiving this notification, the host can re-cache the data in these memory pages. In this embodiment, the function of switching the Bias mode is set to be implemented on the CXL device side, which can achieve better system efficiency. The CXL device has a clearer understanding of the current use of the memory pages. For example, it can evaluate whether these memory pages will be accessed within a period of time according to the service characteristics, etc., so as to easily achieve precise control.
[0042] In an example of this embodiment, the processor maintains the data consistency between the device cache and the device memory, including:
[0043] In the case of using the data stored in the device cache for calculation during the data copy process:
[0044] If the calculation is completed before the Bias mode of the memory pages accessed for this data copy is switched from the device Bias to the host Bias, then after the data copy is completed, according to the source address and the destination address, the updated data in the device cache is written back to the first memory and the second memory, and then the Bias mode of the memory pages accessed for this data copy is switched from the device Bias to the host Bias;
[0045] If the calculation is completed after the Bias mode of the memory pages accessed for this data copy is switched from the device Bias to the host Bias, a request is sent to the host, and after obtaining the permission of the host, the updated data in the device cache is written back to the first memory and the second memory according to the source address and the destination address.
[0046] In this example, when the memory page to be accessed during this data copy is in the Device Bias state, if the CXL device accesses the data being copied using the cache, and if the memory page is still in the Device Bias mode when the calculation is completed, the device maintains consistency. If the current host does not have the current content in the cache, then after the data copy is completed, the updated data in the cache is written to the corresponding DRAM, and then the Bias mode is switched. If the memory page is in the host Bias mode when the calculation is completed, an application needs to be made to the host, and the cache write-back operation is completed under the coherence protocol.
[0047] In an exemplary embodiment of the present disclosure, after receiving the data copy request, the processor first determines whether the data length meets the set conditions. If the set conditions are met, then through operations within the device, the data stored at the source address is copied to the destination address; where the set conditions include any one or more of the following conditions: the data length is greater than or equal to a set threshold, and the data length is an integer multiple of the data length stored in a memory page. The different processing methods adopted in this embodiment according to the size of the data to be copied are as follows: for large chunks of data, such as when the data length is greater than or equal to the set threshold, the data stored at the source address is copied to the destination address through operations within the device; for less data, the data copy can be completed under the leadership of the host, and the CXL device does not need to perform processing such as Bias mode conversion, achieving good comprehensive performance.
[0048] After receiving the data copy instruction between different memories of the device sent by the host, the CXL device in the above embodiment of the present disclosure can complete the data copy inside the device, reducing the data copy path latency, while offloading the workload of the CPU and releasing the computing power of the CPU. In some embodiments, cache coherence during the data copy process can be maintained based on the Bias-based model, and the cache coherence during the data copy process is ensured through the CXL.cache protocol. When using the device cache to complete the data copy operation, large chunks of data can also be prefetched inside the device to improve the data copy efficiency.
[0049] An embodiment of the present disclosure also provides a computer system, which can be seen Figure 3 and includes a host and the CXL device described in any embodiment of the present disclosure.
[0050] An embodiment of the present disclosure also provides a data copy method, which is applied to a CXL device including a first memory and a second memory, as Figure 4 shown, and the data copy method includes:
[0051] Step 110: Receive a data copy request sent by the host. The data copy request carries information about the source address, destination address, and data length. The source address is the address of the first memory, and the destination address is the address of the second memory.
[0052] Step 120: Through operations within this device, copy the data stored at the source address to the destination address, and notify the host that the data copy is complete.
[0053] During this data copy process, there is no need to transfer data between the CXL device and the host, which can shorten the data copy path and reduce latency. At the same time, it can offload the CPU's workload and release the CPU's computing power.
[0054] In an exemplary embodiment of the present disclosure, the step of copying the data stored at the source address to the destination address through operations within this device includes:
[0055] Copy the data in the first memory stored at the source address to the destination address of the second memory in DMA mode; or
[0056] Read the data stored at the source address and write it into the device cache, and then write the data from the device cache into the storage space corresponding to the destination address.
[0057] In an exemplary embodiment of the present disclosure, after receiving the data copy request sent by the host, the method further includes: switching the Bias mode of the memory page to be accessed in this data copy from the host Bias to the device Bias and notifying the host, maintaining the data consistency between the device cache and the device memory, and after completing the data copy, switching the Bias mode of the memory page accessed in this data copy from the device Bias back to the host Bias and notifying the host.
[0058] In an example of this embodiment, the maintaining of the data consistency between the device cache and the device memory includes: in the case of using the data stored in the device cache for calculation during the data copy process:
[0059] If the calculation is completed before the Bias mode of the memory page accessed in this data copy is switched from the device Bias to the host Bias, then after the data copy is completed, update the data in the device cache back to the first memory and the second memory according to the source address and the destination address, and then switch the Bias mode of the memory page accessed in this data copy from the device Bias to the host Bias;
[0060] If the calculation is completed after the Bias mode of the memory page accessed in this data copy is converted from device Bias to host Bias, a request is sent to the host. After obtaining the permission of the host, the updated data in the device cache is written back to the first memory and the second memory according to the source address and the destination address.
[0061] In an exemplary embodiment of the present disclosure, after receiving the data copy request, the method further includes: first determining whether the data length meets a set condition. If the set condition is met, the data stored at the source address is copied to the destination address through an operation inside the device; where the set condition includes any one or more of the following conditions: the data length is greater than or equal to a set threshold, and the data length is an integer multiple of the data length stored in a memory page.
[0062] An embodiment of the present disclosure also provides a non-transitory computer storage medium storing a computer program, characterized in that when the computer program is executed by a processor, the data copy method described in any embodiment of the present disclosure can be implemented.
[0063] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof. In the hardware implementation, the division of the functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component can have multiple functions, or a function or step can be executed by several physical components in cooperation. Some or all components can be implemented as software executed by a processor, such as a digital signal processor or a microprocessor, or be implemented as hardware, or be implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cassette, tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium generally includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.
Claims
1. A CXL device, comprising a processor, a device cache, a first memory, and a second memory, characterized in that, the processor is configured to perform the following processing: receiving a data copy request sent by a host, the data copy request carrying information of a source address, a destination address, and a data length, the source address being an address of the first memory, and the destination address being an address of the second memory; copying the data stored at the source address to the destination address through an operation inside the device, and notifying the host that the data copy is completed.
2. The CXL device according to claim 1, characterized in that: the CXL device further comprises a DMA controller; the processor copying the data stored at the source address to the destination address through an operation inside the device includes: copying the data in the first memory stored at the source address to the destination address of the second memory in a DMA manner through the DMA controller.
3. The CXL device according to claim 1, characterized in that: the processor copying the data stored at the source address to the destination address through an operation inside the device includes: reading the data stored at the source address and writing it into the device cache, and then writing the data from the device cache into the storage space corresponding to the destination address.
4. The CXL device according to claim 3, characterized in that: both the first memory and the second memory are device memories, and the processor reading the data stored at the source address and writing it into the device cache, and then writing the data from the device cache into the storage space corresponding to the destination address includes: when the data volume of the data is greater than or equal to the capacity of a set number of memory pages, starting data copy and parallelly executing a copy operation and a prefetch operation, wherein: the copy operation includes: sequentially searching for data blocks included in the data in the device cache according to the source address, if found, reading the found data blocks, and writing the data blocks into the second memory according to the destination address; if not found, reading the data blocks from the first memory according to the source address and writing them into the device cache, and then writing the data blocks from the device cache into the second memory according to the destination address; the prefetch operation includes: reading data blocks included in the data from a specified address of the first memory and writing them into the device cache until the data reading is completed, the specified address being determined according to the starting address of the source address and a set offset, and the size of the data blocks being an integer multiple of the memory page size.
5. The CXL device according to claim 1, characterized in that: The CXL device further includes a Bias management unit; after receiving a data copy request sent by the host, the processor switches the Bias mode of the memory page to be accessed for this data copy from the host Bias to the device Bias through the Bias management unit and notifies the host, maintains the data consistency between the device cache and the device memory, and after completing the data copy, switches the Bias mode of the memory page accessed for this data copy from the device Bias back to the host Bias through the Bias management unit and notifies the host.
6. The CXL device according to claim 5, wherein: The processor maintains the data consistency between the device cache and the device memory, including: In the case of using the data stored in the device cache for calculation during the data copy process: If the calculation is completed before the Bias mode of the memory page accessed for this data copy is switched from the device Bias to the host Bias, then after the data copy is completed, the updated data in the device cache is written back to the first memory and the second memory according to the source address and the destination address, and then the Bias mode of the memory page accessed for this data copy is switched from the device Bias to the host Bias; If the calculation is completed after the Bias mode of the memory page accessed for this data copy is switched from the device Bias to the host Bias, then a request is sent to the host, and after obtaining permission from the host, the updated data in the device cache is written back to the first memory and the second memory according to the source address and the destination address.
7. The CXL device according to claim 1, wherein: After receiving the data copy request, the processor first determines whether the data length meets a set condition. If the set condition is met, then through an operation within the device, the data stored at the source address is copied to the destination address; wherein, the set condition includes any one or more of the following conditions: the data length is greater than or equal to a set threshold, and the data length is an integer multiple of the data length stored in one memory page.
8. The CXL device according to claim 1, wherein: The CXL device is a CXL Type2 device, and both the first memory and the second memory are device memories managed by the host.
9. A computer system, wherein, it includes a host and a CXL device according to any one of claims 1 to 8.
10. A data copy method, applied to a CXL device including a first memory and a second memory, the data copy method includes: Receiving a data copy request sent by the host, the data copy request carrying information of a source address, a destination address, and a data length, the source address being the address of the first memory, and the destination address being the address of the second memory; Through an operation within the device, copying the data stored at the source address to the destination address, and notifying the host that the data copy is completed.
11. The data copy method according to claim 10, wherein: The operation inside this device to copy the data stored at the source address to the destination address includes: Copying the data in the first memory stored at the source address to the destination address of the second memory in a DMA manner; or Reading the data stored at the source address and writing it into the device cache, and then writing the data from the device cache into the storage space corresponding to the destination address.
12. The data copy method according to claim 10, wherein: After receiving the data copy request sent by the host, the method further includes: switching the Bias mode of the memory page to be accessed in this data copy from the host Bias to the device Bias and notifying the host, maintaining the data consistency between the device cache and the device memory, and after completing the data copy, switching the Bias mode of the memory page accessed in this data copy from the device Bias back to the host Bias and notifying the host.
13. The data copy method according to claim 12, wherein: The maintaining of the data consistency between the device cache and the device memory includes: in the case of using the data stored in the device cache for calculation during the data copy process: If the calculation is completed before the Bias mode of the memory page accessed in this data copy is switched from the device Bias to the host Bias, then after the data copy is completed, according to the source address and the destination address, writing the updated data in the device cache back to the first memory and the second memory, and then switching the Bias mode of the memory page accessed in this data copy from the device Bias to the host Bias; If the calculation is completed after the Bias mode of the memory page accessed in this data copy is switched from the device Bias to the host Bias, then requesting the host, and after obtaining the permission of the host, writing the updated data in the device cache back to the first memory and the second memory according to the source address and the destination address.
14. The data copy method according to claim 10, wherein: After receiving the data copy request, the method further includes: first determining whether the data length meets the set conditions, and if it meets the set conditions, then through the operation inside this device, copying the data stored at the source address to the destination address; wherein, the set conditions include any one or more of the following conditions: the data length is greater than or equal to a set threshold, and the data length is an integer multiple of the length of the data stored in one memory page.
15. A non-transitory computer storage medium storing a computer program, wherein, when the computer program is executed by a processor, it can implement the data copy method according to any one of claims 10 to 14.