Memory expansion and memory access

WO2026189110A1PCT designated stage Publication Date: 2026-09-17CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2026/078517
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-10
Filing Date
2026-02-11
Publication Date
2026-09-17

Smart Images

  • Figure CN2026078517_17092026_PF_FP_ABST
    Figure CN2026078517_17092026_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present disclosure are a memory expansion method, a memory access method, a device, a storage medium and a program product. In the embodiments of the present disclosure, a memory expansion device is added to an acceleration device, and the memory expansion device is interconnected to a processor of the acceleration device and a main unit, such that a memory access channel can be formed. Thus, the processor of the acceleration device uses the memory of the main unit to expand the memory thereof. During the memory expansion, the address of the memory expansion device is mapped to the expanded memory, such that when the processor accesses the memory expansion device, the processor can access the expanded memory via the memory access channel, such that the memory of the main unit is used as a memory resource for expanding the memory of the processor. In this way, the memory implementation cost of the acceleration device can be reduced, the performance of acceleration services provided on the basis of the acceleration device can also be ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Memory Expansion and Memory Access Technical Field

[0001] This disclosure relates to the field of cloud computing technology, and in particular to memory expansion and memory access. Background Technology

[0002] With the continuous development of cloud computing technology, servers are increasingly relying on smart network interface cards (NICs) to provide acceleration services in areas such as networking, storage, and encryption in order to improve service quality and efficiency. As the number of processor cores in servers increases, the number of virtual machines that servers can support also gradually increases. This requires smart NICs to have sufficient memory capacity to ensure the performance of acceleration services.

[0003] To ensure the performance of the acceleration service, the memory capacity of the smart network interface card (NIC) is currently designed based on the maximum memory requirements of the acceleration service. However, in reality, the acceleration service does not reach the maximum memory requirements in most cases, and a portion of the smart NIC's memory remains idle most of the time, resulting in high memory implementation costs for the smart NIC.

[0004] To reduce memory implementation costs, smart network interface cards (NICs) can be designed with relatively low memory specifications. However, this approach carries the risk of insufficient memory, compromising the performance of acceleration services. Therefore, how to reduce memory implementation costs while ensuring the performance of acceleration services is a pressing technical challenge that needs to be addressed. Summary of the Invention

[0005] This disclosure provides a memory expansion method, a memory access method, a device, a storage medium, and a program product to reduce the memory implementation cost of acceleration devices while ensuring the performance of acceleration services provided by the acceleration devices.

[0006] This disclosure provides a computer device, including a host and an acceleration device; the acceleration device includes a processor and a memory expansion device; the memory expansion device is interconnected with both the processor and the host, and is configured to form a memory access channel between the processor and the host, so as to use the host's memory as a memory resource for memory expansion of the processor; the processor is configured to determine the range of memory that the memory expansion device can expand for the processor from the host's memory, and, if the processor meets the memory expansion conditions, request expanded memory from the host's memory not exceeding the memory range; and to perform address mapping on the expanded memory to access the expanded memory through the memory access channel.

[0007] This disclosure also provides a memory expansion method applied to an acceleration device. The processor and memory expansion device in the acceleration device cooperate to form a memory access channel between the processor and a host, so as to use the host's memory as a memory resource for memory expansion of the processor. The method includes: determining the range of memory that the memory expansion device can expand for the processor from the host's memory; if the processor memory expansion conditions are met, requesting expanded memory from the host's memory up to the range of memory; and performing address mapping on the expanded memory to access the expanded memory through the memory access channel.

[0008] This disclosure also provides a memory access method applied to an acceleration device, wherein a processor and a memory expansion device in the acceleration device cooperate to form a memory access channel between the processor and a host, so as to use the host's memory as a memory resource for memory expansion of the processor; the method includes: pre-allocating extended memory for the processor from the host's memory; and accessing the extended memory on the host based on the memory access channel.

[0009] This disclosure also provides an acceleration device, including: a processor, a memory expansion device, and local memory; the memory expansion device is interconnected with the processor and a host respectively, and is used to form a memory access channel between the processor and the host, so as to use the host's memory as a memory resource for memory expansion of the processor; the local memory stores a computer program, and the processor is used to run the computer program to implement the steps in the various methods provided in this disclosure.

[0010] This disclosure also provides a computer-readable storage medium storing a computer program that, when executed by a processor, enables the processor to perform the steps in the methods described above.

[0011] This disclosure also provides a computer program product, which includes a computer program / instructions that, when executed by a processor, enable the processor to perform the steps described in the method embodiments above.

[0012] In this embodiment of the disclosure, a memory expansion device is added to the acceleration device. The memory expansion device is interconnected with the processor and the host of the acceleration device to form a memory access channel. Then, the processor of the acceleration device uses the host's memory as its memory expansion. During the memory expansion process, the address of the memory expansion device is mapped to the expanded memory so that when the processor accesses the memory expansion device, it can access the expanded memory through the memory access channel. This realizes that the host's memory is used as the memory resource for memory expansion, which can not only reduce the memory implementation cost of the acceleration device, but also ensure the performance of the acceleration services provided by the acceleration device. Attached Figure Description

[0013] The accompanying drawings, which are included to provide a further understanding of this disclosure and form part of this disclosure, illustrate exemplary embodiments of this disclosure and are used to explain this disclosure, but do not constitute an undue limitation of this disclosure.

[0014] Figure 1 is a schematic diagram of the structure of a computer device provided in an exemplary embodiment of this disclosure.

[0015] Figure 2 is a schematic diagram of the structure of a computer device provided in another exemplary embodiment of this disclosure.

[0016] Figure 3 is a schematic diagram of the structure of a computer device provided in another exemplary embodiment of this disclosure.

[0017] Figure 4 is a flowchart illustrating a memory expansion method provided in another exemplary embodiment of this disclosure.

[0018] Figure 5 is a flowchart illustrating another memory access method provided in yet another exemplary embodiment of this disclosure.

[0019] Figure 6 is a schematic diagram of the structure of an acceleration device provided in another exemplary embodiment of this disclosure. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of this disclosure clearer, the technical solutions of this disclosure will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.

[0021] It should be noted that, in the cases involving user information in the embodiments of this disclosure, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this disclosure are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse. In addition, the various models involved in this disclosure (including but not limited to language models or large models) comply with relevant laws and standards.

[0022] In this embodiment, the acceleration device includes memory space, which is referred to simply as local memory. The local memory of the acceleration device is a pre-designed memory capacity according to corresponding memory specifications, and there is no limitation on the pre-designed memory capacity size. It should be understood that sufficient memory capacity is a key factor in ensuring the stability and performance of the acceleration service provided by the acceleration device. Nevertheless, in most cases, the acceleration device does not fully utilize the pre-designed maximum memory capacity during the provision of acceleration services, with some local memory remaining idle most of the time, resulting in a high memory resource idle rate. This not only wastes memory resources but also leads to higher memory implementation costs. Therefore, for cost considerations, the memory capacity of the acceleration device can be designed with relatively low memory specifications. However, this cost-reduction design approach carries the risk of insufficient memory.

[0023] To address the aforementioned technical problems, in this embodiment of the disclosure, a memory expansion device is added to the acceleration device. The memory expansion device is interconnected with both the processor and the host of the acceleration device, forming a memory access channel. Furthermore, the processor of the acceleration device utilizes the host's memory as its memory expansion resource. During the memory expansion process, the address of the memory expansion device is mapped to the expanded memory, so that when the processor accesses the memory expansion device, it can access the expanded memory via the memory access channel. This allows the host's memory to be used as the memory resource for memory expansion, which not only reduces the memory implementation cost of the acceleration device but also ensures the performance of the acceleration services provided by the acceleration device.

[0024] The technical solutions provided by the embodiments of this disclosure are described in detail below with reference to the accompanying drawings.

[0025] Figure 1 is a schematic diagram of a computer device provided in an exemplary embodiment of this disclosure. As shown in Figure 1, the computer device includes a host 10 and an acceleration device 11. The host 10 can be any electronic device with computing, storage, and communication functions, such as a computer, mobile phone, tablet, or other terminal devices, or a traditional server, cloud server, server cluster, or workstation, etc., without limitation. The acceleration device 11 refers to a hardware device designed to perform specific tasks, possessing specific hardware acceleration functions, thereby providing acceleration services to the host 10, flexibly offloading specific tasks of the host 10 to the acceleration device 11 for execution. In this embodiment, the type of acceleration device 11 is not limited, and the acceleration device 11 includes, but is not limited to, Smart NIC (Smart Network Interface Card), CIPU (Cloud Infrastructure Processing Unit), GPU (Graphics Processing Unit), and NPU (Neural Network Processing Unit).

[0026] In this embodiment, the acceleration device 11 includes a processor 111 and a memory expansion device 112 newly added to the acceleration device 11. The memory expansion device 112 is used to expand the memory of the acceleration device 11, and is interconnected with both the processor 111 and the host 10 of the acceleration device 11, thereby forming a memory access channel between the processor 111 and the host 10, so that the memory 101 of the host 10 can be used as the memory resource for memory expansion of the processor 111 of the acceleration device 11.

[0027] It should be noted that in the embodiments disclosed herein, when the term "processor" appears directly, it specifically refers to the processor 111 of the acceleration device 11. Unless otherwise specified, such as the processor of a host or similar cases, when the term "processor" appears directly without limitation, the processor specifically refers to the processor 111 of the acceleration device 11.

[0028] In this embodiment, the memory expansion device 112 utilizes the memory 101 of the host 10 to expand the memory of the acceleration device 11. The memory expansion device 112 can determine the range of memory it can expand from the memory 101 of the host 10 for the processor 111. In this embodiment, the implementation method for determining the range of memory that can be expanded from the memory 101 of the host 10 for the processor 111 is not limited. In an optional embodiment, the host memory can be expanded on demand. For example, if the host is a cloud server, in these cases, it is not necessary to obtain the capacity of the host memory; instructions can be directly issued to indicate the range of memory to be expanded for the processor. Alternatively, the memory expansion device can be pre-configured to provide the range of memory to be expanded for the processor. In another optional embodiment, capacity information such as the maximum capacity or free capacity of the memory 101 of the host 10 can be identified. Based on this capacity information, the memory capacity available for memory expansion in the memory 101 of the host 10 can be determined as the memory range in this embodiment. Optionally, the memory 101 of the host 10 and the range of memory available for memory expansion in the memory 101 of the host 10 can be identified by enumerating memory devices.

[0029] Furthermore, when the memory expansion conditions are met, the processor 111 requests extended memory 102 from the host 10's memory 101, up to a previously determined memory range. In some implementations, when the processor 111 meets the memory expansion conditions, the processor 111 can request at least a portion of the memory from the host 10's memory 101, up to a previously determined memory range, as extended memory 102. In other implementations, the processor 111 can request extended memory 102 from the host 10's memory 101 multiple times, up to a previously determined memory range, and the sum of the multiple requests for extended memory 102 does not exceed the previously determined memory range.

[0030] In this embodiment, the specific implementation of the memory expansion condition is not limited. For example, the memory expansion condition can be empty, i.e., no memory expansion condition is set. Alternatively, it can be when the host's memory utilization is less than or equal to a first threshold; where the host's memory utilization is less than or equal to the first threshold, it means the host has sufficient memory. Alternatively, it can be when the memory utilization of the acceleration device's existing memory is greater than or equal to a second threshold; where the memory utilization of the acceleration device's existing memory is greater than or equal to the second threshold, it means the acceleration device has insufficient memory.

[0031] For example, by monitoring the existing memory usage of the acceleration device, including but not limited to memory usage rate, and if the memory usage rate reaches a second threshold, it can be determined that the acceleration device is running out of memory. The second threshold can be, but is not limited to, 70%, 80%, or 90%, etc.

[0032] Furthermore, after the processor requests an extended memory range from the host's memory, it performs address mapping on the extended memory to map the address space of the memory extension device to the address space of the extended memory in the host, thereby enabling the processor to access the extended memory through the memory access channel.

[0033] In this embodiment of the disclosure, a memory expansion device is added to the acceleration device. The memory expansion device is interconnected with the processor and the host of the acceleration device to form a memory access channel. Then, the processor of the acceleration device uses the host's memory as its memory expansion. During the memory expansion process, the address of the memory expansion device is mapped to the expanded memory so that when the processor accesses the memory expansion device, it can access the expanded memory through the memory access channel. This realizes that the host's memory is used as the memory resource for memory expansion, which can not only reduce the memory implementation cost of the acceleration device, but also ensure the performance of the acceleration services provided by the acceleration device.

[0034] In this embodiment of the disclosure, the internal implementation structure of the memory expansion device 112 is not limited. Any implementation structure that can interconnect with the processor in the host and the acceleration device and establish a memory access channel between the host memory and the processor of the acceleration device, so as to enable memory expansion from the host memory to the processor of the acceleration device, is applicable to this embodiment of the disclosure.

[0035] In an optional embodiment, as shown in FIG2, the memory expansion device 112 includes a memory device 113 based on a memory expansion protocol and an interconnect device 114 based on a physical interconnect protocol. In some embodiments, the memory device and the interconnect device can be implemented based on programmable logic devices, such as including but not limited to FPGA (Field Programmable Gate Array), CPLD (Complex Programmable Logic Device), etc. The memory device 113 based on the memory expansion protocol is mainly used to provide memory expansion functionality, supporting memory expansion from the host 10's memory 101 for the processor 111. The interconnect device 114 based on the physical interconnect protocol is mainly used for interconnection with the host 10 to realize data transmission between the host 10 and the acceleration device 11. Because it is a physical interconnect, the data transmission latency is low and the security is high. The interconnection of the interconnect device 114 with the host 10 mainly refers to interconnection with the host's processor 103.

[0036] Based on the above, in one interconnection method between processor 111 and memory expansion device 112, processor 111 is interconnected with memory device 113 via a memory expansion protocol, memory device 113 is interconnected with interconnect device 114, and interconnect device 114 is interconnected with host 10 via a physical interconnect protocol, thereby forming a memory access channel between processor 111 and host 10. The memory access channel can be represented as a channel formed by processor 111 → memory device 113 → interconnect device 114 → host 10. The memory access channel is used at least for data transmission between host 10 and acceleration device 11, so as to use the memory 101 of host 10 as a memory resource for memory expansion of processor 111.

[0037] In this embodiment, the specific implementation of the memory extension protocol and the physical interconnect protocol is not limited. Preferably, a public interconnect protocol can be used to achieve compatibility across various dimensions such as the version and manufacturer of interconnected objects (e.g., processor, host). In an optional embodiment, the physical interconnect protocol can be PCIe (Peripheral Component Interconnect Express, a high-speed peripheral interconnect bus standard), and the memory extension protocol can be CXL (Compute Express Link, an open interconnect protocol). CXL is an extension of PCIe, and both CXL and PCIe are public interconnect protocols. By using these standardized protocols, i.e., public interconnect protocols, the compatibility between the processor and the host is better. Furthermore, in the implementation of the memory access channel, a seamless conversion from the CXL protocol to the PCIe protocol is achieved.

[0038] Furthermore, in this optional embodiment, the interconnect device can be implemented as a PCIe device. The PCIe device interconnects with the host's processor via the PCIe bus, enabling data transmission between the PCIe device and the host. This data transmission offers low latency and high security. The memory device can be implemented as a CXL Type 3 memory device, which is specifically designed for memory expansion within the CXL protocol. CXL includes a PCIe-based input / output (CXL.io) protocol and a (CXL.mem) protocol for supporting memory sharing and cache coherence. In one example, the memory device is implemented based on CXL.io and CXL.mem. For example, corresponding devices can be implemented for CXL.io and CXL.mem respectively, without limitation on the implementation form; it can be software implementation, hardware implementation, or a combination of both. In this case, the memory device includes the devices corresponding to CXL.io and CXL.mem respectively.

[0039] In the above optional embodiments, the memory extension protocol is CXL and the physical interconnect protocol is PCIe, but it is not limited to this.

[0040] In yet another example, the physical interconnect protocol could also be CCIX (Cache Coherent Interconnect for Accelerators, a high-performance interconnect protocol).

[0041] In this disclosure, the extended memory address mapping process may involve address space conversions such as those between the processor and memory devices, memory devices and interconnect devices, and interconnect devices and hosts. For ease of description and distinction, as shown in Figure 2, the address spaces are named in the following order: Processor 111 → Memory Device 113 → Interconnect Device 114 → Host 10. The address spaces corresponding to this order are named as follows: Address Space A1 (i.e., the second address space) → Address Space A2 (i.e., the first address space) → Address Space A3 → Address Space A4. The address segments selected from each address space are named as follows: Address Segment b1 (i.e., the fourth address segment) → Address Segment b2 (i.e., the first address segment) → Address Segment b3 (i.e., the third address segment) → Address Segment b4 (i.e., the second address segment). Correspondingly, the address information within each address segment is named as follows: Address Information c1 (i.e., the fourth address information) → Address Information c2 (i.e., the first address information) → Address Information c3 (i.e., the third address information) → Address Information c4 (i.e., the second address information). In subsequent embodiments, the above naming will be used to refer to different address spaces, address segments, and address information, and will not be repeated in subsequent embodiments.

[0042] During the process of expanding the processor's memory, on the one hand, the memory device is detected and reported to the processor the range of memory it can expand from the host's memory to the processor; on the other hand, upon receiving the reported memory range, the processor maps that memory range to the processor's address space.

[0043] Based on this, in an optional embodiment, when detected by the processor of the accelerated device, the memory device reports to the processor a memory range that can be expanded from the host memory for the processor. This memory range corresponds to a first address space (i.e., address space A2) on the memory device. Through the reporting by the memory device, the processor can understand the memory range in the host memory that can be used for memory expansion. For example, let N represent the memory range reported by the memory device. The size of N is not limited, such as 16G, 64G, etc. In this example, address space A2 is represented as the DPA (Device Physical Address) address range [0 to N-1]. The DPA address range represents a continuous address range from 0 to N-1. The prefix "DPA" before the DPA address range indicates that the DPA address range is the address space corresponding to the memory device.

[0044] Furthermore, when the processor detects a memory device, it receives the memory range reported by the memory device and allocates a second address space (i.e., address space A1) corresponding to that memory range in the processor's address space, thereby incorporating the memory range of the memory device into the processor's address space. The processor also fills the base address of address space A1 into the memory device, and the memory device maintains the correspondence between address space A2 and address space A1. Continuing with the example above, address space A1 refers to the physical address space of the processor in the acceleration device, which can be represented as the HPA (Host Physical Address) address space, for example, as the HPA address range [HPA_0~HPA_N-1], representing a continuous address range from HPA_0 to HPA_N-1. Compared to the DPA address range [0~N-1] in the example above, it can be concluded that address space A2 and address space A1 have a correspondence. As shown in the example, this correspondence can be a one-to-one correspondence between the address information in address space A2 and the address information in address space A1. In this example, where a correspondence exists, the address information between address space A2 and address space A1 is corresponded by setting an offset. In this example, the offset is "HPA_0". The address value HPA_0 + i minus a fixed offset "HPA_0" yields the starting address value "i" of DPA. It should be understood that this method of achieving address correspondence between address space A2 and address space A1 by setting an offset is merely an example and does not constitute a limitation on this embodiment.

[0045] Optionally, based on the processor's current memory usage, the conditions for memory expansion and the amount of memory to be expanded are determined, with this amount corresponding to the expanded memory. The processor's existing memory refers to the memory currently accessible to the processor. If the processor has previously been expanded from the host's memory, the existing memory includes both local memory and expanded memory. Local memory is shown as local memory 116 in Figure 2, and expanded memory is shown as expanded memory 102 in Figure 2. If it has not been expanded or the expanded memory has been released, the existing memory does not include expanded memory. Determining the conditions for memory expansion and the amount of memory to be expanded based on the processor's current memory usage allows for on-demand expansion of the processor's existing memory, enabling more flexible and efficient responses to changes in memory requirements that may arise from accelerated services.

[0046] This embodiment does not limit the method for determining the amount of memory to be expanded based on the processor's current memory usage. Two implementation methods are given below, but are not limited to these.

[0047] In one optional implementation, a second threshold is set for the memory utilization rate of the processor's existing memory. Further, the usage of existing memory, including its utilization rate, is monitored, and if the memory utilization rate reaches the second threshold, it is determined that the memory expansion condition is met. Then, for different values ​​of the second threshold, a corresponding amount of memory to be expanded is set. Optionally, a mapping table between the second threshold and the amount of memory to be expanded is set, thereby obtaining the memory amount corresponding to the second threshold by querying the mapping table. This table-based method of obtaining the amount of memory to be expanded is simple and efficient.

[0048] In another alternative implementation, the processor's current memory requirement is predicted using a memory demand prediction model, which can be trained based on the processor's historical memory usage on the acceleration device. Then, if the difference between the predicted current memory requirement and the remaining existing memory is a predetermined difference, it is determined that the memory expansion condition is met. The predetermined difference is the current memory requirement minus the remaining existing memory. The current memory requirement is at least greater than the processor's current memory. Furthermore, if the memory expansion condition is met, the predetermined difference is used as the amount of memory to be expanded. This model-based prediction method accurately predicts the current memory requirement, and based on this, the amount of memory to be expanded can be calculated relatively accurately.

[0049] Furthermore, in this optional embodiment, the first address segment (i.e., address segment b2) corresponding to the extended memory is selected from address space A2 according to the memory amount. Address space A2 is the address space corresponding to the previously determined memory range on the memory device. Continuing with the above example, when address space A2 is represented as the DPA address range [0~N-1], assuming an address segment of length m is selected from address space A2, the selected address segment is represented as the DPA address range [i~i+m-1]. This address segment can be used as an example of address segment b2, which is represented as an address segment from the starting address "i" to the ending address "i+m-1", with a length of m bytes.

[0050] Further, in this optional embodiment, a memory expansion request, including the amount of memory, is sent to the host. Optionally, one way to send the memory expansion request is through a control channel between the processor and the host, as shown in Figure 2. The control channel is established in advance by the processor interacting with the host based on custom channel establishment instructions. It is a logical channel between the host and the processor, and this logical channel can be established based on the memory expansion device being interconnected with both the processor and the host. In other words, the memory expansion device being interconnected with both the processor and the host can be used not only to form a memory access channel between the processor and the host, but also to form a control channel between them. In one example, taking PCIe as the physical interconnect protocol, the memory expansion device is interconnected with the host based on the PCIe bus, forming a PCIe channel. In some cases, multiple logical channels can be created on this PCIe channel, and these multiple logical channels reuse the PCIe channel. These multiple logical channels include a control channel and a memory access channel. In some cases, the control channel and the memory access channel may not reuse the PCIe channel; for example, the control channel can also communicate bidirectionally through a network interface, thereby enabling the sending of the memory expansion request to the host through the network interface.

[0051] In this optional embodiment, when the host receives a memory expansion request including a memory amount, it allocates a dedicated and contiguous memory space from the host's memory as expanded memory for the processor, based on the memory amount. The host's memory space is referred to as address space A4. The expanded memory has a second address segment (i.e., address segment b4) within address space A4 for address mapping, where the second address segment (i.e., address segment b4) corresponds to the first address segment (i.e., address segment b2). The expanded memory is dedicated to the acceleration device. Other hardware and software on the computer device besides the acceleration device, including but not limited to the host's processor, cannot access the expanded memory, i.e., cannot access address segment b4, thereby ensuring the isolation of the expanded memory from other memory on the host and reducing the probability of data loss or corruption in the expanded memory. Continuing with the above example, when the selected address segment b2 is represented as the DPA address range [i~i+m-1], address segment b4 can be represented as the CN_HPA address range [CN_HPA_i~CN_HPA_i+m-1] as an example of address segment b4. The prefix "CN_HPA" to the CN_HPA address range indicates that this CN_HPA address range belongs to the host's address space, that is, to the physical address of the host's memory. Address segment b4 is represented as the address segment from the starting address "CN_HPA_i" to the ending address "CN_HPA_i+m-1" of address segment b4, with a length of m bytes.

[0052] In some embodiments, during memory allocation based on the memory amount, the host first searches for contiguous blocks of that memory in its own memory. If no contiguous block of that memory exists in the host's memory, alternatively, a page replacement algorithm can be used for page scheduling. This algorithm determines existing pages that should be swapped out to persistent storage media (such as a disk) to make room for the new page. In this embodiment, the specific implementation of the page replacement algorithm is not limited. For example, page replacement algorithms include, but are not limited to, FIFO (First-In-First-Out) or LRU (Least Recently Used). The principle of the FIFO algorithm is to swap out the page that entered memory earliest, while the principle of the LRU algorithm is to swap out the least recently used page.

[0053] In another optional embodiment, a garbage collection (GC) mechanism is employed to reclaim storage space occupied by invalid data and free up memory space from the host's memory by moving valid data. This embodiment does not limit the specific garbage collection strategy. For example, it includes, but is not limited to, the mark-and-sweep algorithm and reference counting. The mark-and-sweep algorithm involves traversing all root objects and marking all reachable objects of the root objects as live objects, thereby reclaiming the memory space occupied by unmarked objects. Reference counting associates a reference counter with each object. When a new reference points to the object, the reference counter is incremented by 1; when a reference no longer points to the object, the reference counter is decremented by 1; if the reference counter drops to 0, it means the object is no longer used and its memory space can be reclaimed.

[0054] In this embodiment, a portion of the host's memory is freed up through methods such as page replacement or garbage collection, and this portion is used as contiguous memory.

[0055] If a contiguous block of memory of that amount is available, it can be marked as dedicated memory allocated to the processor, i.e., extended memory. Optionally, after allocation, the host returns a response message to the acceleration device via a control channel to confirm successful extended memory allocation to the processor.

[0056] Optionally, upon successful allocation of extended memory to the processor, a fourth address segment (address segment b1) corresponding to address segment b2 is allocated in address space A1, and address segment b1 is added to the memory management system of the acceleration device. Before adding address segment b1 to the memory management system, the memory management system is unaware of the extended memory's existence. After addition, the memory management system recognizes the extended memory's existence and can expose it to applications, allowing applications to access and use it. Address space A1 is the portion of the processor's address space corresponding to the memory range that the memory extension device can extend for the processor. Optionally, the allocation of address segment b1 and its addition to the memory management system can be performed through a hot-add operation to expose the extended memory to applications, but this is not a limitation. Following the example above, if address segment b2 can be represented as a DPA address range of [i~+m-1], then the corresponding address segment b1 can be represented as an HPA address range of [HPA_0+i~HPA_0+i+m-1]. Prepending "HPA" to the HPA address range indicates that this HPA address range belongs to the address space of the processor in the acceleration device. [HPA_0+i~HPA_0+i+m-1] can be used as an example of address segment b1, which is represented as an address range from the starting address "HPA_0+i" to the ending address "HPA_0+i+m-1", with a length of m bytes.

[0057] Furthermore, upon successful allocation of extended memory to the processor, the method also includes address mapping of the extended memory. Address mapping refers to the process of converting an address in one address space to an address in another. In this embodiment, address mapping primarily refers to the process of converting the extended memory's address space A2 in the memory device to its address space A4 in the host. This process involves the conversion from address segment b2 corresponding to address space A2 to address segment b4 corresponding to address space A4. Address segment b4 corresponds to address segment b2. The address mapping is unidirectional, i.e., it occurs from processor 111 → memory device 113 → interconnect device 114 → host 10. This may involve multiple address mappings, but the ultimate goal is to map the extended memory to convert address segment b1 in processor 111's address space A1 to its corresponding address segment b4 in the host's address space A4. During the address mapping process, address translation components on the host or acceleration device can be used. The address mapping process differs between different address translation components, which will be described below.

[0058] In an optional embodiment, when using the host's first address translation component 104, a first address translation page table is configured on the first address translation component 104 during the address mapping process. The first address translation page table records the address mapping relationship between a third address segment (i.e., address segment b3) and address segment b4, where address segment b3 is the address segment mapped from address segment b2 on the interconnect device. Optionally, a mapping relationship exists between address segment b3 and address segment b2, and based on this mapping relationship, the conversion from address segment b2 to address segment b3 can be completed. In this embodiment, the mapping relationship between address segment b3 and address segment b2 is not limited. For example, address information in address segment b2 can be superimposed with a certain amount of address offset to obtain address information in address segment b3. Further optionally, the address offset can be 0. In this case, address segment b3 and address segment b2 are the same. In this case, no address translation is required between the interconnect device and the memory device, saving various overheads associated with address translation and improving the efficiency of accessing extended memory. Address segment b3 can be represented as an IOVA address range [i~i+m-1]. The prefix "IOVA" indicates that this IOVA address range belongs to the address space of the interconnect device. This IOVA address range can serve as an example of address segment b3, which is represented as an address segment from the starting address "i" to the ending address "i+m-1", with a length of m bytes. Address segment b4 is represented as [CN_HPA_i~CN_HPA_i+m-1]. In this example, the first address translation page table records the mapping relationship from [i~i+m-1] to [CN_HPA_i~CN_HPA_i+m-1].

[0059] In another optional embodiment, when using the second address translation component 115 of the acceleration device, a second address translation page table is configured on the second address translation component 115 during the address mapping process. The second address translation page table records the address mapping relationship between address segment b2 (i.e., the first address segment) and address segment b4 (i.e., the second address segment). By using the second address translation component of the acceleration device, the dependence on the host's first address translation component can be reduced. Continuing with the above example, address segment b4 is represented as [CN_HPA_i~CN_HPA_i+m-1]. Assuming address segment b2 is the same as address segment b3, then address segment b2 is represented as [i~i+m-1]. The second address page table records the mapping relationship between [i~i+m-1] and [CN_HPA_i~CN_HPA_i+m-1]. Optionally, the first address translation component 104 and the second address translation component 115 can be implemented as an IOMMU (Input / Output Memory Management Unit).

[0060] Furthermore, after address mapping is completed, in response to an application's access operation to extended memory, the processor can perform access to the extended memory on the host based on the memory access channel. This application can run on an acceleration device; alternatively, it can run on the host itself, without limitation. Further optionally, during the access to extended memory, the processor sends a first memory access request to the memory device based on a memory extension protocol. The first memory access request is used to request access to the extended memory, and can be a read request or a write request. Optionally, the first memory access request includes fourth address information (i.e., address information c1) in address segment b1, where address segment b1 is the address segment in address space A1 corresponding to address segment b2, and address space A1 is an address space allocated from the processor's address space.

[0061] Furthermore, during the aforementioned access to extended memory, the memory device receives a first memory access request sent by the processor. Then, the memory device can translate the address information c1 in the first memory access request. The memory device converts the address information c1 in address segment b1 into the first address information (i.e., address information c2) in address segment b2. Based on address information c2, the memory device obtains address information c3 or address information c4, generates a second memory access request based on address information c3 or address information c4, and then sends the second memory access request to the host through the interconnect device. The second memory access request includes address information c3 or c4. For example, the address information c1 in the first memory access request can be replaced with address information c3 or c4 to obtain the second memory access request. The memory device obtains address information c3 or c4 based on address information c2 either by using a second address translation component on the acceleration device, or by using a first address translation component to translate address information c2 in address segment b2 into the second address information (i.e., address information c4) in address segment b4; or the memory device directly translates address information c2 into the third address information (i.e., address information c3) in address segment b3. Whether address information c2 is translated into address information c4 or address information c3 depends on whether the address translation component is on the host or on the acceleration device. These scenarios will be described below.

[0062] If a second address translation component (such as the IOMMU on the acceleration device) is used, the memory device will query the second address translation page table through the second address translation component to translate address information c2 into address information c4 in address segment b4. The second address translation page table records the address mapping relationship between address segment c2 and address segment c4; address segment b4 is the address segment of the extended memory in the host's address space, and address information c4 is used by the host to access the extended memory.

[0063] If a first address translation component (such as an IOMMU on the host) is used, then on the acceleration device side, the memory device will translate address information c2 into address information c3 within address segment b3, where address segment b3 is the address segment mapped from address segment b2 on the interconnect device. There is a mapping relationship between address information c3 and address information c4, and the first address translation component can complete the conversion from address information c3 to address information c4 based on this mapping relationship. In this embodiment, the mapping relationship between address information c3 and address information c4 is not limited; for example, address information c4 can be obtained by adding a certain address offset to address information c3. Further optionally, the address offset can be 0. In this case, address information c3 and address information c4 are the same, resulting in higher address translation efficiency.

[0064] Then, the interconnect device sends a second memory access request carrying address information c4 or address information c3 to the host via the physical interconnect protocol.

[0065] Upon receiving a second memory access request, the host accesses the extended memory. Optionally, upon receiving a second memory access request including address information c3, the host queries the first address translation page table through the first address translation component to translate the address information c3 in address segment b3 to the address information c4 in address segment b4; wherein, the first address translation page table records the address mapping relationship between address segment b3 and address segment b4; then, the host accesses the extended memory according to the address information c4. Alternatively, if the host receives a second memory access request including address information c4, address translation is not required, and the host directly accesses the extended memory according to the address information c4.

[0066] Following the example above, an example of the processor's access process to extended memory in an accelerator device is given, and the access processes of the first address translation component on the host and the second address translation component on the accelerator device are described separately.

[0067] If the first address translation component on the host is used, and the first memory access request is a write request, assuming the address information c1 included in the first memory access request is represented as "HPA_0+i", where HPA_0+i represents the starting address of address segment b1 and a dataset [S1data1, S1data2, ..., S1dataL] of length length bytes, the memory device first translates the address information c1 into address information c2, such as translating address information c1 "HPA_0+i" into address information c2 "i". Further, the memory device or the second address translation component translates address information c2 into address information c3. In this example, address information c2 and address information c3 are the same, so no translation is needed between address information c2 and address information c3, but this is not a limitation. In this process, when the second address translation component converts address information c2 to address information c3, address information c3 can be returned to the memory device. Upon receiving address information c3, the memory device converts the first memory access request into a second memory access request, that is, replacing address information c1 in the first memory access request with address information c3 to obtain the second memory access request. Then, the interconnect device sends the second memory access request to the host via the physical interconnect protocol. The second memory access request includes address information c3, such as "i"; and a dataset [S1data1, S1data2, ..., S1dataL] of length length bytes. Upon receiving the second memory access request, the host queries the first address translation page table through the first address translation component to convert address information c3 into address information c4, such as converting address information c3 "i" into address information c4 "CN_HPA_i". Then, the host writes the dataset [S1data1, S1data2, ..., S1dataL] into an address space starting at address CN_HPA_i and of length length bytes.

[0068] If the first address translation component on the host is used, and the first memory access request is a read request, assuming the address information c1 included in the first memory access request is "HPA_0+i" for example, and the length of the read request is length bytes, firstly, the memory device will convert the address information c1 to address information c2, such as converting address information c1 "HPA_0+i" to address information c2 "i". Further, the memory device or the second address translation component will convert address information c2 to address information c3. In this example, address information c2 and address information c3 are the same, so no conversion is needed between address information c2 and address information c3, but this is not a limitation. Where the second address translation component converts address information c2 to address information c3, address information c3 can be returned to the memory device. After receiving address information c3, the memory device will convert the first memory access request to a second memory access request, that is, replace the address information c1 in the first memory access request with address information c3 to obtain the second memory access request. Then, the interconnect device sends a second memory access request to the host via the physical interconnect protocol; the second memory access request includes address information c3, such as "i". The first address translation component queries the first address translation page table to convert address information c3 to address information c4, such as converting address information c3 "i" to address information c4 "CN_HPA_i". The host reads a dataset [S2data1, S2data2, ..., S2dataL] of length from the starting address "CN_HPA_i" and returns it.

[0069] If the second address translation component on the acceleration device is used, assuming the first memory access request is a write request, and the address information c1 included in the first memory access request is "HPA_0+i" for example, where HPA_0+i is the starting address of address segment b1; and the first memory access request also includes a dataset [S1data1, S1data2, ..., S1dataL] of length length bytes. First, the memory device will convert the address information c1 into address information c2, such as converting address information c1 "HPA_0+i" into address information c2 "i". Further, the second address translation component queries the second address translation page table to convert address information c2 into address information c4 in address segment b4, such as converting address information c2 "i" into address information c4 "CN_HPA_i". In this process, the second address translation component converts address information c2 to address information c4 and returns c4 to the memory device. Upon receiving c4, the memory device converts the first memory access request into a second memory access request, replacing address information c1 in the first request with c4 to obtain the second memory access request. Then, the interconnect device sends the second memory access request to the host via the physical interconnect protocol. The second memory access request includes address information c4, such as "CN_HPA_i", and a dataset [S1data1, S1data2, ..., S1dataL] of length length bytes. Upon receiving the second memory access request, the host writes the dataset [S1data1, S1data2, ..., S1dataL] into an address space starting at address CN_HPA_i and of length length bytes.

[0070] If the second address translation component on the acceleration device is used, and the first memory access request is a read request, assuming the address information c1 included in the first memory access request is "HPA_0+i" for example, and the length of the read request is length bytes, the memory device first converts the address information c1 to address information c2, such as converting address information c1 "HPA_0+i" to address information c2 "i". Further, the second address translation component queries the second address translation page table to convert address information c2 to address information c4 in address segment b4, such as converting address information c2 "i" to address information c4 "CN_HPA_i". After converting address information c2 to address information c4, the second address translation component can return address information c4 to the memory device. Upon receiving address information c4, the memory device converts the first memory access request to a second memory access request, that is, replacing address information c1 in the first memory access request with address information c4 to obtain the second memory access request. Subsequently, the interconnect device sends a second memory access request to the host via the physical interconnect protocol; the second memory access request includes address information c4, such as "CN_HPA_i". Upon receiving the second memory access request, the host reads a dataset [S2data1, S2data2, ..., S2dataL] of length from the starting address "CN_HPA_i" and returns it.

[0071] In one optional embodiment, a background daemon runs in the processor's on-memory management system of the acceleration device to monitor the popularity information of data stored in memory. This memory includes extended memory on the host for the processor and / or the processor's local memory. During the monitoring of the popularity information of data in memory, for example, the access frequency of the data in memory can be used for monitoring. Data that is frequently accessed within a certain period can be assigned a higher popularity rating; data with higher popularity ratings can be called "hot data." Conversely, memory accessed less frequently within a certain period can be assigned a lower popularity rating; data with lower popularity ratings can be called "cold data."

[0072] Based on this, in one optional embodiment, the popularity information of data stored in extended memory is monitored, and first data with popularity information greater than a first popularity threshold in extended memory is migrated to the processor's local memory. This first data is data with high popularity information, also known as hot data. And / or, the popularity information of data stored in the processor's local memory is monitored, and second data with popularity information less than a second popularity threshold in local memory is migrated to extended memory. This second data is data with low popularity information, also known as cold data. The first popularity threshold is greater than the second popularity threshold.

[0073] The above embodiments provide a method for expanding processor memory, which has diverse application scenarios and is applicable to a variety of situations. The following application scenarios will provide a detailed description of the technical solutions provided by the embodiments of this disclosure.

[0074] The application scenarios can be determined based on the actual environment and usage requirements. In this embodiment, the technical solution of this disclosure is illustrated using a scenario where a server expands the memory of a smart network card as an example. However, this does not mean that the method provided by this disclosure is only applicable to this application scenario. In these scenario embodiments, the specific implementation of some components or devices will also be limited. It should be understood that the examples in these scenario embodiments do not constitute a limitation on the embodiments of this disclosure. In this scenario embodiment, as shown in FIG3, a computer device is provided, including: a server and a smart network card; the smart network card includes a programmable logic device (an example of a memory expansion device) and a CPU; wherein, the programmable logic device can be implemented as an ASIC (Application Specific Integrated Circuit) or an FPGA.

[0075] Based on the above scenario, this embodiment includes a preparation phase, a memory expansion phase, a memory access phase, and an optimization management phase. These phases will be described below. Preparation Phase

[0076] Step 1: Implement a PCIe device (an example of an interconnect device) and "CXL.mem / io" (an example of a memory device) based on the programming logic device. "CXL.mem / io" refers to the hardware and software co-implementation of CXL.io and CXL.mem as corresponding devices. For details regarding "CXL.mem" and CXL.io," please refer to the above embodiments; they will not be repeated here. For ease of description, CXL.mem / io will be referred to as "memory device" in subsequent scenario embodiments. The PCIe device interconnects with the server based on the PCIe protocol; the CPU of the smart network card interconnects with the memory device via the CXL protocol; the memory device interconnects with the server via PCIe, forming a memory access channel between the CPU of the smart network card and the server.

[0077] Step Two: The memory device presents the CXL Type 3 memory device to the CPU of the smart network card. For details regarding the CXL Type 3 memory device, please refer to the above embodiment. The memory device reports to the CPU of the smart network card the memory range that can be expanded from the server's memory for the CPU. This memory range, as an example of address space A2 on the CXL Type 3 memory device, can be represented as the DPA address range [0~N-1], where the DPA address range represents a continuous address range from 0 to N-1.

[0078] Step 3: The CPU of the smart network interface card (NIC) receives the memory range reported by the memory device and allocates an address space corresponding to that memory range in the CPU's address space. This allocated address space can be used as an example of address space A1. The CPU of the smart network interface card enumerates the CXL Type 3 memory device through CXL.io. The address space A1 allocated to this CXL Type 3 memory device in the CPU's address space corresponding to the memory range can be represented as the HPA address range [HPA_0~HPA_N-1]. This HPA address range represents a continuous address range from HPA_0 to HPA_N-1, corresponding one-to-one with address space A2 [0~N-1]. By allocating address space A1 to this memory range, it is incorporated into the processor's address space. However, since there is no corresponding physical memory on the server at this time, it is not added to the memory management system.

[0079] This concludes the preparation phase. The next phase is the memory expansion phase, which will be described step-by-step below.

[0080] Step 4: If the memory expansion conditions are met, determine the amount of memory that needs to be expanded by m bytes from the server's memory. This amount of m bytes shall not exceed the expansion range reported by the memory device.

[0081] Step 5: The CPU of the smart network card selects an address segment of length m bytes from the address space of the CXL Type 3 memory device (one example of address space A2). This address segment serves as one example of address segment b2. Further, the selected address segment b2 is represented as the DPA address range [i~i+m-1], that is, the address segment from the starting address "i" to the ending address "i+m-1". This address segment, as an example of address segment b2, has a length of m bytes. The corresponding processor address segment b1 can be represented as [HPA_0+i~[HPA_0+i+m-1], that is, the address segment from the starting address "HPA_0+i" to the ending address "HPA_0+i+m-1", with a length of m bytes.

[0082] Step Six: The CPU of the smart network card sends a memory expansion request to the server through the control channel between the two servers, in order to actually request a continuous amount of m bytes of memory (e.g., granularity of 128M) from the server.

[0083] Step 7: The server receives the memory expansion request and allocates memory according to the requested amount. First, it searches for a contiguous m-byte block of memory in the server's memory. If there is no contiguous m-byte block of memory in the server's memory, it attempts to free up the contiguous m-byte block of memory through page swapping or garbage collection, and uses this m-byte block of memory as memory allocated to the processor. Assume that the address segment of this contiguous m-byte block of memory in the server's address space is represented as the CN_HPA address range [CN_HPA_i~CN_HPA_i+m-1]. This address segment can be used as an example of address segment b4. Address segment b4 is represented as an address segment from the starting address "CN_HPA_i" to the ending address "CN_HPA_i+m-1", with a length of m. Furthermore, a customized memory management service on the server can be located in the hypervisor. Through this memory management service, other hardware and software, including the server's CPU, can be configured to prevent access to the CN_HPA address range [CN_HPA_i~CN_HPA_i+m-1], so that the CN_HPA address range can be used as extended memory for the smart network card.

[0084] Step 8: Configure the server's first address translation component or a second address translation component based on the programmable logic device.

[0085] a. If using the server's IOMMU (an example of the first address translation component), configure the first address translation page table on the server's IOMMU to set the mapping relationship between the IOVA address range [i~i+m-1] and the CN_HPA address range [CN_HPA_i~CN_HPA_i+m-1]. The IOVA address range can be used as an example of address segment b3, which is represented as an address segment from the starting address "i" to the ending address "i+m-1", with a length of m bytes.

[0086] In this scenario embodiment, the example is that address segment b2 and address segment b3 have the same address. That is, the same address space is used on the CXL Type3 memory device and the PCIe device. No address translation or mapping processing is required between the two devices, which can greatly improve the access efficiency of extended memory.

[0087] b. If using the IOMMU (an example of a second address translation component) of a smart network interface card, configure the second address translation page table on the IOMMU of the smart network interface card to set the mapping relationship between address segment b2[i~i+m-1] and address segment b4[CN_HPA_i~CN_HPA_i+m-1].

[0088] Step 9: The server returns a response message to the processor of the smart network card through the control channel defined with the smart network card to notify the CPU of the smart network card that the request for contiguous memory m bytes has been successfully completed.

[0089] Step 10: The CPU of the smart network card performs a hot-addition operation on the memory. It allocates address segment b1 corresponding to address segment b2 in the CPU's address space, then allocates a management structure (such as a page structure), adds address segment b1 to the memory management system, and sets its status to idle and available. In this example, address segment b2 can be represented as the DPA address range [i+m-1], and the corresponding address segment b1 can be represented as the HPA address range [HPA_0+i~HPA_0+i+m-1]. [HPA_0+iHPA_0+i+m-1] can be used as an example of address segment b1, representing an address range from the starting address "HPA_0+i" to the ending address "HPA_0+i+m-1", with a length of m bytes.

[0090] At this point, the memory expansion process is complete, and the expanded HPA address space [HPA_0+i, HPA_0+i+m-1] for the processor can now be accessed. The memory access process is described step by step below.

[0091] Step 11: Describe the memory access process, which includes the following sub-steps: ad.

[0092] a. The CPU of the smart network card sends a first memory access request to the CXL Type3 memory device simulated by the programmable logic device via the CXL.mem protocol. It can be a read request or a write request. The first memory access request includes address information c1. Assume that the address information c1 includes the starting HPA address as HPA_0+i.

[0093] After receiving the first memory access request, the b.CXL Type3 memory device extracts information such as the starting address HPA_0+i, the read / write direction of the message, and the length in bytes. If the first memory access request is a write request, it also includes the dataset, denoted as [S1data1, S1data2, ..., S1dataL].

[0094] c. Based on the programming logic device, subtract a fixed offset HPA_0 from the address value HPA_0+i to obtain the starting address value "i" of DPA.

[0095] d. The access process will vary depending on whether a server or an IOMMU based on a programmable logic device is used.

[0096] If the server's IOMMU is used: Based on the programming logic device, a second memory access request is sent to the server via the PCIe protocol through the PCIe device. This second memory access request includes address information c3, such as "i"; and a dataset [S1data1, S1data2, ..., S1dataL] of length `length` bytes. Upon receiving the second memory access request, the server queries the first address translation page table through the IOMMU to convert the address information c3 "i" to the address information c4 "CN_HPA_i". Then, if it is a write request, the server will write the dataset [S1data1, S1data2, ..., S1dataL] to the address space starting at address CN_HPA_i and with a length of `length` bytes. If it is a read request, the server, through the IOMMU, will return the dataset [S2data1, S2data2, ..., S2dataL] starting at address CN_HPA_i and with a length of `length` bytes.

[0097] If the IOMMU on the smart network card is used, the programming logic device maps the address information c2 "i" to the address information c4CN_HPA_i through the second address page table on its IOMMU. Then, a second memory access request is sent to the server via the PCIe protocol through the PCIe device. The second memory access request includes the address information c4 "CN_HPA_i" and a dataset [S1data1, S1data2, ..., S1dataL] of length length bytes. If it is a write request, upon receiving the second memory access request, the server writes the dataset [S1data1, S1data2, ..., S1dataL] to the address space starting at CN_HPA_i with a length of length bytes. If it is a read request, the server returns the dataset [S2data1, S2data2, ..., S2dataL] starting at CN_HPA_i with a length of length bytes.

[0098] The example above illustrates the process of accessing extended memory. The following section describes how to optimize the management of extended memory.

[0099] Step 12: A background daemon runs in the CPU memory management system of the smart network interface card (NIC) to monitor memory usage. Alternatively, the application layer can specify whether data should be migrated to local memory or extended memory. Memory usage is monitored periodically. This includes extended memory on the server for the NIC's CPU and / or the NIC's local memory. The memory usage frequency is determined by scanning the access bits of the process page table, and the frequency determines the corresponding usage frequency. After multiple scans, an LRU (Least Recently Used) linked list is created, with the head of the list containing hot data and the tail containing cold data. Hot data in extended memory can then be migrated to the NIC's local memory, and cold data in the NIC's local memory can be migrated to extended memory. The swap unit between extended memory and local memory is not limited; it can be, for example, 4KB or 16KB, and can be flexibly configured.

[0100] In this scenario embodiment, the process of releasing extended memory may also be included. The release process is as follows.

[0101] Step 13: When the smart network card finishes using m bytes of memory, it can choose to release that m bytes of memory, including: a. The CPU of the smart network card will perform a hot memory removal operation. First, it can check whether the status of the m bytes of memory is free and available. If it is free and available, it will remove the address segment c1 corresponding to the m bytes of memory from the memory management system, then release the related management structures (such as the page structure), and finally release the previously allocated address space (such as address space A1); b. Using the control channel between the CPU of the smart network card and the server, it can notify the server to release the address segment (such as address segment c4) corresponding to the m bytes of memory in the server's address space. Different address reporting methods can be selected; for example: if the server's IOMMU is used, the address segment c3 is notified, which can be represented as [i ~ i + m - 1]; if the smart network card's IOMMU is used, the address segment c4 is notified, which can be represented as [CN_HPA_i ~ CN_HPA_i + m - 1]; c. In response to the notification from the smart network interface card (NIC), the server determines the corresponding address range c4, which can be represented as [CN_HPA_i~CN_HPA_i+m-1]. For example, if the notification is for address c3, the server needs to translate address range c3 into address range c4; d. The server's memory management service configures other hardware and software besides the smart NIC, including the server's CPU, to access address range c4 as needed, which can be represented as [CN_HPA_i~CN_HPA_i+m-1].

[0102] At this point, the process of releasing the extended memory is complete.

[0103] This disclosure also provides a memory expansion method applied to an acceleration device. The processor and memory expansion device in the acceleration device cooperate to form a memory access channel between the processor and the host, so as to use the host's memory as a memory resource for memory expansion of the processor. As shown in FIG4, the method includes: S401: determining the range of memory that the memory expansion device can expand for the processor from the host's memory; S402: if the processor meets the memory expansion conditions, requesting expanded memory from the host's memory not exceeding the memory range; S403: performing address mapping on the expanded memory to access the expanded memory through the memory access channel.

[0104] In one optional embodiment, the memory expansion device includes: a memory device based on a memory expansion protocol and an interconnect device based on a physical interconnect protocol; determining the memory range that the memory expansion device can expand for the processor from the host's memory includes: upon detecting the memory device, receiving a memory range reported by the memory device that it can expand for the processor from the host's memory, the memory range corresponding to a first address space on the memory device; and allocating a second address space corresponding to the memory range in the processor's address space.

[0105] In an optional embodiment, the memory expansion device includes: a memory device based on a memory expansion protocol and an interconnect device based on a physical interconnect protocol; when the processor meets the memory expansion conditions, requesting extended memory from the host's memory up to the specified memory range includes: determining, based on the current memory usage of the processor, whether the processor meets the memory expansion conditions and the amount of memory to be expanded, wherein the amount of memory corresponds to the extended memory; selecting, based on the amount of memory, a first address segment corresponding to the extended memory from a first address space, wherein the first address space is the address space corresponding to the memory range on the memory device; sending a memory expansion request including the amount of memory to the host, so that the host can allocate dedicated and contiguous memory space for the processor as the extended memory based on the amount of memory; the extended memory has a second address segment in the host's address space for address mapping of the extended memory, wherein the second address segment corresponds to the first address segment.

[0106] In one optional embodiment, sending a memory expansion request, including the amount of memory, to the host so that the host can allocate a dedicated and contiguous memory space for the processor from its memory as the expanded memory, includes: sending the memory expansion request to the host via a control channel between the host and the host so that the host can allocate a dedicated and contiguous memory space for the processor from its memory as the expanded memory; wherein the control channel is pre-established by the processor with the host based on a custom channel establishment instruction.

[0107] In one optional embodiment, address mapping of the extended memory to access it via the memory access channel includes: configuring a first address translation page table on a first address translation component of the host to access the extended memory via the memory access channel; the first address translation page table records the address mapping relationship between a second address segment and a third address segment, wherein the third address segment is the address segment mapped by the first address segment on the interconnect device; or, configuring a second address translation page table on a second address translation component of the acceleration device to access the extended memory via the memory access channel; the second address translation page table records the address mapping relationship between the second address segment and the first address segment.

[0108] In an optional embodiment, the method further includes: if the extended memory is successfully allocated to the processor, allocating a fourth address segment corresponding to the first address segment in a second address space; the second address space is a portion of the address space corresponding to the memory range in the processor's address space; and adding the fourth address segment to the memory management system of the acceleration device to expose the extended memory to the application.

[0109] In an optional embodiment, the method further includes: monitoring the popularity information of data stored in the extended memory, and migrating first data in the extended memory with popularity information greater than a first popularity threshold to the local memory of the processor; and / or, monitoring the popularity information of data stored in the local memory of the processor, and migrating second data in the local memory with popularity information less than a second popularity threshold to the extended memory; wherein the first popularity threshold is greater than the second popularity threshold.

[0110] This disclosure also provides a memory access method applied to an acceleration device. The processor and memory expansion device in the acceleration device cooperate to form a memory access channel between the processor and the host, so as to use the host's memory as a memory resource for memory expansion of the processor. As shown in FIG5, the method includes: S501: pre-allocating extended memory for the processor from the host's memory; S502: accessing the extended memory on the host based on the memory access channel.

[0111] In this embodiment, the implementation method of pre-allocating extended memory for the processor from the host's memory is not limited. For example, it can be the process described in the above embodiments. Alternatively, other methods may be used.

[0112] In one optional embodiment, the memory expansion device includes: a memory device based on a memory expansion protocol and an interconnect device based on a physical interconnect protocol; accessing the expanded memory on the host based on the memory access channel includes: sending a first memory access request based on the memory expansion protocol, the first memory access request being used to request access to the expanded memory; converting the first memory access request into a second memory access request based on the physical interconnect protocol and sending it to the host, so that the host can access the expanded memory.

[0113] In an optional embodiment, the first memory access request includes fourth address information in a fourth address segment, wherein the fourth address segment is an address segment in a second address space corresponding to the first address segment, the second address space is a portion of the address space corresponding to the memory range in the processor's address space, the first address segment is the address segment corresponding to the extended memory in the first address space, and the first address space is the address space corresponding to the memory range on the memory device; converting the first memory access request into a second memory access request based on the physical interconnect protocol and sending it to the host for the host to access the extended memory includes: converting the fourth address information into first address information in the first address segment; obtaining the second address information based on the first address information. Address information or third address information; generate the second memory access request based on the second address information or third address information, wherein the second address information or third address information is address information in the second address segment or third address segment obtained by converting the first address information; wherein the second address segment is the address segment that the extended memory has in the address space of the host, and the third address segment is the address segment that the first address segment is mapped on the interconnect device; send the second memory access request to the host via the interconnect device, wherein the second memory access request includes the second address information or the third address information, wherein the third address information is used for the host to convert to obtain the second address information, and the second address information is used for the host to access the extended memory.

[0114] The detailed implementation methods and beneficial effects of each step in this embodiment have been described in detail in the foregoing embodiments, and will not be elaborated here.

[0115] It should be noted that some processes described in the above embodiments and accompanying drawings include multiple operations that appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order they appear in this document, or they may be executed in parallel. The operation numbers, such as S401, S402, etc., are merely used to distinguish different operations and do not represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should also be noted that the descriptions such as "first" and "second" in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.

[0116] Figure 6 is a schematic diagram of an acceleration device provided in another exemplary embodiment of this disclosure. As shown in Figure 6, the acceleration device includes: a processor 65, a memory expansion device 63, and local memory 64; the local memory 64 is used to store computer programs and can be configured to store various other data to support operation on the acceleration device. Examples of this data include instructions, data structures, images, videos, etc., for any application or method operating on the acceleration device.

[0117] Processor 65, coupled to local memory 64, is configured to execute a computer program in local memory 64 for: determining a memory range that the memory expansion device can expand for the processor from the host memory; requesting extended memory from the host memory up to the memory range if the processor satisfies the memory expansion conditions; and performing address mapping on the extended memory to access the extended memory through the memory access channel.

[0118] In an optional embodiment, the memory expansion device includes: a memory device based on a memory expansion protocol and an interconnect device based on a physical interconnect protocol; when the processor 65 determines the memory range that the memory expansion device can expand for the processor from the host's memory, it is specifically configured to: upon detecting the memory device, receive a memory range reported by the memory device that it can expand for the processor from the host's memory, the memory range corresponding to a first address space on the memory device; and allocate a second address space corresponding to the memory range in the processor's address space.

[0119] In an optional embodiment, the memory expansion device includes: a memory device based on a memory expansion protocol and an interconnect device based on a physical interconnect protocol; when the processor 65 meets the memory expansion conditions and requests expanded memory from the host's memory not exceeding the memory range, the specific steps are as follows: based on the current memory usage of the processor, determine that the processor 65 meets the memory expansion conditions and the amount of memory to be expanded, the amount of memory corresponding to the expanded memory; select a first address segment corresponding to the expanded memory from a first address space based on the amount of memory, the first address space being the address space corresponding to the memory range on the memory device; send a memory expansion request including the amount of memory to the host, so that the host can allocate dedicated and contiguous memory space for the processor as the expanded memory based on the amount of memory; the expanded memory has a second address segment in the host's address space for address mapping of the expanded memory, the second address segment corresponding to the first address segment.

[0120] In an optional embodiment, the processor 65 sends a memory expansion request, including the amount of memory, to the host, so that the host can allocate a dedicated and contiguous memory space for the processor from its memory as the expanded memory based on the amount of memory. Specifically, the processor 65 sends the memory expansion request to the host via a control channel with the host, so that the host can allocate a dedicated and contiguous memory space for the processor from its memory as the expanded memory based on the amount of memory. The control channel is pre-established by the processor with the host based on a custom channel establishment instruction.

[0121] In an optional embodiment, when the processor 65 performs address mapping on the extended memory to access the extended memory through the memory access channel, it specifically performs the following: configuring a first address translation page table on the first address translation component of the host to access the extended memory through the memory access channel; the first address translation page table records the address mapping relationship between the second address segment and the third address segment, wherein the third address segment is the address segment mapped by the first address segment on the interconnect device; or, configuring a second address translation page table on the second address translation component of the acceleration device to access the extended memory through the memory access channel; the second address translation page table records the address mapping relationship between the second address segment and the first address segment.

[0122] In an optional embodiment, the processor 65 is further configured to: allocate a fourth address segment corresponding to the first address segment in a second address space when the extended memory is successfully allocated to the processor 65; the second address space is a portion of the address space corresponding to the memory range in the processor's address space; and add the fourth address segment to the memory management system of the acceleration device to expose the extended memory to the application.

[0123] In an optional embodiment, the processor 65 is further configured to: monitor the popularity information of data stored in the extended memory, and migrate first data in the extended memory whose popularity information is greater than a first popularity threshold to the processor's local memory; and / or, monitor the popularity information of data stored in the processor's local memory, and migrate second data in the local memory whose popularity information is less than a second popularity threshold to the extended memory; wherein the first popularity threshold is greater than the second popularity threshold.

[0124] Furthermore, as shown in Figure 6, the acceleration device also includes other components such as a communication component 66, an audio component 67, and a power supply component 68. Figure 6 only schematically shows some components and does not imply that the acceleration device only includes the components shown in Figure 6. Additionally, the components within the dashed boxes in Figure 6 are optional, not mandatory, and their specific implementation depends on the product form of the acceleration device. The acceleration device can be implemented as various chips, such as ASICs, CPLDs, and FPGAs; or it can be implemented as various boards, such as GPUs, smart network cards, and CIPUs.

[0125] This disclosure also provides an acceleration device, the implementation structure of which is the same as or similar to that of the acceleration device shown in FIG. 6, and can be implemented with reference to the structure of the acceleration device shown in FIG. 6. In the acceleration device provided in this embodiment, the memory expansion device pre-allocates extended memory from the host's memory for the processor; the processor accesses the extended memory on the host based on the memory access channel.

[0126] In an optional embodiment, the memory expansion device includes: a memory device based on a memory expansion protocol and an interconnect device based on a physical interconnect protocol; the processor is specifically configured to: send a first memory access request to the memory device based on the memory expansion protocol, the first memory access request being used to request access to the expanded memory; the memory device is specifically configured to: convert the first memory access request into a second memory access request based on the physical interconnect protocol and send it to the host, so that the host can access the expanded memory.

[0127] In an optional embodiment, the first memory access request includes fourth address information in a fourth address segment, wherein the fourth address segment is an address segment in a second address space corresponding to the first address segment, the second address space is a portion of the address space corresponding to the memory range in the processor's address space, the first address segment is the address segment corresponding to the extended memory in the first address space, and the first address space is the address space corresponding to the memory range on the memory device; when the memory device converts the first memory access request into a second memory access request based on the physical interconnect protocol and sends it to the host, it is specifically used to: convert the fourth address information into first address information in the first address segment; obtain the second address information based on the first address information or The third address information; the second memory access request is generated based on the second address information or the third address information, wherein the second address information or the third address information is address information in the second address segment or the third address segment obtained by converting the first address information; wherein the second address segment is the address segment that the extended memory has in the address space of the host, and the third address segment is the address segment that the first address segment is mapped on the interconnect device; the second memory access request is sent to the host via the interconnect device, the second memory access request including the second address information or the third address information, wherein the third address information is used for the host to convert to obtain the second address information, and the second address information is used for the host to access the extended memory.

[0128] The aforementioned local memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0129] The aforementioned communication components are configured to facilitate wired or wireless communication between the device containing the communication components and other devices. The device containing the communication components can access wireless networks based on communication standards, such as WiFi, 2G, 3G, 4G / LTE, 5G, or combinations thereof. In one exemplary embodiment, the communication components receive broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, the communication components also include a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on Radio Frequency Identification (RFID), Infrared Data Association (IrDA), Ultra Wide Band (UWB), Bluetooth (BT), and other technologies.

[0130] The aforementioned display includes a screen, which may include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a Touch Panel, the screen can be implemented as a touchscreen to receive input signals from the user. The Touch Panel includes one or more touch sensors to sense touches, swipes, and gestures on the Touch Panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation.

[0131] The aforementioned power supply components provide power to various components within the device in which they reside. These power supply components may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device in which they reside.

[0132] The aforementioned audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) configured to receive external audio signals when the device containing the audio component is in an operating mode, such as call mode, recording mode, or voice recognition mode. The received audio signals can be further stored in local memory or transmitted via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.

[0133] Accordingly, this disclosure also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement the steps in the above-described method embodiments. The computer-readable storage medium includes volatile or non-volatile or a combination thereof, and may be removable or non-removable. Examples of computer-readable storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random-access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technologies, CD-ROM, Digital Video Disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium.

[0134] Accordingly, this disclosure also provides a computer program product, which includes a computer program or instructions that, when executed by a processor, cause the processor to implement the steps in the above-described method embodiments. It should be understood that each step or combination of steps in the above-described method flow can be implemented by the computer program or instructions. Furthermore, these computer programs or instructions can be applied to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device, enabling the processor of the general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to function as an apparatus for implementing the corresponding functions in the above-described method embodiments.

[0135] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0136] The above are merely embodiments of this disclosure and are not intended to limit the scope of this disclosure. Various modifications and variations can be made to this disclosure by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of the claims of this disclosure.

Claims

1. A computer device, comprising: Host and acceleration devices; The acceleration device includes: a processor and a memory expansion device; The memory expansion device is interconnected with the processor and the host respectively, and is used to form a memory access channel between the processor and the host, so as to use the host's memory as a memory resource for memory expansion of the processor; The processor is configured to determine the range of memory that the memory expansion device can expand for the processor from the host's memory, and, if the processor meets the memory expansion conditions, request expanded memory from the host's memory up to the range of memory; and perform address mapping on the expanded memory to access the expanded memory through the memory access channel.

2. The device according to claim 1, wherein, The memory expansion device includes: a memory device based on a memory expansion protocol and an interconnect device based on a physical interconnect protocol; The processor is interconnected with the memory device; the memory device is interconnected with the host through the interconnection device, and is used to form the memory access channel between the processor and the host, so as to use the host's memory as a memory resource for memory expansion of the processor.

3. The device according to claim 2, wherein, The memory device is configured to, when detected by the processor, report to the processor a memory range that can be expanded for the processor from the host memory, the memory range corresponding to a first address space on the memory device; The processor is configured to: upon detecting the memory device, receive a memory range reported by the memory device, and allocate a second address space corresponding to the memory range in the processor's address space.

4. The device according to claim 2, wherein, The processor is used for: Based on the current memory usage of the processor, determine whether the processor meets the memory expansion conditions and the amount of memory that needs to be expanded, wherein the amount of memory corresponds to the expanded memory; Based on the amount of memory, a first address segment corresponding to the extended memory is selected from the first address space, where the first address space is the address space corresponding to the memory range on the memory device; Send a memory expansion request, including the amount of memory, to the host so that the host can allocate a dedicated and contiguous memory space for the processor from its memory as the expanded memory, based on the amount of memory. The extended memory has a second address segment in the host's address space for address mapping of the extended memory, and the second address segment corresponds to the first address segment.

5. The device according to claim 4, wherein, The processor is used for: A first address translation page table is configured on the first address translation component of the host. The first address translation page table records the address mapping relationship between the second address segment and the third address segment. The third address segment is the address segment that the first address segment maps to on the interconnect device. or A second address translation page table is configured on the second address translation component of the acceleration device. The second address translation page table records the address mapping relationship between the second address segment and the first address segment.

6. The device according to claim 4, wherein, The processor is further configured to: upon successful allocation of the extended memory to the processor, allocate a fourth address segment in the second address space corresponding to the first address segment, and add the fourth address segment to the memory management system of the acceleration device to expose the extended memory to the application; the second address space is a portion of the address space corresponding to the memory range in the processor's address space.

7. The device according to any one of claims 1-6, wherein, The memory expansion device includes: a memory device based on a memory expansion protocol and an interconnect device based on a physical interconnect protocol; The processor is configured to: send a first memory access request to the memory device based on the memory extension protocol, wherein the first memory access request is used to request access to the extended memory; The memory device is further configured to: convert the first memory access request into a second memory access request based on the physical interconnect protocol, and then send the second memory access request to the host via the interconnect device so that the host can access the extended memory.

8. The device according to any one of claims 2-6, wherein, The memory interconnect protocol is an open interconnect protocol, and the physical interconnect protocol is a high-speed peripheral interconnect bus standard.

9. A memory expansion method applied to an acceleration device, wherein a processor and a memory expansion device in the acceleration device cooperate to form a memory access channel between the processor and a host, so as to use the host's memory as a memory resource for memory expansion of the processor; the method includes: Determine the range of memory that the memory expansion device can expand for the processor from the host's memory; If the processor meets the memory expansion conditions, it requests additional memory from the host's memory up to the specified memory range. The extended memory is address-mapped so that it can be accessed through the memory access channel.

10. The method according to claim 9, wherein, The memory expansion device includes: a memory device based on a memory expansion protocol and an interconnect device based on a physical interconnect protocol; Determining the range of memory that the memory expansion device can expand for the processor from the host's memory includes: If the memory device is detected, the processor receives a memory range reported by the memory device that it can expand from the host's memory, the memory range corresponding to a first address space on the memory device; In the processor's address space, a second address space corresponding to the memory range is allocated.

11. The method according to claim 9, wherein, The memory expansion device includes: a memory device based on a memory expansion protocol and an interconnect device based on a physical interconnect protocol; If the processor meets the memory expansion conditions, requesting extended memory from the host's memory up to the specified memory range includes: Based on the current memory usage of the processor, determine whether the processor meets the memory expansion conditions and the amount of memory that needs to be expanded, wherein the amount of memory corresponds to the expanded memory; Based on the amount of memory, a first address segment corresponding to the extended memory is selected from the first address space, where the first address space is the address space corresponding to the memory range on the memory device; A memory expansion request, including the amount of memory, is sent to the host so that the host can allocate a dedicated and contiguous memory space for the processor from its memory as the expanded memory, based on the amount of memory; the expanded memory has a second address segment in the host's address space for address mapping of the expanded memory, and the second address segment corresponds to the first address segment.

12. The method according to claim 11, wherein, Sending a memory expansion request, including the amount of memory, to the host so that the host can allocate a dedicated and contiguous memory space from its memory for the processor as the expanded memory, includes: A memory expansion request is sent to the host via a control channel, so that the host can allocate a dedicated and contiguous memory space for the processor from its memory as the expanded memory, based on the amount of memory. The control channel is pre-established by the processor with the host based on a custom channel establishment instruction.

13. The method according to claim 11, wherein, Mapping the extended memory to its address so that it can be accessed through the memory access channel includes: A first address translation page table is configured on the first address translation component of the host to access the extended memory through the memory access channel; the first address translation page table records the address mapping relationship between the second address segment and the third address segment, wherein the third address segment is the address segment mapped by the first address segment on the interconnect device; or A second address translation page table is configured on the second address translation component of the acceleration device to access the extended memory through the memory access channel; the second address translation page table records the address mapping relationship between the second address segment and the first address segment.

14. The method of claim 11, further comprising: If the extended memory is successfully allocated to the processor, a fourth address segment corresponding to the first address segment is allocated in the second address space; The second address space is the portion of the memory range within the processor's address space. The fourth address segment is added to the memory management system of the acceleration device to expose the extended memory to the application.

15. The method according to any one of claims 11-14, further comprising: Monitor the popularity information of the data stored in the extended memory, and migrate the first data in the extended memory whose popularity information is greater than a first popularity threshold to the local memory of the processor; And / or, Monitor the popularity information of the data stored in the processor's local memory, and migrate the second data whose popularity information in the local memory is less than a second popularity threshold to the extended memory; The first heat threshold is greater than the second heat threshold.

16. A memory access method applied to an acceleration device, wherein a processor and a memory expansion device in the acceleration device cooperate to form a memory access channel between the processor and a host, so as to use the host's memory as a memory resource for memory expansion of the processor; the method includes: Request extended memory for the processor from the host's memory in advance; The extended memory on the host is accessed based on the memory access channel.

17. The method according to claim 16, wherein, The memory expansion device includes: a memory device based on a memory expansion protocol and an interconnect device based on a physical interconnect protocol; Accessing the extended memory on the host based on the memory access channel includes: A first memory access request is sent based on the memory extension protocol, the first memory access request being used to request access to the extended memory; The first memory access request is converted into a second memory access request based on the physical interconnect protocol and then sent to the host so that the host can access the extended memory.

18. The method according to claim 17, wherein, The first memory access request includes fourth address information in the fourth address segment, the fourth address segment being the address segment in the second address space corresponding to the first address segment, the second address space being the partial address space corresponding to the memory range in the processor's address space, the first address segment being the address segment corresponding to the extended memory in the first address space, and the first address space being the address space corresponding to the memory range on the memory device; After converting the first memory access request into a second memory access request based on the physical interconnect protocol, it is sent to the host so that the host can access the extended memory, including: The fourth address information is converted into first address information in the first address segment; second address information or third address information is obtained based on the first address information; a second memory access request is generated according to the second address information or third address information, wherein the second address information or third address information is address information in the second address segment or third address segment obtained by converting the first address information; wherein the second address segment is the address segment that the extended memory has in the address space of the host, and the third address segment is the address segment that the first address segment is mapped to on the interconnect device; The interconnect device sends the second memory access request to the host. The second memory access request includes the second address information or the third address information. The third address information is used by the host to convert the second address information, and the second address information is used by the host to access the extended memory.

19. An acceleration device, comprising: Processor, memory expansion devices, and local memory; The memory expansion device is interconnected with the processor and the host respectively, and is used to form a memory access channel between the processor and the host, so as to use the host's memory as a memory resource for memory expansion of the processor; The local memory stores a computer program, and the processor is used to run the computer program to perform the steps of the method according to any one of claims 9-15 and 16-18.

20. A computer-readable storage medium storing a computer program / instructions, wherein, When the computer program is executed by a processor, it causes the processor to perform the steps of the method according to any one of claims 9-15 and 16-18.

21. A computer program product, comprising: A computer program / instruction that, when executed by a processor, causes the processor to perform the steps of the method according to any one of claims 9-15 and 16-18.