Memory access method and apparatus, server, and storage medium

CN116126742BActive Publication Date: 2026-09-04INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310073814.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-30
Publication Date
2026-09-04
Estimated Expiration
2043-01-30

AI Technical Summary

Technical Problem

[0004]由于Host设备的主存空间是有限的,基于上述技术方案进行内存访问,会占用Host设备原有的主存,进而影响Host设备的性能

Benefits of technology

[0018]在第一Host设备需要对第二Host设备中目标数据进行访问时,直接访问内存扩展卡,内存扩展卡通过分配给第二Host设备的目标扩展内存处理第一Host设备的内存访问请求,生成内存访问结果,并将内存访问结果反馈给第一Host设备,从而在实现Host设备间相互访问内存的同时,不占用Host设备(如CPU、GPU)的本地内存空间,提升Host设备的处理能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116126742B_ABST
    Figure CN116126742B_ABST
Patent Text Reader

Abstract

The application relates to a memory access method and device, a server and a storage medium, and particularly relates to the technical field of computers. The method is applied to a server, the server comprises a Host device and a memory expansion card, and the method comprises the following steps: a first Host device generates a memory access request for target data in a second Host device; in response to the memory access request, the memory expansion card calls target expansion memory to process the memory access request, generates a memory access result, the memory expansion card comprises multiple parts of expansion memory allocated to different Host devices, and the target expansion memory is expansion memory allocated to the second Host device; and the memory expansion card feeds back the memory access result to the first Host device. Based on the above technical scheme, the memory of the Host devices can be accessed simultaneously without occupying the local memory space of the Host devices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, specifically to a memory access method, apparatus, server, and storage medium. Background Technology

[0002] In today's rapidly developing information and digital industry, the demand for memory in servers is also increasing as the amount of information processed increases.

[0003] In relevant technical solutions, if a host device needs to access data from other host devices, the data can be copied in batches to the host device's local main memory for local processing, or it can use the PCIe channel to directly access the main memory of other host devices for processing.

[0004] Since the main memory space of the host device is limited, memory access based on the above technical solution will occupy the original main memory of the host device, thereby affecting the performance of the host device. Summary of the Invention

[0005] This application provides a memory access method, apparatus, server, and storage medium. The technical solution is as follows.

[0006] On one hand, a memory access method is provided, the method being applied in a server, the server including: a host device and a memory expansion card, the method comprising:

[0007] The first host device generates a memory access request for the target data in the second host device;

[0008] In response to the memory access request, the memory expansion card calls the target extended memory to process the memory access request and generates a memory access result. The memory expansion card includes multiple extended memory parts allocated to different host devices, and the target extended memory is the extended memory allocated to the second host device.

[0009] The memory expansion card feeds back the memory access results to the first host device.

[0010] In another aspect, a memory access device is provided, the device comprising:

[0011] The access request generation module is used to enable the first host device to generate memory access requests for target data in the second host device.

[0012] An access request processing module is used to enable the memory expansion card to respond to the memory access request, call the target extended memory to process the memory access request, and generate a memory access result. The memory expansion card includes multiple parts of extended memory allocated to different host devices. The target extended memory is the extended memory allocated to the second host device.

[0013] The access result feedback module is used to allow the memory expansion card to feed back the memory access result to the first Host device.

[0014] In another aspect, a server is provided, which includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, code set, or instruction set, and the at least one instruction, at least one program, code set, or instruction set is loaded and executed by the processor to implement the above-described memory access method.

[0015] In another aspect, a computer-readable storage medium is provided, wherein at least one instruction, at least one program, code set, or instruction set is stored therein, wherein the at least one instruction, at least one program, code set, or instruction set is loaded and executed by a processor to implement the memory access method described above.

[0016] Furthermore, a computer program product or computer program is also provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the aforementioned memory access method.

[0017] The technical solution provided in this application may include the following beneficial effects:

[0018] When the first host device needs to access target data in the second host device, it directly accesses the memory expansion card. The memory expansion card processes the memory access request of the first host device through the target expanded memory allocated to the second host device, generates the memory access result, and feeds the memory access result back to the first host device. In this way, while realizing mutual memory access between host devices, it does not occupy the local memory space of the host device (such as CPU or GPU), thereby improving the processing capability of the host device. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the specific embodiments of this application or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0020] Figure 1 This is a schematic diagram illustrating the interior of a server according to an exemplary embodiment.

[0021] Figure 2 This is a flowchart illustrating a memory access method according to an exemplary embodiment.

[0022] Figure 3 This is a schematic diagram of a memory expansion card according to an exemplary embodiment.

[0023] Figure 4 This is a flowchart illustrating a memory access method according to an exemplary embodiment.

[0024] Figure 5 This is a flowchart illustrating a memory access method according to an exemplary embodiment.

[0025] Figure 6 This is a schematic diagram of a memory expansion card according to an exemplary embodiment.

[0026] Figure 7 This is a structural block diagram of a memory access device according to an exemplary embodiment.

[0027] Figure 8 This is a schematic diagram of a server provided according to an exemplary embodiment. Detailed Implementation

[0028] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0029] It should be understood that the term "instruction" mentioned in the embodiments of this application can be a direct instruction, an indirect instruction, or an indication of a relationship. For example, A instructing B can mean that A directly instructs B, such as B being able to obtain information through A; it can also mean that A indirectly instructs B, such as A instructing C, so B can obtain information through C; or it can mean that there is a relationship between A and B.

[0030] In the description of the embodiments of this application, the term "correspondence" may indicate that there is a direct or indirect correspondence between two things, or that there is an association between two things, or that there is a relationship of instruction and being instructed, configuration and being configured, etc.

[0031] In the embodiments of this application, "predefined" can be achieved by pre-storing corresponding codes, tables or other means that can be used to indicate relevant information in the device (e.g., including terminal devices and network devices). This application does not limit the specific implementation method.

[0032] With the rapid development of the information and digital industry in today's society, the demand for memory in data center servers and AI servers is increasing rapidly due to the increasing number of CPU (Central Processing Unit) cores and the rising volume of information processing. Simultaneously, the surge in data processing by GPUs (Graphics Processing Units) is also leading to a significant increase in GPU memory requirements. In cloud services, AI (Artificial Intelligence) servers, and other computationally intensive scenarios, insufficient memory can become a bottleneck limiting server performance. How to efficiently expand memory has been a persistent research topic in the industry. Expanding CPU memory can improve overall server performance. However, for AI servers, the computational load of GPUs is high, resulting in similarly high memory requirements; insufficient memory will limit the data processing efficiency of GPUs.

[0033] In the above context, combined with references Figure 1 The memory access methods provided in the related technologies include the following two:

[0034] One method is to copy data in batches to the device's local main memory. For example, after the CPU finishes processing, the data moves from the CPU to the Cache and then to the Memory. The GPU then copies the data from the CPU's main memory to its own main memory for local processing.

[0035] One method is to directly access the host's main memory to process data. For example, the GPU uses the PCIe channel to directly access the CPU's main memory.

[0036] Of the two memory access methods mentioned above, the former has a bandwidth of 50GT / s and a latency of approximately 50ns, while the latter has only 32GT / s (PCIe 5.0) and a latency of approximately 100ns+.

[0037] Since the main memory space of the CPU and GPU is limited, the two memory access methods mentioned above will occupy the original main memory of the CPU and GPU, which will affect the performance of the CPU and GPU.

[0038] To avoid the above-mentioned drawbacks, this application proposes a memory expansion and management scheme based on CXL (Compute ExpressLink, an open interconnect standard), which expands the memory of the CPU and GPU without occupying the local memory space of the CPU and GPU, and greatly improves the processing power of the CPU and GPU.

[0039] The technical solution provided in this application will be further described below with reference to the following embodiments.

[0040] Figure 2 This is a flowchart illustrating a memory access method according to an exemplary embodiment. The method is applied in a server, which includes a host device and a memory expansion card, such as... Figure 2 As shown, the memory access method may include the following steps:

[0041] Step 210: The first host device generates a memory access request for the target data in the second host device.

[0042] In this embodiment, the server includes a Host device and a memory expansion card. The Host device, also known as the master device, refers to the device that takes the initiative; for example, in the case of a CPU and memory, the CPU is the master device, and the memory is the slave device. The memory expansion card refers to a card with memory expansion capabilities.

[0043] In this embodiment, there can be multiple host devices, and the host devices have mutual memory access requirements. Here, we will use the example of the first host device needing to access the target data of the second host device as an example. Therefore, the first host device generates a memory access request for the target data in the second host device.

[0044] In one possible implementation, the host device includes: a CPU; and a GPU. For example, the first host device is a CPU, the second host device is a GPU, and the CPU needs to access memory from the GPU. Alternatively, the first host device is a GPU, the second host device is a CPU, and the GPU needs to access memory from the CPU.

[0045] In one possible implementation, the hardware structure of the memory expansion card includes the following components connected by wiring: an expansion memory that plugs into a memory slot; a CXL-based expansion chip for managing the expansion memory; board contacts for power supply and connection to the host device; and a connector interface for connecting to the host device.

[0046] In this implementation, a memory expansion card is provided. This memory expansion card is installed in a server and includes an expansion chip and expansion memory. Because the CXL bus is physically layer compatible with the PCIe (Peripheral Component Interconnect Express) bus, the memory expansion card can be connected uplink to a PCIe slot within the server chassis, just like a PCIe device. For example, as shown... Figure 3 As shown, in the memory expansion card, 1 is the expansion chip (CXL SWITCH); 2 is the expansion memory slot, into which the expansion memory can be plugged, and the expansion memory is connected to the expansion chip; 3 is the board's gold fingers, which connect to the motherboard and are used for board power supply and CPU connection; 4 is the MCIO (Mini Cool edge Input / Output) connector interface, which connects to the GPU via cable.

[0047] Understandably, the expansion chip uses the CXL protocol, which can speed up memory access. Therefore, memory access performance can be improved by using a CXL-based memory expansion card.

[0048] Step 220: In response to the memory access request, the memory expansion card calls the target extended memory to process the memory access request and generates the memory access result.

[0049] In this embodiment, the extended memory in the memory expansion card may include multiple parts, and different parts of the extended memory can be used to store data from the same or different host devices. Specifically, the target extended memory is the extended memory allocated to the second host device, and the target extended memory stores the target data of the second host device.

[0050] In this embodiment, the memory expansion card is accessed directly. The memory expansion card processes memory access requests through the target extended memory allocated to the second host device and generates memory access results.

[0051] Step 230: The memory expansion card sends the memory access results back to the first host device.

[0052] In this embodiment, the memory access result is fed back to the first host device through the memory expansion card to complete the memory access.

[0053] In summary, the memory access method provided in this application allows the first host device to directly access the memory expansion card when it needs to access target data in the second host device. The memory expansion card processes the memory access request from the first host device through the target extended memory allocated to the second host device, generates a memory access result, and feeds the memory access result back to the first host device. This enables mutual memory access between host devices without occupying the local memory space of the host device (such as CPU or GPU), thereby improving the processing capability of the host device.

[0054] In addition, the expansion chip in the memory expansion card uses the CXL protocol, which can speed up memory access, thereby improving the access performance of memory access through the memory expansion card.

[0055] In an illustrative embodiment, when the amount of data to be processed is not large and the memory access frequency is low, memory access is performed by directly accessing extended memory.

[0056] Figure 4 This is a flowchart illustrating a memory access method according to an exemplary embodiment. The method is applied in a server, which includes a host device and a memory expansion card, such as... Figure 4 As shown, step 220 above can be replaced by the following steps:

[0057] Step 410: In response to the memory access request, determine the access attributes of the target data.

[0058] Among them, access attributes are used to indicate the attributes related to data and memory access. Access attributes include, but are not limited to, the amount of data and the frequency of memory access.

[0059] In this embodiment, the data in the Host device can have different memory access paths depending on the access attributes. Therefore, after the first Host device generates a memory access request for the target data in the second Host device, the access attributes of the target data are queried and determined.

[0060] Step 420: If the access attribute of the target data is a low-performance requirement attribute, the memory access request is passed from the first host device to the memory expansion card.

[0061] Among them, the low performance requirement attribute refers to the target data volume being lower than the quantity threshold and the memory access frequency being lower than the frequency threshold.

[0062] Step 430: The memory expansion card calls the target extended memory to process the memory access request and generates the memory access result.

[0063] In this embodiment, data in the host device that has a data volume lower than the quantity threshold and a memory access frequency lower than the frequency threshold is placed in the memory expansion card. When a memory access request requests access to such data, memory access is performed through the memory expansion card based on the access attributes of such data.

[0064] In one possible implementation, if the access attribute of the target data is not a low-performance requirement attribute, the target data is obtained from the first host device and copied to the main memory of the second host device; the second host device uses the main memory to process memory access requests and generates and obtains memory access results.

[0065] In this implementation, data from other host devices that have a data volume exceeding a quantity threshold or a memory access frequency exceeding a frequency threshold are copied to the main memory of this host device. When a memory access request requests access to this type of data, memory access is performed locally based on the access attributes of this type of data.

[0066] For example, based on the above description, when the amount of data to be processed is large, the number of iterations is high, and memory access is frequent, such as in AI training and graphics rendering, the data should be copied in batches to the device's internal storage for local processing, according to the processing methods provided in the relevant technologies. When the amount of data to be processed is not very large and the memory access frequency is low, the extended memory can be accessed directly.

[0067] Understandably, directly accessing the device's internal memory and processing data locally is a memory access method with high bandwidth, low latency, and good access performance.

[0068] In summary, the memory access method provided in this application can directly access extended memory when the amount of data to be processed is not large and the memory access frequency is low. When the amount of data to be processed is large, the number of iterations is high, and memory access is frequent, the data is copied in batches to the device's internal storage for local processing. In this way, by using different memory access paths, the processing performance of local memory and the access performance of memory access can be guaranteed as much as possible.

[0069] In an illustrative embodiment, pooled resources are released and allocated for the extended memory in the memory expansion card.

[0070] Figure 5 This is a flowchart illustrating a memory access method according to an exemplary embodiment. The method is applied in a server, which includes a host device and a memory expansion card, such as... Figure 5 As shown, in addition to the steps described above, the following steps may also be included:

[0071] Step 510: The memory expansion card obtains the memory usage requirements of each host device.

[0072] Among them, the extended memory usage requirement refers to the host device's usage requirement for the extended memory in the memory expansion card.

[0073] Step 520: The memory expansion card dynamically allocates the extended memory in the expansion memory card based on the extended memory usage requirements of each host device.

[0074] In this embodiment, the memory expansion card allocates the free expansion memory in the expansion memory card to the host device that currently has an expansion memory usage requirement, based on the expansion memory usage needs of each host device. Alternatively, it releases the expansion memory in the expansion memory card that was originally allocated to the host device and that the host device currently does not have an expansion memory usage requirement.

[0075] In one possible implementation, step 510 includes: the memory expansion card obtaining the current main memory utilization rate of each host device; step 520 includes: releasing the extended memory allocated to the target host device when the current main memory utilization rate of the target host device is lower than a first utilization rate threshold; and allocating the free extended memory to the target host device when the current main memory utilization rate of the target host device is higher than a second utilization rate threshold.

[0076] In this implementation, the extended memory usage requirements of each host device are measured by the main memory utilization rate of each host device. If the current main memory utilization rate of the target host device is lower than the first utilization rate threshold, it is considered that the target host device does not currently have an extended memory usage requirement, and therefore, the extended memory allocated to the target host device is released. If the current main memory utilization rate of the target host device is higher than the second utilization rate threshold, it is considered that the target host device currently has an extended memory usage requirement, and therefore, the free extended memory is allocated to the target host device.

[0077] For example, in conjunction with reference Figure 6 Dn(D1, D2...) represents extended memory, and Hn(H1, H2...) represents the host device, such as a CPU or GPU. Extended memory (Dn) can be dynamically allocated to host devices (CPU and GPU) on demand. For example, D1 and D2 belong to H1, and D3 belongs to H2. Suppose that after a period of time, H1 no longer needs device D2, and D2 can be released. Suppose that after another period of time, H2 needs more acceleration devices, it checks the extended memory, finds D2 is free, and D2 will be allocated to H2. As a result, H2 now has D2 and D3, thus completing a full cycle of pooled resource release and allocation.

[0078] In summary, the memory access method provided in this application obtains the extended memory usage requirements of each host device, releases and allocates pooled resources of the extended memory in the memory expansion card, thereby effectively utilizing the extended memory in the memory expansion card.

[0079] It should be noted that the above method embodiments can be implemented individually or in combination, and this application does not limit them in this regard.

[0080] Figure 7 This is a structural block diagram of a memory access device according to an exemplary embodiment.

[0081] The device includes:

[0082] The access request generation module 701 is used to enable the first host device to generate a memory access request for target data in the second host device.

[0083] The access request processing module 702 is used to enable the memory expansion card to respond to the memory access request, call the target extended memory to process the memory access request, and generate a memory access result. The memory expansion card includes multiple parts of extended memory allocated to different host devices. The target extended memory is the extended memory allocated to the second host device.

[0084] The access result feedback module 703 is used to allow the memory expansion card to feed back the memory access result to the first Host device.

[0085] In one possible implementation, the access request processing module 702 is configured to:

[0086] In response to the memory access request, determine the access attributes of the target data;

[0087] If the access attribute of the target data is a low-performance requirement attribute, the memory access request is passed from the first Host device to the memory expansion card. The low-performance requirement attribute means that the amount of the target data is lower than the quantity threshold and the memory access frequency is lower than the frequency threshold.

[0088] The memory expansion card calls the target extended memory to process the memory access request and generates the memory access result.

[0089] In one possible implementation, the access request processing module 702 is configured to:

[0090] If the access attribute of the target data does not belong to the low performance requirement attribute, the target data is obtained from the first host device and copied to the main memory of the second host device;

[0091] The second host device uses main memory to process the memory access request, and generates and obtains the memory access result.

[0092] In one possible implementation, the apparatus further includes: an extended memory management module; the extended memory management module is configured to:

[0093] The memory expansion card obtains the memory usage requirements of each host device;

[0094] The memory expansion card dynamically allocates the expanded memory in the expansion memory card based on the expanded memory usage requirements of each host device.

[0095] In one possible implementation, the extended memory management module is used for:

[0096] The memory expansion card obtains the current main memory utilization rate of each host device;

[0097] The memory expansion card dynamically allocates the expanded memory in the expansion memory card based on the expanded memory usage requirements of each host device, including:

[0098] If the current main memory utilization of the target host device is lower than the first utilization threshold, the extended memory allocated to the target host device will be released.

[0099] If the current main memory utilization of the target host device is higher than the second utilization threshold, the free extended memory is allocated to the target host device.

[0100] In one possible implementation, the hardware structure of the memory expansion card includes the following components connected by wiring:

[0101] Extended memory that plugs into a memory slot;

[0102] CXL-based expansion chip for managing extended memory;

[0103] Gold fingers of the board used for power supply and connection to the host device;

[0104] Connector interface for connecting to host devices.

[0105] In one possible implementation, the Host device includes:

[0106] Central Processing Unit (CPU);

[0107] Graphics Processor (GPU).

[0108] It should be noted that the memory access device provided in the above embodiments is only an example of the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.

[0109] Please see Figure 8 This is a schematic diagram of a server provided according to an exemplary embodiment of the present application. The server includes a memory and a processor. The memory is used to store a computer program. When the computer program is executed by the processor, it implements the memory access method described above.

[0110] The processor can be a central processing unit (CPU). It can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or combinations thereof.

[0111] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the methods in the embodiments of this application. The processor executes various functional applications and data processing by running the non-transitory software programs, instructions, and modules stored in the memory, thereby implementing the methods in the above-described embodiments.

[0112] The memory may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created by the processor, etc. Furthermore, the memory may include high-speed random access memory and non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory may optionally include memory remotely located relative to the processor, which can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0113] In one exemplary embodiment, a computer-readable storage medium is also provided for storing at least one computer program, which is loaded and executed by a processor to implement all or part of the steps in the above-described method. For example, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, or optical data storage device, etc.

[0114] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.

[0115] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. A memory access method, characterized in that, The method is applied to a server, the server including a host device and a memory expansion card, the method comprising: The first host device generates a memory access request for the target data in the second host device; In response to the memory access request, the memory expansion card calls the target extended memory to process the memory access request and generates a memory access result. The memory expansion card includes multiple extended memory portions allocated to different host devices, and the target extended memory is the extended memory allocated to the second host device. The memory expansion card is based on the CXL protocol. The memory expansion card feeds back the memory access results to the first Host device; The memory expansion card obtains the memory usage requirements of each host device; The memory expansion card dynamically allocates the expanded memory in the memory expansion card based on the expanded memory usage requirements of each host device; The memory expansion card obtains the expanded memory usage requirements of each host device, including: The memory expansion card obtains the current main memory utilization rate of each host device; The memory expansion card dynamically allocates the expanded memory in the memory expansion card based on the expanded memory usage requirements of each host device, including: If the current main memory utilization of the target host device is lower than the first utilization threshold, the extended memory allocated to the target host device will be released. If the current main memory utilization of the target host device is higher than the second utilization threshold, the free extended memory is allocated to the target host device.

2. The method according to claim 1, characterized in that, In response to the memory access request, the memory expansion card invokes the target extended memory to process the memory access request and generates a memory access result, including: In response to the memory access request, determine the access attributes of the target data; If the access attribute of the target data is a low-performance requirement attribute, the memory access request is passed from the first Host device to the memory expansion card. The low-performance requirement attribute means that the amount of the target data is lower than the quantity threshold and the memory access frequency is lower than the frequency threshold. The memory expansion card calls the target extended memory to process the memory access request and generates the memory access result.

3. The method according to claim 2, characterized in that, The method further includes: If the access attribute of the target data does not belong to the low performance requirement attribute, the target data is obtained from the first host device and copied to the main memory of the second host device; The second host device uses main memory to process the memory access request, and generates and obtains the memory access result.

4. The method according to any one of claims 1 to 3, characterized in that, The hardware structure of the memory expansion card includes the following components connected by wiring: Extended memory that plugs into a memory slot; CXL-based expansion chip for managing extended memory; Gold fingers of the board used for power supply and connection to the host device; Connector interface for connecting to host devices.

5. The method according to any one of claims 1 to 3, characterized in that, The Host device includes: Central Processing Unit (CPU); Graphics Processor (GPU).

6. A memory access device, characterized in that, The device includes: The access request generation module is used to enable the first host device to generate memory access requests for target data in the second host device. An access request processing module is used for the memory expansion card to respond to the memory access request, call the target extended memory to process the memory access request, and generate a memory access result. The memory expansion card includes multiple parts of extended memory allocated to different host devices, and the target extended memory is the extended memory allocated to the second host device. The memory expansion card is based on the CXL protocol. An access result feedback module is used to allow the memory expansion card to feed back the memory access result to the first Host device; The access request processing module is further configured to: obtain the extended memory usage requirements of each host device through the memory expansion card; dynamically allocate extended memory in the memory expansion card based on the extended memory usage requirements of each host device; and obtain the extended memory usage requirements of each host device through the memory expansion card, including: obtaining the current main memory utilization rate of each host device; releasing the extended memory allocated to the target host device if the current main memory utilization rate of the target host device is lower than a first utilization rate threshold; and allocating the free extended memory to the target host device if the current main memory utilization rate of the target host device is higher than a second utilization rate threshold.

7. A server, characterized in that, The server includes a processor and a memory, the memory storing at least one instruction, at least one program, code set, or instruction set, the at least one instruction, at least one program, code set, or instruction set being loaded and executed by the processor to implement the memory access method as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, at least one program, code set, or instruction set is loaded and executed by a processor to implement the memory access method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Method and device for expanding synchronous memory bus function

    CN105095138A