System and memory access method
Patent Information
- Application Number
- PCT/CN2026/080141
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-26
- Filing Date
- 2026-02-26
- Publication Date
- 2026-10-01
Smart Images

Figure CN2026080141_01102026_PF_FP_ABST
Abstract
Description
A system and a memory access method
[0001] This application claims priority to Chinese Patent Application No. 202510370685.7, filed on March 26, 2025, entitled "A System and a Memory Access Method", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of computing, and more particularly to a system and a memory access method. Background Technology
[0003] As data center businesses grow, memory costs account for an increasingly larger proportion of data center server costs. However, the average memory utilization rate in data centers is less than 50%, indicating significant room for improvement in memory application efficiency.
[0004] Existing technologies can enable memory reuse among multiple virtual machines (VMs) within a single server to improve memory utilization efficiency. However, this improvement in memory utilization is limited to within a single server and cannot achieve cross-server (e.g., data center) memory utilization improvement. Summary of the Invention
[0005] This application provides a system and a memory access method that enables cross-server memory access, reducing memory waste at the system level (e.g., data center level) and thus lowering costs.
[0006] In a first aspect, this application provides a system, wherein the system may be a data unit comprising multiple servers, the system including a first computing unit and a second computing unit that are communicatively connected to each other, and the first computing unit and the second computing unit are units on different servers, wherein in order to access the target memory space on the second computing unit, the first computing unit (e.g., the central processing unit (CPU) of the first computing unit) can obtain the physical address of the target memory space on the second computing unit, and construct a page table indicating the mapping relationship between the physical address and the virtual address based on the physical address, and then the first computing unit can determine the corresponding target physical address based on the target virtual address (i.e., the address of the target memory space to be accessed) through the page table, and access the target memory space on the second computing unit (i.e., access the memory space corresponding to the target virtual address) based on the target physical address.
[0007] In this embodiment, the first computing unit can construct a page table indicating the mapping relationship between virtual addresses and physical addresses based on the physical addresses of the memory spaces of computing units on other servers. Based on this page table, the first computing unit can directly access the memory spaces of computing units on other servers through virtual addresses in the same way as local memory access, thereby realizing cross-server memory access. This can improve the memory utilization of a system containing multiple servers. For a single server, since it can use the memory resources of other servers, it can increase the upper limit of the memory resources that the server can use and the upper limit of the access bandwidth.
[0008] The page table indicates the physical address corresponding to each virtual address in a virtual memory space, and the target virtual address can be one or more virtual addresses in the virtual memory space.
[0009] In one possible implementation, the second computing unit and the first computing unit are connected via a high-speed interconnect bus.
[0010] In one possible implementation, the first computing unit is a computing unit on a single server, and the second computing unit is a computing unit on multiple servers. That is, the first computing unit can access the memory resources of multiple servers. When the memory space requirement of the first computing unit is large, and the available memory space on the computing units of other servers is insufficient to meet the memory space requirement of the first computing unit, the memory space on multiple computing units can be allocated to the first computing unit for use, thus enabling cross-server memory access in the above scenario.
[0011] In one possible implementation, the first computing unit is specifically configured to: receive the physical address of the target memory space sent by the second computing unit or the management unit.
[0012] In one possible implementation, the system further includes: a management unit; the management unit is configured to acquire first information; the first information indicates the size of allocatable memory space on the second computing unit; the first computing unit is further configured to send a first request to the management unit, the first request indicating the amount of memory to be used; the management unit is further configured to, if it is determined from the first information that there is allocatable memory space on the second computing unit that meets the memory requirement, send a second request to the second computing unit, the second request instructing the second computing unit to allocate memory space of the specified memory amount; wherein, the target memory space is the memory space of the specified memory amount. In this embodiment, a server (a management unit within the server) is used as the manager to maintain memory information on other computing units in the system (e.g., the first information in this embodiment), the first information indicating the size of allocatable memory space on each computing unit. When the management unit receives a request from a computing unit to use memory (e.g., a first request from the first computing unit), it can determine, based on the memory space requirement in the request, which (or several) computing units in the system have allocable memory space that meets the memory space requirement in the request (e.g., the second computing unit has memory space that meets the memory requirement indicated in the first request). Then, it can request the second computing unit to allocate memory space that meets the memory requirement for the first computing unit to use. This achieves unified management of memory resources across servers, which can improve the memory utilization of a system containing multiple servers. For a single server, since it can use the memory resources of other servers, it can increase the upper limit of the memory resources that the server can use and the upper limit of the access bandwidth.
[0013] In one possible implementation, the management unit is further configured to send a third request to the first computing unit when, based on the first information, it is determined that there is allocatable memory space on the second computing unit that meets the memory requirement. The third request instructs the first computing unit to use the memory space on the second computing unit. This is equivalent to the management unit informing the first computing unit which computing unit's memory it can access. Furthermore, after receiving the physical address of the target memory space, the first computing unit proactively constructs a page table indicating the mapping relationship between the physical address and the virtual address based on the physical address.
[0014] In one possible implementation, the second computing unit is further configured to send memory update information to the management unit; the memory update information indicates that the allocatable memory space on the second computing unit has been reduced by the size of the memory; the management unit updates the size of the allocatable memory space on the second computing unit as indicated in the first information according to the memory update information; or, the management unit may store third information in addition to the first information, the third information indicating that the allocatable memory space on the second computing unit has been reduced by the size of the target memory space.
[0015] In one possible implementation, the allocatable memory space on the second computing unit indicated in the first information includes a plurality of first memory pages, wherein the first memory pages are zero-page memory; the first information also includes a first identifier of the same operation record corresponding to the plurality of first memory pages; the second information indicates the size of the allocatable memory space on the third computing unit; the allocatable memory space on the third computing unit indicated in the second information includes a plurality of second memory pages, wherein the second memory pages are zero-page memory; the second information also includes a second identifier of the same operation record corresponding to the plurality of second memory pages; the first identifier and the second identifier are the same.
[0016] The memory update information may include the operation identifier (or operation record) of the memory release or occupation operation, such as the page frame number (PFN). In existing implementations, operation records are recorded at the memory page level. For example, for a memory space containing memory page 1, memory page 2, ..., memory page n, releasing it requires n records (i.e., n PFNs, each corresponding to one record), which leads to excessively large memory requirements. Since memory pages in the allocatable memory space are often zero-page memory, in this embodiment, multiple memory pages in the allocatable memory space can be mapped to the same operation record identifier using the same operation record, thereby reducing storage requirements and the difficulty of information maintenance.
[0017] In one possible implementation, the identifier is the page frame number (PFN).
[0018] In one possible implementation, before the second computing unit receives the second request, the allocatable memory space on the second computing unit is the space occupied by the kernel driver (e.g., but not limited to the balloon driver) on the second computing unit. The method further includes: after the second computing unit receives the second request, the kernel driver on the second computing unit releases the amount of memory space from the allocatable memory space.
[0019] In one possible implementation, the first information is further used to indicate the access bandwidth on the second computing unit; when the management unit determines, based on the first information, that there is allocatable memory space on the second computing unit that meets the memory requirement, it sends a second request to the second computing unit, including: when the management unit determines, based on the first information, that there is allocatable memory space on the second computing unit that meets the memory requirement, and that the access bandwidth on the second computing unit meets a preset condition, it sends a second request to the second computing unit.
[0020] For a computing unit, allocating too much of its memory space to other computing units can lead to excessive pressure on memory access bandwidth (memory access bandwidth refers to the rate at which a memory module transfers data per unit of time; excessive access bandwidth pressure essentially means approaching the maximum access bandwidth of the computing unit). Therefore, when allocating memory, the management unit can also consider the access bandwidth of each computing unit. When the access bandwidth meets preset conditions (e.g., when the access bandwidth is less than a preset value), memory can be allocated from that computing unit for use by other computing units.
[0021] When a computing unit is allocated a large amount of memory space, the memory access bandwidth is often large. Even if there is still remaining allocable memory space, it is not suitable to allocate the remaining allocable memory space due to the limitation of memory access bandwidth. Instead, some computing units with remaining allocable memory space and lower memory access bandwidth can be selected for memory space allocation. This can achieve the balance of memory access bandwidth in the system, which is equivalent to expanding the memory bandwidth.
[0022] Secondly, this application provides a memory access method applied to a first computing unit, wherein the first computing unit and a second computing unit are communicatively connected; wherein the first computing unit and the second computing unit are units on different servers; the method includes: obtaining the physical address of a target memory space on the second computing unit; constructing a page table based on the physical address; the page table indicating the mapping relationship between the physical address and a virtual address; determining the corresponding target physical address based on the target virtual address through the page table; and accessing the target memory space on the second computing unit based on the target physical address.
[0023] In one possible implementation, the second computing unit and the first computing unit are connected via a high-speed interconnect bus.
[0024] In one possible implementation, the first computing unit is a computing unit on a single server, and the second computing unit is a computing unit on multiple servers.
[0025] In one possible implementation, the first computing unit obtains the physical address of the target memory space on the second computing unit, including receiving the physical address of the target memory space sent by the second computing unit or the management unit.
[0026] In one possible implementation, the first information is further used to indicate the access bandwidth on the second computing unit; when the management unit determines, based on the first information, that there is allocatable memory space on the second computing unit that meets the memory requirement, it sends a second request to the second computing unit, including: when the management unit determines, based on the first information, that there is allocatable memory space on the second computing unit that meets the memory requirement, and that the access bandwidth on the second computing unit meets a preset condition, it sends a second request to the second computing unit.
[0027] In one possible implementation, when the memory space allocated to a computing unit is large, the memory access bandwidth is often large. Even if there is still remaining allocatable memory space, it is not suitable to allocate the remaining allocatable memory space due to the limitation of memory access bandwidth. Instead, some computing units with remaining allocatable memory space and lower memory access bandwidth can be selected for memory space allocation, thereby achieving a balance of memory access bandwidth in the system, which is equivalent to achieving memory bandwidth expansion.
[0028] Thirdly, this application provides a memory access method applied to a management unit, wherein the management unit, a second computing unit, and a first computing unit are communicatively connected to each other; wherein the first computing unit and the second computing unit are units on different servers. The method includes: obtaining first information; the first information indicating the size of allocatable memory space on the second computing unit; receiving a first request sent by the first computing unit, the first request indicating the amount of memory to be used; and, if the management unit determines, based on the first information, that there is allocatable memory space on the second computing unit that meets the memory requirement, sending a second request to the second computing unit, the second request instructing the second computing unit to allocate memory space for the specified amount of memory.
[0029] In one possible implementation, the second computing unit and the first computing unit are connected via a high-speed interconnect bus.
[0030] In one possible implementation, the first computing unit is a computing unit on a server, and the second computing unit is a computing unit on one or more servers.
[0031] In one possible implementation, the method further includes:
[0032] When the management unit determines, based on the first information, that there is allocatable memory space on the second computing unit that meets the memory requirement, it sends a third request to the first computing unit. The third request is used to instruct the first computing unit to use the memory space on the second computing unit.
[0033] In one possible implementation, the method further includes:
[0034] The management unit receives memory update information sent by the second computing unit; the memory update information indicates that the allocatable memory space of the memory size has been reduced on the second computing unit;
[0035] The management unit updates the size of the allocatable memory space on the second computing unit as indicated in the first information, based on the memory update information; or,
[0036] The management unit may store third information in addition to the first information, the third information indicating that the allocatable memory space on the second computing unit has been reduced to the size of the memory.
[0037] In one possible implementation, the memory space allocated by the second computing unit includes multiple memory pages; the memory update information also includes an identifier of the operation record of the memory space allocated by the second computing unit, and the multiple memory pages correspond to the same identifier.
[0038] In one possible implementation, the identifier is the page frame number (PFN).
[0039] In one possible implementation, the first information is further used to indicate the access bandwidth on the second computing unit; if the management unit determines, based on the first information, that there is allocatable memory space on the second computing unit that meets the memory requirement, it sends a second request to the second computing unit, including:
[0040] When the management unit determines, based on the first information, that there is allocatable memory space on the second computing unit that meets the memory requirement and that the access bandwidth on the second computing unit meets preset conditions, it sends a second request to the second computing unit.
[0041] In one possible implementation, the system further includes a third computing unit, wherein the third computing unit, the first computing unit, and the second computing unit are units on different servers;
[0042] The method further includes:
[0043] The management unit obtains second information, which indicates the access bandwidth on the first computing unit;
[0044] The access bandwidth on the second computing unit meets the preset conditions, including: the relationship between the access bandwidth on the second computing unit and the access bandwidth on the first computing unit meets the preset conditions.
[0045] Fourthly, this application provides a memory access method applied to a second computing unit, wherein the second computing unit, a management unit, and a first computing unit are communicatively connected; wherein the first computing unit and the second computing unit are units on different servers. The method includes: the second computing unit sending first information to the management unit; the first information indicating the size of allocatable memory space on the second computing unit; the second computing unit receiving a second request sent by the management unit, the second request being used to instruct the second computing unit to allocate a certain amount of memory space; and the second computing unit sending the address of a target memory space to the first computing unit, the target memory space being the amount of memory space.
[0046] In one possible implementation, the second computing unit and the first computing unit are connected via a high-speed interconnect bus.
[0047] In one possible implementation, the first computing unit is a computing unit on a server, and the second computing unit is a computing unit on one or more servers.
[0048] In one possible implementation, the second computing unit sending the address of the target memory space to the first computing unit includes: the second computing unit sending the physical address of the target memory space to the first computing unit.
[0049] In one possible implementation, the method further includes: the second computing unit sending memory update information to the management unit; the memory update information indicating that the allocatable memory space on the second computing unit has been reduced by the memory size.
[0050] In one possible implementation, the memory space allocated by the second computing unit includes multiple memory pages; the memory update information also includes an identifier of the operation record of the memory space allocated by the second computing unit, and the multiple memory pages correspond to the same identifier.
[0051] In one possible implementation, the identifier is the page frame number (PFN).
[0052] In one possible implementation, before receiving the second request, the allocatable memory space on the second computing unit is the space occupied by the kernel driver on the second computing unit, and the method further includes:
[0053] After receiving the second request, the kernel driver on the second computing unit releases the amount of memory space from the allocatable memory space.
[0054] In one possible implementation, the first information is also used to indicate the access bandwidth on the second computing unit.
[0055] Fifthly, this application provides a computing unit, which is a first computing unit and a second computing unit communicatively connected to each other; wherein the first computing unit and the second computing unit are units on different servers, comprising: a transceiver module for obtaining the physical address of a target memory space on the second computing unit; a page table construction module for constructing a page table based on the physical address; the page table indicating the mapping relationship between the physical address and the virtual address; and a memory access module for determining the corresponding target physical address based on the target virtual address through the page table, and accessing the target memory space on the second computing unit based on the target physical address.
[0056] In one possible implementation, the second computing unit and the first computing unit are connected via a high-speed interconnect bus.
[0057] In one possible implementation, the first computing unit is a computing unit on a server, and the second computing unit is a computing unit on one or more servers.
[0058] In one possible implementation, the transceiver module is specifically used for:
[0059] The physical address of the target memory space is received from the second computing unit or the management unit.
[0060] In a sixth aspect, this application provides a data processing device, comprising: a processor, a memory, and a bus, wherein: the processor and the memory are connected via the bus;
[0061] The memory is used to store computer programs or instructions;
[0062] The processor is configured to call or execute programs or instructions stored in the memory to implement the steps described in the first aspect and any possible implementation of the first aspect, the steps described in the second aspect and any possible implementation of the second aspect, the steps described in the third aspect and any possible implementation of the third aspect, or the steps described in the fourth aspect and any possible implementation of the first aspect.
[0063] In a seventh aspect, this application provides a computer storage medium including computer instructions that, when executed on an electronic device or server, perform the steps described in the second aspect and any possible implementation thereof, the third aspect and any possible implementation thereof, or the fourth aspect and any possible implementation thereof.
[0064] Eighthly, this application provides a computer program product that, when running on an electronic device or server, performs the steps described in the second aspect and any of the possible implementations of the second aspect, the third aspect and any of the possible implementations of the third aspect, or the fourth aspect and any of the possible implementations of the first aspect.
[0065] Ninthly, this application provides a chip system including a processor for supporting devices in implementing the functions involved in the foregoing aspects, such as transmitting or processing data involved in the foregoing methods; or, information. In one possible design, the chip system further includes a memory for storing program instructions and data necessary for the data processing device. The chip system may be composed of chips or may include chips and other discrete devices.
[0066] Based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods. Attached Figure Description
[0067] Figure 1 is a schematic diagram of the application system framework according to an embodiment of this application;
[0068] Figure 2 is a schematic diagram of the application system framework of an embodiment of this application;
[0069] Figure 3 is a flowchart illustrating a memory access method provided in an embodiment of this application;
[0070] Figures 4 to 7 are schematic diagrams of the application system framework of the embodiments of this application;
[0071] Figure 8 is a detailed flowchart of a memory access method according to an embodiment of this application;
[0072] Figure 9 is a schematic diagram of a data processing device provided in an embodiment of this application;
[0073] Figure 10 is a schematic diagram of an architecture provided in an embodiment of this application;
[0074] Figure 11 is a conceptual partial view of a computer program product provided in an embodiment of this application. Detailed Implementation
[0075] The embodiments of this application are described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. As those skilled in the art will understand, with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0076] The memory access method provided in this application is applied to a system that may include multiple servers. Each server may be internally divided into one or more computing units, such as one or more virtual machines or containers. The system architecture can be understood with reference to Figure 1. Figure 1 is a schematic diagram of a computer system architecture, including server 1, server 2, and a server acting as a manager, with communication connections between server 1, server 2, and the manager server.
[0077] Figure 2 is a schematic diagram of an architecture provided in an embodiment of this application. The management unit shown in Figure 2, acting as a manager, may include at least one processor 110, memory 120, and transceiver module 130.
[0078] Optionally, the management unit may also include a system bus, wherein the processor 110, memory 120, and transceiver module 130 are respectively connected to the system bus. The processor 110 can access the memory 120 through the system bus; for example, the processor 110 can perform data reading and writing or code execution in the memory 120 through the system bus.
[0079] The processor 110 primarily interprets the instructions (or code) of the computer program and processes the data in the computer software. The instructions of the computer program and the data in the computer software can be stored in memory 120 or cache unit 116.
[0080] In this embodiment, processor 110 may be an integrated circuit chip with signal processing capabilities. By way of example and not limitation, processor 110 may be a general-purpose processor, a system-on-chip (SOC), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or any combination of the above devices. The general-purpose processor may be a microprocessor, etc. Each processor 110 includes a memory allocation unit 112 and a processing unit (not shown in the figure). The processing unit is the core, kernel, processing core, or central processing unit (CPU) of processor 110.
[0081] The memory allocation unit 112 is used to control the data interaction between the memory 120 and the processing unit. Specifically, the memory allocation unit 112 can receive memory access requests from the processing unit and control access to memory based on the memory access requests. By way of example and not limitation, in the embodiments of this application, the memory control unit may be a device such as a memory management unit (MMU).
[0082] In this embodiment of the application, the memory allocation unit 112 may be a kernel driver (e.g., but not limited to a balloon driver), for example, it may perform memory allocation tasks based on memory requests from computing unit 1 and computing unit 2.
[0083] In this embodiment, the processing unit and the memory allocation unit 112 can be connected via internal chip lines, such as address lines, to enable communication between the processing unit and the memory allocation unit 112.
[0084] Optionally, each processor 110 may also include a cache unit 116, where the cache is a buffer for data exchange (called a cache). When a processing unit needs to read data, it first looks for the required data in the cache. If the data is found, it is executed directly; otherwise, it looks for the data in memory. Since the cache operates much faster than memory, its role is to help the processing unit run faster.
[0085] Memory 120 provides running space for processes in a computing device. For example, memory 120 can store computer programs (specifically, program code) used to create processes. Memory, also known as internal memory, is used to temporarily store computational data from the processor 110, as well as data exchanged with external storage devices such as hard drives. As long as the computer is running, the processor 110 loads the data that needs to be processed into memory for computation, and after the computation is completed, the processing unit sends the result back out.
[0086] By way of example and not limitation, in this embodiment, memory 120 may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory may be random access memory (RAM), which serves as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory 120 of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0087] The transceiver module 130 can be an interface for communication between servers. For example, the transceiver module 130 can communicate via a high-speed interconnect bus.
[0088] The computing unit 1 and computing unit 2 also include an access module 113 and a memory management module 114. The access module 113 can be used for business processes to access the memory of other computing units, and the memory management module 114 can control the release and occupation of its own memory.
[0089] It should be understood that the structures of the computing devices listed above are merely illustrative examples and are not limited thereto. The computing devices in the embodiments of this application may include various hardware in computer systems in the prior art. For example, the computing device may also include other storage devices besides memory 120, such as disk storage.
[0090] To facilitate understanding of the embodiments of this application, some basic concepts involved in this application will be briefly explained.
[0091] 1. Virtual Address: A virtual address is a specific address in the address space that the operating system can recognize or generate. Its size range is determined by the bitness of the operating system running on the processor. For example, if the operating system running on the processor is 32-bit, the virtual address is also 32-bit, and its virtual address range is 0-0xFFFFFFFF (4GB); if the operating system running on the processor is 64-bit, the virtual address is also 64-bit, and its address space is 0-0xFFFFFFFFFFFFFFFF (16EB).
[0092] Virtual addresses can be divided into multiple virtual address spaces according to actual needs, such as user-mode address space and kernel-mode address space. User-mode address space can be accessed by user-mode programs (such as reading, writing, opening, closing, or drawing) and kernel-mode programs (such as process management, memory management, file management, or device management), while kernel-mode address space can only be accessed by kernel-mode programs at runtime.
[0093] 2. Physical address: This can be a specific address within the address space of a hardware storage device such as internal memory.
[0094] 3. Virtual Page: Also known as a page. The MMU can use a paging mechanism to manage the virtual address space in units of virtual pages. Each page can include a virtual address space of a preset size. A virtual address space can include more than one virtual address, and one virtual address space can correspond to a set of page tables. This set of page tables is used to determine the physical address corresponding to the virtual address in the virtual address space. The set of page tables can include a first-level page table, or it can include a first-level page table and at least one second-level page table. The second-level page table can include second-level page tables, third-level page tables, or even lower-level page tables.
[0095] 4. Physical Pages: In the Linux kernel, the primary function of the MMU is to map "virtual addresses" to real "physical addresses." Specifically, the MMU uses physical pages as the basic unit of memory management. Different architectures support different physical page sizes. For example, in a 32-bit architecture, the supported physical page size is 4 kilobytes (KB), in a 64-bit architecture, the supported physical page size is 8 KB, and in a 128-bit architecture, the supported physical page size is 16 KB, and so on.
[0096] Physical page partitioning usually follows certain requirements. A 4KB page is typically called a regular page, while pages larger than 4KB are typically called "large pages." For example, a 2-megabyte (MB) page and a 1-gigabyte (GB) page are both called large pages. Of course, large pages can have other sizes, but they are usually multiples of 4KB.
[0097] 5. Balloon Driver: The kernel uses a special driver (called a "balloon") to flexibly manage memory. Its core goal is to optimize physical memory utilization and support on-demand memory allocation and reclamation. When memory needs to be reclaimed, the balloon driver requests memory ("inflating"), occupying a portion of the memory space. If it detects insufficient memory, it may trigger internal memory reclamation (such as releasing cache). The physical memory occupied by the balloon is reclaimed and allocated to other computing units or for its own use. When a computing unit needs more memory, the balloon driver releases the previously occupied memory ("deflating"). The computing unit can then regain available memory without needing to allocate additional physical memory.
[0098] In addition, the following points are provided to facilitate understanding of the embodiments of this application.
[0099] First, in this application, "at least one" refers to one or more, and "more than one" refers to two or more. Furthermore, in the embodiments of this application, "first," "second," and various numerical designations (e.g., "#1," "#2," etc.) are merely for descriptive convenience and are not intended to limit the scope of the embodiments of this application. The sequence numbers of the processes below do not imply an order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. It should be understood that the objects described in this way can be interchanged where appropriate to describe solutions other than those in the embodiments of this application. In addition, in the embodiments of this application, the term "201," etc., is merely an identifier for descriptive convenience and does not limit the order of execution steps.
[0100] Second, in the embodiments of this application, the words "exemplary" or "for example" are used to indicate that they are examples, illustrations, or descriptions. Any embodiment or design that is described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design options. Specifically, the use of the words "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0101] Third, the term "storage" in the embodiments of this application can refer to storage in one or more memories. These memories can be separate installations or integrated into an encoder, decoder, processor, or communication device. Alternatively, some memories can be separately installed, while others are integrated into the decoder, processor, or communication device. The type of memory can be any form of storage medium, and this application does not limit this.
[0102] Fourth, the term "comprising" (also referred to as "includes", "including", "comprises" and / or "comprising") used in the embodiments of this application, when used in this specification, specifies the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0103] Fifth, the word "if" in the embodiments of this application can be interpreted as meaning "when" or "upon" or "in response to determination" or "in response to detection". Similarly, depending on the context, the phrase "if it is determined..." or "if [the stated condition or event] is detected" can be interpreted as meaning "when it is determined..." or "in response to determination..." or "when [the stated condition or event] is detected" or "in response to detection of [the stated condition or event]".
[0104] Sixth, the terminology used in the description of the various examples in the embodiments of this application is for the purpose of describing particular examples only and is not intended to be limiting. As used in the description of the various examples and in the appended claims, the numerical forms “a” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise.
[0105] Seventh, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0106] As data center businesses grow, memory costs account for an increasingly larger proportion of data center server costs. However, the average memory utilization rate in data centers is less than 50%, indicating significant room for improvement in memory application efficiency.
[0107] Existing technologies can enable memory reuse among multiple virtual machines (VMs) within a single server to improve memory utilization efficiency. However, this improvement in memory utilization is limited to within a single server and cannot achieve cross-server (e.g., data center) memory utilization improvement.
[0108] The memory access method provided in this application will be described in detail below with reference to the accompanying drawings.
[0109] It should be understood that the memory access method provided in the embodiments of this application can be applied to computer systems, such as the systems shown in Figures 1 and 2.
[0110] It should also be understood that the embodiments shown below do not particularly limit the specific structure of the execution subject of the method provided in the embodiments of this application. As long as the method provided in the embodiments of this application can be implemented by running a program that records the code of the method provided in the embodiments of this application. For example, the execution subject of the method provided in the embodiments of this application can be a device, or a functional module in the device that can call and execute a program.
[0111] Figure 3 is a schematic flowchart of a memory access method provided in this application. This method is applied to a system comprising a management unit, a second computing unit, and a first computing unit that are communicatively connected to each other; wherein the first computing unit and the second computing unit are units on different servers.
[0112] For example, the second computing unit could be a server, and the first computing unit could be another server;
[0113] For example, the second computing unit can be multiple servers, and the first computing unit can be other servers not included in the multiple servers;
[0114] For example, the second computing unit can be one or more virtual machines within the server, and the first computing unit can be one or more virtual machines within the server.
[0115] In one possible implementation, the second computing unit and the first computing unit (optionally, also including a management unit) are connected via a high-speed interconnect bus.
[0116] A high-speed interconnect bus is a high-speed communication channel used in a computer system to connect processors, memory, storage, peripherals, and other components, aiming to achieve low-latency, high-bandwidth data transmission. High-speed interconnect buses can achieve high bandwidth and low latency transmission. Functionally, they can be classified as system buses (Front-Side Bus, FSB), memory buses, I / O buses, etc. Based on topology, they can be classified as shared buses (such as traditional PCI), point-to-point buses (such as PCIe), and switching buses (such as InfiniBand). This application does not limit the specific type of high-speed interconnect bus.
[0117] This memory access method includes the following steps:
[0118] 301. The first computing unit obtains the physical address of the target memory space on the second computing unit.
[0119] This application uses the example of a first computing unit accessing the memory on a second computing unit to illustrate how to enable a computing unit on one server to access the memory of a computing unit on another server.
[0120] In one possible implementation, the first computing unit can obtain the physical address of the target memory space on the second computing unit. The target memory space is the memory space on the second computing unit that the first computing unit can access.
[0121] The target memory space can be the memory space on the second computing unit that can be allocated to other computing units, i.e., the allocatable memory space. Here, "allocatable memory space" can be understood as the memory space of a computing unit that can be occupied by other computing units. In a kernel-driven implementation (e.g., but not limited to balloon drivers), the memory space occupied by the kernel driver can be considered allocatable memory space. When the memory space occupied by the kernel driver needs to be allocated to other computing units, the kernel driver can release the occupied memory space, and the released memory space can then be occupied and used by other computing units. Specifically, the target memory space can be the memory space occupied by the kernel driver of the second computing unit. Before the second computing unit allocates the target memory space to the first computing unit, the kernel driver of the second computing unit can release the target memory space.
[0122] The following describes how the first computing unit obtains the physical address of the target memory space on the second computing unit:
[0123] In one possible implementation, the second computing unit can send the physical address of the target memory space to the first computing unit. Then, the first computing unit can obtain the physical address of the target memory space sent by the second computing unit. Specifically, the second computing unit can determine the memory space (i.e., the target memory space) that it can allocate to other computing units and inform the first computing unit of the physical address of the target memory space.
[0124] In one possible implementation, the management node can send the physical address of the target memory space to the first computing unit. Then, the first computing unit can obtain the physical address of the target memory space sent by the management unit. Specifically, the second computing unit can determine the memory space (including the target memory space) that it can allocate to other computing units and inform the management unit of the physical address of the target memory space. Then, the management unit can send the physical address of the target memory space to the first computing unit.
[0125] In this embodiment, a server (a management unit within the server) can act as the manager to maintain memory information on other computing units in the system (e.g., the first information in this embodiment). The first information can indicate the size of allocable memory space on each computing unit. When the management unit receives a request from a computing unit to use memory (e.g., a first request from the first computing unit, which may include the amount of memory space required), it can determine, based on the memory space requirement in the request, which (or which computing units) in the system have allocable memory space that meets the memory space requirement in the request (e.g., the second computing unit has memory space that meets the memory requirement indicated in the first request). Then, it can request the second computing unit to allocate memory space that meets the memory requirement for the first computing unit to use. Subsequently, the second computing unit can release the memory space that meets the memory requirement (i.e., the target memory space) and send the address of the target memory space to the first computing unit.
[0126] The following section describes how to construct and maintain the memory information of each computing unit (for example, the memory information of the second computing unit is the first information).
[0127] In one possible implementation, each computing unit can send information about itself and the size of its allocatable memory space to the administrator for maintenance. For example, when the server of a computing unit is powered on, all memory within the server can be provided as allocatable memory space to the administrator. For example, when the allocatable memory space of a computing unit is subsequently occupied by its own business process or the business process of other computing units, the size of the allocated memory space can be provided to the administrator.
[0128] For example, when the server powers on, the second computing unit can occupy all the memory within the second computing unit through a kernel driver (e.g., but not limited to a balloon driver). After the kernel driver on the second computing unit occupies all the memory, it releases it to the management node (that is, it sends the memory space occupied by the balloon driver on the second computing unit to the management unit, and all subsequent allocation of this memory is done by the management node, which is equivalent to releasing it to the management node). The management node can maintain a global memory mapping table, which can contain the size of the allocatable memory space of each computing unit in multiple computing units, including the second computing unit.
[0129] For example, when all or part of the allocatable memory space on the second computing unit is occupied by its own business processes or business processes of other computing units, the size of the allocated memory space can be informed to the management unit.
[0130] For example, after the memory space on the second computing unit that was occupied by its own business process or the business process of other computing units is released, it can be reoccupied by the kernel driver. Then, the second computing unit can inform the management unit of the size of the memory space reoccupied by the kernel driver.
[0131] Referring to Figure 4, server 1 includes a management unit, server 2 and server 3 are child nodes (e.g., the second computing unit and the first computing unit), the kernel driver (e.g., but not limited to the balloon driver) maintains the memory mapping table of its own node and the global memory mapping table, the management unit maintains the memory usage of all physical server nodes in the system, the child nodes maintain their own memory usage, the initial balloon driver requests memory for its own node, marks the balloon driver usage, and uniformly hands it over to the balloon driver of the management node for management.
[0132] In one possible implementation, the second computing unit can inform the management unit of changes in the size of the allocatable memory space through memory update information.
[0133] The memory update information may include the operation identifier (or operation record) of the memory release or occupation operation, such as the page frame number (PFN). In existing implementations, operation records are recorded at the memory page level. For example, for a memory space containing memory page 1, memory page 2, ..., memory page n, releasing it requires n records (i.e., n PFNs, each corresponding to one record), which leads to excessively large memory requirements. Since memory pages in the allocatable memory space are often zero-page memory, in this embodiment, multiple memory pages in the allocatable memory space can be mapped to the same operation record identifier using the same operation record, thereby reducing storage requirements and the difficulty of information maintenance.
[0134] The operation identifier can be mapped to a memory page. Based on the index of the operation identifier, the corresponding memory page can be found and data can be read from it.
[0135] In one possible implementation, the allocatable memory space on the second computing unit indicated in the first information includes a plurality of first memory pages, wherein the first memory pages are zero-page memory; the first information also includes a first identifier of the same operation record corresponding to the plurality of first memory pages; the second information indicates the size of the allocatable memory space on the third computing unit; the allocatable memory space on the third computing unit indicated in the second information includes a plurality of second memory pages, wherein the second memory pages are zero-page memory; the second information also includes a second identifier of the same operation record corresponding to the plurality of second memory pages; the first identifier and the second identifier are the same.
[0136] For example, referring to Figure 5, the second computing unit can inform the management node that the currently allocable memory space has not been allocated, meaning this page of memory is a zero-page memory. After receiving this information, the management node can allocate a page of available memory on any computing unit in the system, mapping all zero-page memory on all servers to this page of memory, i.e., PFN2 in Figure 5. Optionally, when a computing unit requests to use this page of memory, the kernel driver can send a notification to the kernel driver of the management node. The management node will then request a new page of memory for that node and unmap the zero-page memory, remapping it to the newly requested memory page.
[0137] In one possible implementation, the first computing unit is a computing unit on a single server, and the second computing unit is a computing unit on one or more servers. That is, the management unit can allocate the memory resources of multiple servers to the first computing unit, thereby further improving the system's memory utilization.
[0138] Referring to Figure 6, which is a schematic diagram of the system architecture implemented in this application, the management unit is located on the server acting as the manager and maintains global memory mapping information, including the allocatable memory size of each server, such as the allocatable memory size of server 1 and server 2.
[0139] In one possible implementation, the second computing unit itself can also maintain the memory usage of the server.
[0140] 302. The first computing unit constructs a page table based on the physical address; the page table indicates the mapping relationship between the physical address and the virtual address.
[0141] In one possible implementation, if the management unit determines, based on the first information, that there is allocatable memory space on the second computing unit that meets the memory requirement, it may send a third request to the first computing unit. The third request is used to instruct the first computing unit to use the memory space on the second computing unit. Then, after receiving the physical address of the target memory space, the first computing unit actively constructs a page table indicating the mapping relationship between the physical address and the virtual address based on the physical address.
[0142] 303. The first computing unit determines the corresponding target physical address through the page table based on the target virtual address, and accesses the target memory space on the second computing unit based on the target physical address.
[0143] In one possible implementation, the first computing unit (e.g., through the CPU's memory management module) constructs a page table indicating the mapping relationship between the physical address and the virtual address; then, the first computing unit can map the virtual address to the target physical address through the page table, and access the target memory space according to the target physical address.
[0144] The first computing unit can construct a page table indicating the mapping relationship between virtual addresses and physical addresses based on the physical addresses of computing units on other servers. Based on this page table, the first computing unit can directly access the memory space of computing units on other servers through virtual addresses in the same way as local memory access, thereby realizing cross-server memory access. This can improve the memory utilization of a system containing multiple servers. For a single server, since it can use the memory resources of other servers, it can increase the upper limit of the memory resources that the server can use and the upper limit of the access bandwidth.
[0145] In one possible implementation, the second computing unit sends memory update information to the management unit; the memory update information indicates that the allocatable memory space on the second computing unit has been reduced by the memory size; the management unit updates the allocatable memory space size on the second computing unit as indicated in the first information according to the memory update information.
[0146] For example, referring to Figure 7, in the upper mapping relationship of Figure 7, if the newly received memory update information indicates that the allocatable memory size of server 1 has decreased by c, then the allocatable memory size of server 1 can be directly updated in the record corresponding to PFN2 (updated from a to ac).
[0147] In one possible implementation, the second computing unit sends memory update information to the management unit; the memory update information indicates that the allocatable memory space of the memory size has been reduced on the second computing unit; the management unit may store third information in addition to the first information, the third information indicating that the allocatable memory space of the memory size has been reduced on the second computing unit.
[0148] For example, referring to Figure 7, in the next mapping relationship in Figure 7, if the newly received memory update information indicates that the allocatable memory size of server 1 has decreased by c, then an operation record can be added, that is, an operation record corresponding to PFN4 can be added, which indicates that the allocatable memory size of server 1 has decreased by c.
[0149] For a computing unit, allocating too much of its memory space to other computing units can lead to excessive pressure on memory access bandwidth (memory access bandwidth refers to the rate at which a memory module transfers data per unit of time; excessive access bandwidth pressure essentially means approaching the maximum access bandwidth of the computing unit). Therefore, when allocating memory, the management unit can also consider the access bandwidth of each computing unit. When the access bandwidth meets preset conditions (e.g., when the access bandwidth is less than a preset value), memory can be allocated from that computing unit for use by other computing units.
[0150] When a computing unit is allocated a large amount of memory space, the memory access bandwidth is often large. Even if there is still remaining allocable memory space, it is not suitable to allocate the remaining allocable memory space due to the limitation of memory access bandwidth. Instead, some computing units with remaining allocable memory space and lower memory access bandwidth can be selected for memory space allocation. This can achieve the balance of memory access bandwidth in the system, which is equivalent to expanding the memory bandwidth.
[0151] Referring to Figure 8, which is a schematic flowchart of a memory access method according to an embodiment of this application, including:
[0152] 8011. The balloon driver of the first computing unit occupies memory;
[0153] 8012. The balloon driver of the second computing unit occupies memory;
[0154] 8021. The first computing unit sends memory update information to the management unit; the memory update information indicates the amount of memory occupied by the kernel driver on the first computing unit;
[0155] 8022. The second computing unit sends memory update information to the management unit; the memory update information indicates the amount of memory occupied by the kernel driver on the second computing unit;
[0156] 803. The management unit constructs global memory mapping information (including the first information) based on memory update information;
[0157] 804. The first computing unit sends a first request to the management unit; the first request indicates the amount of memory to be used;
[0158] 8051. Based on the first information, the management unit determines that there is allocatable memory space on the first computing unit that meets the memory requirement;
[0159] 8052. The management unit sends a second request to the second computing unit; the second request is used to instruct the second computing unit to allocate memory space for the specified amount of memory.
[0160] 806. The balloon driver of the second computing unit releases memory;
[0161] 8071. The second computing unit sends the address of the target memory space to the first computing unit, wherein the target memory space is the memory space of the specified amount of memory;
[0162] 8072. The second computing unit sends memory update information to the management unit; the memory update information indicates that the allocatable memory space of the memory size has been reduced on the second computing unit;
[0163] 808. The management unit updates the global memory mapping information based on the memory update information;
[0164] 809. The first computing unit constructs a page table indicating the mapping relationship between the physical address and the virtual address;
[0165] 8010. The first computing unit determines the corresponding target physical address through the page table based on the target virtual address, and accesses the target memory space based on the target physical address.
[0166] It should be understood that the specific examples shown in Figures 3 to 8 of the embodiments of this application are only for the purpose of helping those skilled in the art to better understand the embodiments of this application, and are not intended to limit the scope of the embodiments of this application. It should also be understood that the sequence number of each process does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0167] It should also be understood that, in the various embodiments of this application, unless otherwise specified or in case of logical conflict, the terms and / or descriptions between different embodiments are consistent and can be referenced by each other, and the technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationships.
[0168] The foregoing mainly describes the solutions provided by the embodiments of this application from a methodological perspective. To achieve the above functions, it includes corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should readily recognize that, based on the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0169] The data processing apparatus provided in the embodiments of this application will now be described in detail with reference to Figures 9 and 10. It should be understood that the description of the apparatus embodiments corresponds to the description of the method embodiments. Therefore, for content not described in detail, please refer to the method embodiments above. For the sake of brevity, some content will not be repeated.
[0170] Referring to Figure 9, which is a schematic diagram of the structure of a computing unit provided in an embodiment of this application, the computing unit 900 can be the first computing unit in the above embodiment, and the first computing unit and the second computing unit are communicatively connected to each other; wherein the first computing unit and the second computing unit are units on different servers, including: a transceiver module 901, used to obtain the physical address of the target memory space on the second computing unit;
[0171] The transceiver module 901 can be the interface circuit of the server's processor (e.g., CPU). For a detailed description of the transceiver module 901, please refer to the description of step 301 in the above embodiment. The similarities will not be repeated here.
[0172] Page table construction module 902 is used to construct a page table based on the physical address; the page table indicates the mapping relationship between the physical address and the virtual address; memory access module 903 is used to determine the corresponding target physical address based on the target virtual address through the page table, and to access the target memory space on the second computing unit based on the target physical address.
[0173] The page table construction module 902 and the memory access module 903 can be the memory management unit of the server's processor (e.g., CPU). For a detailed description of the page table construction module 902 and the memory access module 903, please refer to the description of steps 302 and 303 in the above embodiments. The similarities will not be repeated here.
[0174] In one possible implementation, the second computing unit and the first computing unit are connected via a high-speed interconnect bus.
[0175] In one possible implementation, the first computing unit is a computing unit on a server, and the second computing unit is a computing unit on one or more servers.
[0176] In one possible implementation, the transceiver module is specifically used for:
[0177] The physical address of the target memory space is received from the second computing unit or the management unit.
[0178] This application provides a chip system 800, as shown in FIG10. The chip system 800 includes at least one processor and at least one interface circuit. As an example, when the chip system 800 includes a processor and an interface circuit, the processor may be the processor 810 shown in the solid box in FIG10 (or the processor 810 shown in the dashed box), and the interface circuit may be the receiving circuit 820 shown in the solid box in FIG10 (or the receiving circuit 820 shown in the dashed box).
[0179] When the chip system 800 includes two processors and two interface circuits, the two processors include processor 810 shown in the solid box and processor 810 shown in the dashed box in Figure 10, and the two interface circuits include receiving circuit 820 shown in the solid box and receiving circuit 820 shown in the dashed box in Figure 10. This is not limited. Processor 810 and receiving circuit 820 can be interconnected via lines. For example, receiving circuit 820 can be used to receive signals (e.g., instructions stored in memory). As another example, receiving circuit 820 can be used to send signals to other devices (e.g., processor 810).
[0180] For example, the receiving circuit 820 can read instructions stored in the memory and send the instructions to the processor 810. When the instruction is executed by the processor 810, it can cause a data processing device or a device accessing memory to perform the steps in the above embodiments. Of course, the chip system 800 may also include other discrete devices, and this application embodiment does not specifically limit this.
[0181] Another embodiment of this application provides a computer-readable storage medium storing instructions that, when executed on a data processing apparatus, perform the various steps of the method flow shown in the above-described method embodiments. In some embodiments, the disclosed method can be implemented as computer program instructions encoded in a machine-readable format on a computer-readable storage medium or on other non-transitory media or articles of art.
[0182] Figure 11 schematically illustrates a conceptual partial view of a computer program product provided in an embodiment of this application, the computer program product including a computer program for executing computer processes on a computing device.
[0183] In one embodiment, a computer program product is provided using a signal-bearing medium 1100. The signal-bearing medium 1100 may include one or more program instructions that, when executed by one or more processors, provide the functions or portions thereof described above with reference to FIG3. Thus, for example, one or more features of 301-303 in FIG3 may be embodied by one or more instructions associated with the signal-bearing medium 1100. Furthermore, example instructions are also described in FIG11.
[0184] In some examples, the signal carrying medium 1100 may include a computer-readable medium 1101, such as, but not limited to, a hard disk drive, a compact disc (CD), a digital video disc (DVD), a digital magnetic tape, a memory, a read-only memory (ROM), or a random access memory (RAM), etc.
[0185] In some implementations, the signal carrying medium 1100 may include a computer recordable medium 1102, such as, but not limited to, a memory, a read / write (R / W) CD, a R / W DVD, and so on.
[0186] In some implementations, the signal-bearing medium 1100 may include a communication medium 1103, such as, but not limited to, digital and / or analog communication media (e.g., fiber optic cables, waveguides, wired communication links, wireless communication links, etc.). The signal-bearing medium 1100 may be transmitted by a wireless communication medium 1103 (e.g., a wireless communication medium conforming to the IEEE 1502.11 standard or other transmission protocols). One or more program instructions may be, for example, computer-executable instructions or logical implementation instructions.
[0187] In some examples, such as the memory access method for Figure 3, it can be configured to provide various operations, functions, or actions in response to one or more program instructions in a computer-readable medium 1101, a computer-recordable medium 1102, and / or a communication medium 1103.
[0188] It should be understood that the arrangements described herein are for illustrative purposes only. Therefore, those skilled in the art will understand that other arrangements and other elements (e.g., machines, interfaces, functions, sequences, and functional groups, etc.) can be used instead, and some elements may be omitted depending on the desired outcome. Furthermore, many of the described elements are functional entities that can be implemented as discrete or distributed components, or in any suitable combination and location with other components.
[0189] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software programs, it can be implemented, in whole or in part, in the form of a computer program product. This computer program product includes one or more computer instructions. When executed on a computer and when the computer execution instructions are executed, all or part of the processes or functions according to the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device.
[0190] Computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. A computer-readable storage medium can be any available medium that a computer can access, or it can include one or more data storage devices such as servers or data centers that can be integrated with media. Available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).
[0191] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this invention should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A system, characterized in that, The system includes a first computing unit and a second computing unit that are interconnected; wherein the first computing unit and the second computing unit are units on different servers; The first computing unit is used to obtain the physical address of the target memory space on the second computing unit; The first computing unit is configured to construct a page table based on the physical address; the page table indicates the mapping relationship between the physical address and the virtual address; The first computing unit is configured to determine the corresponding target physical address based on the target virtual address through the page table, and to access the target memory space on the second computing unit based on the target physical address.
2. The system according to claim 1, characterized in that, The second computing unit and the first computing unit are connected via a high-speed interconnect bus.
3. The system according to claim 1 or 2, characterized in that, The first computing unit is a computing unit on a single server, and the second computing unit is a computing unit on multiple servers.
4. The system according to any one of claims 1 to 3, characterized in that, The first computing unit is specifically used for: The physical address of the target memory space is received from the second computing unit or the management unit.
5. The system according to any one of claims 1 to 4, characterized in that, The system also includes: a management unit; The management unit is used to acquire first information; the first information indicates the size of the allocatable memory space on the second computing unit. The first computing unit is further configured to send a first request to the management unit, the first request indicating the amount of memory to be used; The management unit is further configured to send a second request to the second computing unit when it is determined, based on the first information, that there is allocatable memory space on the second computing unit that meets the memory requirement. The second request is used to instruct the second computing unit to allocate memory space of the specified memory amount; wherein the target memory space is memory space of the specified memory amount.
6. The system according to claim 5, characterized in that, The management unit is further configured to send a third request to the first computing unit when it is determined, based on the first information, that there is allocatable memory space on the second computing unit that meets the memory requirement. The third request is used to instruct the first computing unit to use the memory space on the second computing unit.
7. The system according to any one of claims 1 to 6, characterized in that, The second computing unit is further configured to send memory update information to the management unit; the memory update information indicates that the allocatable memory space on the second computing unit has been reduced by the size of the memory. The management unit updates the size of the allocatable memory space on the second computing unit as indicated in the first information, based on the memory update information; or, The management unit may store third information in addition to the first information, the third information indicating that the allocatable memory space on the second computing unit has been reduced to the size of the target memory space.
8. The system according to any one of claims 5 to 7, characterized in that, The memory space available for allocation on the second computing unit indicated in the first information includes multiple first memory pages, where each first memory page is a zero-page memory; the first information also includes a first identifier of the same operation record corresponding to the multiple first memory pages; The system also includes: The management unit acquires second information, which indicates the size of the allocatable memory space on the third computing unit; the allocatable memory space on the third computing unit indicated in the second information includes multiple second memory pages, each of which is a zero-page memory; the second information also includes a second identifier of the same operation record corresponding to the multiple second memory pages; the first identifier and the second identifier are the same.
9. The system according to claim 8, characterized in that, The first identifier and the second identifier are page frame numbers (PFNs).
10. The system according to any one of claims 5 to 9, characterized in that, The allocatable memory space on the second computing unit is the space occupied by the balloon driver on the second computing unit; The second computing unit is further configured to, upon receiving the second request, release the amount of memory space from the allocable memory space via a balloon driver on the second computing unit.
11. The system according to any one of claims 5 to 10, characterized in that, The first information is also used to indicate the access bandwidth on the second computing unit; the management unit is specifically used to: send a second request to the second computing unit when it is determined, based on the first information, that there is allocable memory space on the second computing unit that meets the memory requirement and the access bandwidth on the second computing unit meets the preset conditions.
12. A memory access method, characterized in that, The method is applied to a first computing unit, wherein the first computing unit and the second computing unit are communicatively connected; wherein the first computing unit and the second computing unit are units on different servers; the method includes: Obtain the physical address of the target memory space on the second computing unit; A page table is constructed based on the physical address; the page table indicates the mapping relationship between the physical address and the virtual address. Based on the target virtual address, the corresponding target physical address is determined through the page table, and the target memory space on the second computing unit is accessed based on the target physical address.
13. The method according to claim 12, characterized in that, The second computing unit and the first computing unit are connected via a high-speed interconnect bus.
14. The method according to claim 12 or 13, characterized in that, The first computing unit is a computing unit on a single server, and the second computing unit is a computing unit on multiple servers.
15. The method according to any one of claims 12 to 14, characterized in that, The first computing unit obtains the physical address of the target memory space on the second computing unit, including: The physical address of the target memory space is received from the second computing unit or the management unit.
16. A computing unit, characterized in that, The computing unit is a first computing unit, and the first computing unit and the second computing unit are communicatively connected to each other; wherein the first computing unit and the second computing unit are units on different servers, including: The transceiver module is used to obtain the physical address of the target memory space on the second computing unit; A page table construction module is used to construct a page table based on the physical address; the page table indicates the mapping relationship between the physical address and the virtual address. The memory access module is used to determine the corresponding target physical address based on the target virtual address through the page table, and to access the target memory space on the second computing unit based on the target physical address.
17. The computing unit according to claim 16, characterized in that, The second computing unit and the first computing unit are connected via a high-speed interconnect bus.
18. The computing unit according to claim 16 or 17, characterized in that, The first computing unit is a computing unit on a single server, and the second computing unit is a computing unit on multiple servers.
19. The computing unit according to any one of claims 16 to 18, characterized in that, The transceiver module is specifically used for: The physical address of the target memory space is received from the second computing unit or the management unit.
20. A server, characterized in that, Including processor and memory; The processor is configured to execute instructions stored in the memory to cause the computing device to perform the method as described in any one of claims 12 to 15.
21. A computer program product containing instructions, characterized in that, When the instructions are executed by the computing device, the computing device performs the method as described in any one of claims 12 to 15.
22. A computer-readable storage medium, characterized in that, It includes computer program instructions, which, when executed by a computing device, cause the computing device to perform the method as described in any one of claims 12 to 15.
23. A chip, characterized in that, It includes at least one processing unit and an interface circuit, the interface circuit being used to provide program instructions or data to the at least one processing unit, the at least one processing unit being used to execute the program instructions to implement the method of any one of claims 12 to 15.