Video memory management method and device, electronic equipment and storage medium

By identifying and replacing hot and cold pages in the GPU page table, memory management is optimized, solving the problems of limited memory resources and low utilization, and improving GPU performance and system stability.

CN120994305APending Publication Date: 2025-11-21MOORE THREADS TECHNOLOGY (SHANGHAI) CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511095027.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing GPU virtualization technologies suffer from limited memory resources, low utilization, and a lack of hot and cold page identification and replacement mechanisms, leading to a decline in GPU performance.

Method used

By identifying cold and hot pages in the GPU page table, cold pages are swapped into system memory and hot pages are swapped into GPU memory. The system optimizes memory usage by periodically scanning threads and manages memory resources by setting margin thresholds and memory page pools.

Benefits of technology

This improves the utilization of video memory and reduces the need to use system memory as a substitute due to insufficient video memory, thereby enhancing the overall performance of the GPU and the stability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120994305A_ABST
    Figure CN120994305A_ABST
Patent Text Reader

Abstract

The invention relates to a video memory management method and device, electronic equipment and a storage medium, the method is applied to a GPU drive in a virtual machine, and the method comprises the steps that when it is detected that the remaining amount of a GPU video memory is lower than a set remaining amount threshold value, cold pages in a GPU page table are recognized; wherein the cold page is a page table item which is not accessed in a recent preset time interval; and under the condition that the cold page is a video memory, replacing a physical address corresponding to the virtual address in the cold page to a physical address in a system memory so as to replace video memory data corresponding to the cold page to the system memory. The embodiment of the invention can improve the overall utilization rate of the video memory.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of computer, and particularly relates to a GPU memory management method and device, electronic equipment and storage medium. BACKGROUND

[0002] With the rapid development of information technology, GPU virtualization technology is increasingly widely used in virtual display scenarios. For example, cloud desktop technology VDI, as a technology of deploying desktop environment on a cloud server, can effectively improve resource utilization, reduce operation and maintenance costs, and provide higher data security. In the VDI environment, multiple virtual machines share the same physical GPU resources, and through GPU virtualization technology, a virtual GPU (vGPU) can be allocated to each virtual machine, thereby realizing flexible scheduling and management of GPU resources.

[0003] However, the existing GPU virtualization technology has some deficiencies in GPU memory management. In the virtual display scenario, the vGPU device memory resources used by the virtual machine are usually limited, such as 512MB, 1GB or 2GB, etc. When the memory resources are tight, the system can only use the system memory to replace the GPU memory, which will cause the GPU performance to decline, because the performance of the GPU accessing the system memory is much lower than that of accessing the GPU memory. In addition, the existing GPU memory management method lacks effective cold and hot page identification and replacement mechanism, which cannot fully utilize the GPU memory resources, resulting in low GPU memory utilization and affecting the overall system performance. SUMMARY

[0004] The present disclosure provides a GPU memory management technical solution.

[0005] According to an aspect of the present disclosure, a GPU memory management method is provided, applied to a GPU driver in a virtual machine, comprising:

[0006] When it is detected that the remaining amount of GPU memory is lower than a set amount threshold, a cold page in a GPU page table is identified; wherein the cold page is a page table entry that has not been accessed in a recent preset time interval;

[0007] In the case that the cold page is GPU memory, the physical address corresponding to the virtual address in the cold page is replaced to a physical address in the system memory, so as to replace the GPU memory data corresponding to the cold page to the system memory.

[0008] In a possible implementation manner, the method further comprises:

[0009] A hot page in the GPU page table is identified; wherein the hot page is a page table entry that has been accessed in a recent preset time interval;

[0010] The physical address corresponding to the virtual address in the hot page is replaced by a physical address in GPU display memory, so as to replace the memory data corresponding to the hot page into the GPU display memory.

[0011] In a possible implementation, the replacing the physical address corresponding to the virtual address in the hot page by a physical address in GPU display memory comprises:

[0012] allocating display memory from a display memory idle pool;

[0013] modifying the physical address in the hot page to the newly allocated display memory address, and resetting the access flag in the hot page to a first identifier; the first identifier is used to indicate that the GPU page table is a cold page.

[0014] In a possible implementation, after the physical address in the hot page is modified to the newly allocated display memory address, the method further comprises:

[0015] saving the original system memory page to a preset system memory page pool; the original system memory page is a system memory page originally storing the memory data corresponding to the hot page before the memory data corresponding to the hot page is replaced into the GPU display memory;

[0016] when the capacity of the system memory page pool exceeds a set capacity threshold, releasing the excess memory pages.

[0017] In a possible implementation, the replacing the physical address corresponding to the virtual address in the cold page by a physical address in system memory comprises:

[0018] obtaining an idle system memory address from a preset system memory page pool;

[0019] modifying the physical address corresponding to the virtual address in the cold page to the obtained system memory address, and maintaining the access flag in the cold page as the first identifier;

[0020] releasing the display memory corresponding to the cold page.

[0021] In a possible implementation, the method further comprises:

[0022] when the GPU accesses the GPU page table, setting the access flag in the corresponding page table entry to a second identifier; the second identifier is used to indicate that the GPU page table is a hot page.

[0023] In a possible implementation, the identifying the cold page in the GPU page table comprises:

[0024] preferentially traversing the GPU page table of a process that is not active in a preset time interval, to identify the cold page in the GPU page table.

[0025] In a possible implementation, the method further includes:

[0026] In response to a GPU driver loading instruction of the virtual machine, initializing the GPU driver;

[0027] In a case of successful initialization, creating a periodically running scanning thread to perform a cold page and hot page identification operation.

[0028] According to an aspect of the present disclosure, a GPU driver management apparatus is provided, which is applied to a GPU driver in a virtual machine, and includes:

[0029] A cold page identification module, configured to identify a cold page in a GPU page table when it is detected that a GPU video memory remaining amount is lower than a set remaining amount threshold; wherein the cold page is a page table entry that has not been accessed in a preset time interval.

[0030] A replacement module, configured to, in a case that the cold page is a video memory, replace a physical address corresponding to a virtual address in the cold page to a physical address in a system memory, so as to replace video memory data corresponding to the cold page to the system memory.

[0031] According to an aspect of the present disclosure, an electronic device is provided, which includes a processor, and a memory for storing processor-executable instructions; wherein the processor is configured to invoke the instructions stored in the memory to perform the above method.

[0032] According to an aspect of the present disclosure, a computer readable storage medium is provided, which stores computer program instructions, and the computer program instructions are executed by a processor to implement the above method.

[0033] In the embodiments of the present disclosure, during the running of a virtual machine, when video memory resources are in shortage, a cold page is identified and replaced in time, so as to release video memory space. In this way, the originally limited video memory can be more efficiently utilized, and the case of using system memory instead due to insufficient video memory is reduced, thereby improving the overall utilization rate of the video memory.

[0034] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, but not limiting the present disclosure. Other features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS

[0035] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the present disclosure and, together with the specification, serve to explain the technical solutions of the present disclosure.

[0036] Figure 1 A flowchart of a GPU driver loading process according to an embodiment of the present disclosure is shown.

[0037] Figure 2 A flow chart of a method for managing a display memory according to an embodiment of the present disclosure is shown.

[0038] Figure 3 A specific application scenario of an embodiment of the present disclosure is shown.

[0039] Figure 4 A block diagram of a device for managing a display memory according to an embodiment of the present disclosure is shown.

[0040] Figure 5 A block diagram of an electronic device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0041] Various exemplary embodiments, features and aspects of the present disclosure will be explained in detail below with reference to the accompanying drawings. The same reference numerals in the drawings denote the same or similar elements. Although various aspects of the embodiments are illustrated in the drawings, the drawings are not necessarily drawn to scale unless specifically indicated.

[0042] The word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any implementation described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other implementations.

[0043] The term "and / or", merely describes association relationship of associated objects, and means that three relationships can exist, for example, A and / or B can mean that A exists alone, A and B exist together, and B exists alone. In addition, the term "at least one" herein means any one of multiple or any combination of at least two of multiple, for example, including at least one of A, B and C can mean including any one or more elements selected from the set consisting of A, B and C.

[0044] In addition, in order to better illustrate the present disclosure, numerous specific details are given in the specific embodiments below. Those skilled in the art should understand that the present disclosure can also be implemented without some specific details. In some examples, methods, means, elements and circuits well known to those skilled in the art are not described in detail, in order to highlight the main idea of the present disclosure.

[0045] The present disclosure provides a method for managing a display memory applied to GPU (Graphics Processing Unit) driver in a virtual machine. In the context of the increasing popularity of cloud computing and virtualization technology, GPU driver in a virtual machine plays a crucial role. It not only affects the graphics performance of the virtual machine, but also largely determines the utilization efficiency of GPU resources.

[0046] When the host machine starts, the GPU driver performs a series of initialization operations, including checking and configuring the GPU hardware, preparing for the operation of the virtual machine. For example, the driver detects the hardware version and functional characteristics of the GPU, configures the register settings of the GPU, and pre-allocates resources for subsequent virtual machine startup and graphics rendering. The host machine provides a unified framework to manage multiple virtual machines, and the GPU driver is part of this framework, coordinating the allocation and use of GPU resources, and allocating GPU resources to different virtual machines.

[0047] The GPU driver in the virtual machine is responsible for managing GPU hardware resources, including allocating and releasing video memory, managing GPU computing units and memory bandwidth, etc.

[0048] When the virtual machine starts, the GPU driver allocates a certain amount of video memory from the host machine or GPU hardware to the virtual machine. The size of the video memory depends on the configuration of the virtual machine and the capabilities of the GPU hardware. For example, for a graphics-intensive virtual machine, it may need to allocate 2GB or more of video memory.

[0049] In one possible implementation, the GPU driver in the virtual machine initializes the GPU driver in response to the virtual machine's GPU driver loading instruction; in the case of successful initialization, a periodic running scan thread is created to perform cold page and hot page identification operations.

[0050] When the virtual machine starts, the corresponding GPU driver program is loaded. This process usually involves hardware initialization and software configuration to ensure that the GPU can interact normally with the virtual machine operating system and other software. The loading instruction can be triggered automatically by the operating system or manually started by the user.

[0051] Figure 1 A flowchart showing the GPU driver loading process according to an embodiment of the present disclosure. As shown in Figure 1 The GPU driver program starts the initialization process after receiving the loading instruction. This includes checking whether the GPU hardware is available, allocating video memory resources, configuring the working mode of the GPU, etc. For example, the GPU driver may communicate with the GPU hardware to set the size and allocation strategy of the video memory, and initialize the computing unit and memory bandwidth of the GPU.

[0052] After the GPU driver initialization succeeds, a periodic scan thread can be created. The thread is used to scan the GPU page table, identify cold pages and hot pages, and further perform the replacement work of cold and hot pages. For details, refer to the memory management method provided in the present disclosure. The cold page refers to a page that has not been accessed in a recent period of time, and the hot page refers to a page that has been frequently accessed in a recent period of time. By identifying the cold and hot pages, the GPU driver can optimize the use of the video memory, replace the cold pages to the system memory, and retain the hot pages in the video memory, thereby improving the utilization of the video memory and the performance of the GPU.

[0053] The scan thread is periodically run at a preset time interval. For example, the thread can be run once every 5 seconds to scan the GPU page table and check the access state of each page. In this way, changes in video memory resources can be discovered in a timely manner, and the allocation of video memory and memory can be adjusted in a timely manner.

[0054] The specific identification of cold pages and hot pages and the corresponding replacement process are described below.

[0055] Figure 2 A flowchart of a memory management method according to an embodiment of the present disclosure is shown. The method is applied to a GPU driver in a virtual machine, as shown in Figure 2 The method includes the following steps.

[0056] In step S11, when it is detected that the GPU video memory remaining amount is lower than a set amount threshold, a cold page in the GPU page table is identified; wherein the cold page is a page table entry that has not been accessed in a preset time interval.

[0057] The GPU video memory remaining amount is the part of the video memory allocated to the current virtual machine that has not been occupied for storing data. The GPU video memory is used to temporarily store information required by the GPU for graphics operation or data processing, such as pixel data, vertex data, and texture data of an image. For example, for a certain virtual machine, a total of 2 GB of video memory is allocated. When the virtual machine runs multiple graphics-intensive applications, these applications start to request video memory resources from the GPU to store their data. Assuming that 1.6 GB of video memory has been allocated to store the graphics data of each application at this time, the GPU video memory remaining amount is 2 GB-1.6 GB=0.4 GB. This 0.4 GB of video memory space has not been allocated and can be used for subsequent possible graphics tasks or data storage tasks.

[0058] The set amount threshold is a reference value for measuring whether the GPU video memory remaining amount is sufficient, which is set in advance. When the GPU video memory remaining amount is lower than this value, the cold pages in the video memory can be identified and replaced to the system memory. The amount threshold can be set according to past running data and test experience, which is not limited in the present disclosure. For example, the threshold can be 20% of the total video memory amount of the virtual machine.

[0059] When the remaining GPU memory is below the set threshold, cold pages in the GPU page table can be identified. A cold page is a GPU page table entry that has not been accessed within a recent preset time interval. Cold pages are relative to hot pages, which are frequently accessed page table entries. The identification of cold pages helps optimize the use of memory resources. By replacing the memory data of cold pages, memory space can be released.

[0060] The specific definition of a cold page can be determined according to the actual application scenario and requirements. Generally, the access flag in the page table entry can be used to determine whether a page is a cold page. If the access flag in the page table entry has not been set to an accessed state within a recent period of time, the page is considered a cold page.

[0061] In the GPU page table, the data within the virtual address range corresponding to a page table entry has not been accessed by the GPU within the past 10 seconds. According to the setting of the access flag, the access flag of this page table entry is 0 (not accessed), so the page is identified as a cold page. This indicates that within the past 10 seconds, the GPU has not performed read or write operations on the data of this page, so the memory data of the page can be replaced into the system memory to free up memory space.

[0062] In step S12, in the case where the cold page is in the memory, the physical address corresponding to the virtual address in the cold page is replaced to a physical address in the system memory to replace the memory data corresponding to the cold page into the system memory.

[0063] The GPU page table is a data structure used to implement the mapping of GPU virtual addresses to physical addresses. In a GPU system, the GPU page table is composed of a series of page table entries, each of which records the mapping relationship between a GPU virtual address range and the corresponding physical address, as well as related access control information. Through the GPU page table, the GPU can convert the virtual addresses used in the program to actual physical addresses, thereby accessing data in the memory.

[0064] Suppose the data of a cold page in the memory is located at the physical address 0x60000000 in the memory. When this cold page is replaced into the system memory, this data will be copied to an area in the system memory, such as physical address 0x70000000. At the same time, in the GPU page table, the physical address of this page table entry will be updated to 0x70000000, indicating that the data of this page is now stored in the system memory. In this way, when the GPU needs to access the data of this page, it can find the physical address in the system memory through the page table to perform data access.

[0065] In the embodiments of the present disclosure, when the video memory resource is in shortage during the running of the virtual machine, the cold page is identified and replaced in time to release the video memory space. In this way, the originally limited video memory can be more efficiently utilized, the case of using the system memory instead due to the shortage of the video memory is reduced, and the overall utilization rate of the video memory is improved.

[0066] In a possible implementation, the method further includes: when the GPU accesses the GPU page table, setting an access flag in the corresponding page table entry to a second identifier (reserving the access bit as 0); and the second identifier is used to indicate that the GPU page table is a hot page.

[0067] The access flag is a state flag bit in the GPU page table, and an initial value of the access flag is a first identifier (for example, 0), indicating that the page in the page table entry (that is, the memory page) has not been accessed in a recent time; when the GPU accesses the data in the memory through the page table, the access flag of the corresponding page table entry is set to a second identifier (for example, 1) by the MMU. The flag is used to indicate that the page in the page table entry (that is, the memory page) has been accessed in a recent time.

[0068] In a possible implementation, the method further includes: obtaining an idle system memory address from a preset system memory page pool (if there is no idle page in the memory page pool, a new page is dynamically allocated from the system memory); modifying the physical address corresponding to the virtual address in the cold page to the obtained system memory address, and maintaining the access flag in the cold page as the first identifier (reserving the access bit as 0); and releasing the video memory corresponding to the cold page.

[0069] When the physical address corresponding to the virtual address in the cold page is replaced by the physical address in the system memory, the GPU driver first checks whether there is an idle memory page in the system memory page pool. If there is an idle page, a usable memory page is directly selected from the idle page, and the physical address of the memory page is used to replace the physical address of the cold page. For example, if there is an idle page with a physical address of 0xB0000000 in the page pool, the address is used as a candidate target.

[0070] If there is no idle page in the system memory page pool, the GPU driver dynamically allocates a new memory page from the system memory. For example, by calling a memory allocation function of the system, a part of the physical address space that is not used in the system memory is divided to be used as a new system memory page, and the physical address of the new system memory page is recorded for subsequent use.

[0071] In the GPU driver, a page table management module can be used to maintain the mapping relationship between the virtual address and the physical address. After the cold page to be replaced is determined, the module can find the corresponding page table entry. Then the physical address field in the corresponding page table entry is modified, replacing the original video memory physical address with the newly obtained system memory physical address. For example, the physical address in the original page table entry is 0xC0000000 in the video memory, which is now modified to 0xB0000000 in the system memory.

[0072] The cold page data is then copied from the video memory physical address to the system memory physical address. Direct memory access (DMA) operations can be used to quickly transfer the data block in the video memory to the target address in the system memory.

[0073] The access bit is used to indicate whether the page has been accessed in the recent time. For a cold page, the access bit has been set to the first identification (e.g., access bit is 0), indicating that the page has not been accessed in the recent time. Thus, by keeping the access bit as the first identification, the subsequent scanning thread can still accurately identify the page as a cold page. For example, when the scanning thread scans the page table entry again, the state of the access bit can be used to quickly determine whether the page has been accessed, so as to decide whether to maintain the current state or perform the replacement operation again.

[0074] When the data of the cold page is successfully migrated to the system memory, the GPU driver can release the original physical address space in the video memory. Specifically, the video memory management module can be used to mark the corresponding part in the video memory as available, for example, marking the video memory physical addresses 0xC0000000 to 0xC000FFFF as unallocated, so that the subsequent other tasks can reuse this part of the video memory, thereby improving the utilization of the video memory.

[0075] In one possible implementation, the method further includes: identifying a hot page in the GPU page table; wherein the hot page is a page table entry that has been accessed in the recent preset time interval; and replacing the physical address corresponding to the virtual address in the hot page to a physical address in the GPU video memory to replace the memory data corresponding to the hot page to the GPU video memory.

[0076] The GPU driver periodically scans the GPU page table to identify the hot page in the GPU page table. For example, the access bit of each page table entry can be checked. If the access bit of a page table entry is 1, it indicates that the page has been accessed in the recent preset time interval, and thus it can be identified as a hot page. By identifying the hot page, the GPU driver can understand which page table entries have data that is currently active, important, and needs to be kept in the video memory for fast access, thereby improving the performance of the GPU and the utilization of the video memory.

[0077] When the hot page is identified, the GPU driver program replaces the physical address corresponding to the virtual address in the hot page to the physical address in the GPU display memory.

[0078] In the embodiments of the present disclosure, considering that the access speed of the GPU display memory is much higher than that of the system memory, replacing the data of the hot page to the display memory can significantly improve the efficiency of the GPU accessing the data. Therefore, when the hot page in the GPU page table is identified and the page is located in the system memory, the hot page is replaced to the display memory, so that the GPU can quickly access the frequently used data.

[0079] In a possible implementation, the replacing the physical address corresponding to the virtual address in the hot page to the physical address in the GPU display memory comprises: allocating display memory from a display memory idle pool; modifying the physical address in the hot page to the newly allocated display memory address, and resetting the access flag in the hot page to a first identifier; the first identifier is used to indicate that the GPU page table is a cold page.

[0080] In this implementation, a page of display memory can be allocated from the available display memory. Then the page table entry of the hot page is changed, and the physical address of the display memory is updated to the page table entry of the hot page of the GPU. That is, the physical address field in the page table entry is updated to point to the physical address of the new area in the display memory.

[0081] In addition, the access flag in the hot page can also be reset to the first identifier, for example, the access bit is updated to 0. This operation is used to reset the access flag to 0, that is, if the access flag is always 0, the cold page will be identified in the subsequent operation, and the cold page can be replaced to the system memory, and if it becomes 1 again, the subsequent operation will not be performed, that is, the hot page is still retained in the display memory.

[0082] In the embodiments of the present disclosure, after the hot page data is replaced to the display memory, the delay of the GPU accessing the data is significantly reduced, thereby improving the execution efficiency of the graphic rendering and computing tasks. By resetting the access flag in the hot page to the first identifier, the state of the hot page and the cold page is dynamically identified, and the GPU driver program can identify the cold page and the hot page in the GPU page table based on the access flag bit, so as to more efficiently utilize the display memory and system memory resources, ensure that the display memory is mainly used to store frequently accessed data, and at the same time, store the infrequently used data in the system memory, thereby realizing the optimal configuration of the resources.

[0083] In a possible implementation, after the physical address in the hot page is modified to the newly allocated GPU memory address, the method further includes: saving the original system memory page to a preset system memory page pool, the original system memory page being a system memory page originally storing memory data corresponding to the hot page before the memory data corresponding to the hot page is replaced to the GPU memory; and releasing excess memory pages when a capacity of the system memory page pool exceeds a set capacity threshold.

[0084] In this implementation, after the physical address in the hot page is modified to the newly allocated GPU memory address, the original system memory page is saved to the preset system memory page pool, the original system memory page being a system memory page originally storing memory data corresponding to the hot page before the memory data corresponding to the hot page is replaced to the GPU memory. The system memory page pool is a preset memory area for saving the system memory pages replaced from the GPU memory. These memory pages are temporarily stored in the system memory page pool after being replaced from the GPU memory, for possible subsequent reuse.

[0085] By saving the original system memory page to the system memory page pool, frequent allocation and release of memory pages from the system memory can be avoided, and the efficiency of memory management can be improved. Meanwhile, the system memory page pool can also serve as a buffer area, to ensure that sufficient memory pages can be quickly replaced when GPU memory resources are in shortage.

[0086] The capacity threshold of the system memory page pool is a preset upper limit value for controlling the size of the system memory page pool. When the capacity of the system memory page pool exceeds the threshold, excess memory pages need to be released, to ensure that the system memory page pool does not occupy too much system memory resource.

[0087] By setting the capacity threshold and releasing the excess memory pages, unlimited growth of the system memory page pool can be avoided, and waste of system memory resources can be avoided. Meanwhile, releasing the excess memory pages can also ensure that the memory pages in the system memory page pool have a certain timeliness, to avoid that the memory pages not used for a long time occupy too much space.

[0088] For example, assuming that the capacity upper limit of the system memory page pool is set to 200 MB. When the total amount of memory pages in the system memory page pool exceeds 200 MB, the GPU driver program automatically releases the excess memory pages. For example, if 250 MB of memory pages are saved in the system memory page pool, 50 MB of memory pages are released, to ensure that the capacity of the system memory page pool does not exceed 200 MB. The released memory pages can be used by other parts of the system, to improve the utilization rate of system memory.

[0089] In the embodiments of the present disclosure, by saving the original system memory page to the system memory page pool, frequent memory allocation and release operations can be reduced, and the efficiency of memory management can be improved. At the same time, as a buffer, the system memory page pool can quickly respond to the data replacement demand between the video memory and the system memory. In addition, by setting a capacity threshold and releasing excess memory pages, the system memory page pool can be prevented from growing indefinitely, and the rational use of system memory resources can be ensured. This helps to improve the overall performance and stability of the system and avoid performance degradation or system crashes due to insufficient memory resources.

[0090] In a possible implementation, the identifying the cold page in the GPU page table comprises: preferentially traversing the GPU page table of a process that is not active in a preset time interval to identify the cold page in the GPU page table.

[0091] The GPU driver can monitor the GPU activity status of each process in the virtual machine and determine whether the process is active. An active process is a process that frequently calls GPU resources for graphics rendering, calculation, and other operations in a recent preset time interval. For example, if a process running a graphics-intensive game sends rendering commands to the GPU multiple times in the past 10 seconds, it is considered an active process. An inactive process is a process that does not call GPU resources in a recent preset time interval.

[0092] The GPU driver can monitor the activity status of the process with the help of hardware-assisted functions and software timers. On the hardware level, the GPU hardware performance counting component can capture the activity signals of the GPU in real time, such as the usage of the rendering pipeline and the workload of the calculation unit. On the software level, the GPU driver sets a timer to record the time stamp of the last access of the GPU by each process.

[0093] According to the access frequency and time interval in the recent preset time interval, the processes are divided into active processes and inactive processes. For example, if a process does not access the GPU at all in the last 10 seconds or the access interval exceeds a certain threshold (such as 5 seconds), it is classified as an inactive process.

[0094] When the remaining amount of GPU video memory is lower than the set amount threshold, the GPU driver program starts the cold page identification process. At this time, the scanning thread preferentially traverses the GPU page table of the process that is not active in the preset time period. For example, if there are multiple processes in a virtual machine, process A and process B have been inactive in the past 10 seconds, and process C is active, the scanning thread will preferentially scan the GPU page table of process A and process B.

[0095] Since the inactive process has not accessed the GPU resource in recent time, the occupied GPU memory is likely to contain a large number of cold pages. By preferentially identifying these cold pages and replacing them to the system memory, the GPU memory space can be quickly released to meet the demand of the current active process for GPU memory. For example, assuming that the inactive process occupies 100MB of GPU memory, by preferentially replacing the cold pages, a large amount of GPU memory space can be vacated for other processes in a short time.

[0096] Therefore, the overall efficiency of the system can be improved, and blind search of cold pages in the GPU page table of all processes is avoided, and instead, efforts are concentrated on processing those inactive processes that are most likely to contain cold pages. In this way, the workload of the scanning thread can be reduced, the speed of cold page identification can be accelerated, and the overall performance can be improved.

[0097] Figure 3 A specific application scenario of an embodiment of the present disclosure is shown in a schematic diagram as shown in Figure 3 As shown, the virtual machine runs a graphics application, and then the GPU driver allocates memory, sets the GPU page table, and the GPU hardware accesses the memory. At this time, the MMU of the GPU sets the memory bit of the GPU page table item to 1. When the thread detects that the available GPU memory is less, the GPU page table of the process is traversed to identify cold pages and hot pages. If a hot page is identified and is system memory, the hot page is replaced to the GPU memory; if a cold page is identified and is GPU memory, the cold page is replaced to the system memory.

[0098] In a possible implementation, the GPU memory management method can be executed by electronic devices such as terminal devices and servers. The terminal device can be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. The method can be executed by a processor invoking computer-readable instructions stored in a memory.

[0099] In addition, the present disclosure also provides a GPU memory management apparatus, an electronic device, a computer-readable storage medium, and a program, which can be used to implement any of the GPU memory management methods provided by the present disclosure. The corresponding technical solutions and descriptions are described in the method section and are not repeated here.

[0100] Figure 4 A block diagram of a GPU memory management apparatus according to an embodiment of the present disclosure is shown in Figure 4 As shown, the apparatus 20 includes:

[0101] According to an aspect of the present disclosure, a GPU memory management apparatus is provided, which is applied to a GPU driver in a virtual machine and includes:

[0102] The cold page identification module 21 is configured to identify a cold page in the GPU page table when it is detected that the GPU memory remaining amount is lower than the set remaining amount threshold; the cold page is a page table entry that has not been accessed in a preset time interval.

[0103] The replacement module 22 is configured to, in the case that the cold page is GPU memory, replace a physical address corresponding to a virtual address in the cold page with a physical address in the system memory, so as to replace GPU memory data corresponding to the cold page into the system memory.

[0104] In a possible implementation, the device further includes:

[0105] The hot page identification module is configured to identify a hot page in the GPU page table; the hot page is a page table entry that has been accessed in a preset time interval.

[0106] The replacement module is configured to replace a physical address corresponding to a virtual address in the hot page with a physical address in the GPU memory, so as to replace memory data corresponding to the hot page into the GPU memory.

[0107] In a possible implementation, the replacement module is configured to:

[0108] allocate GPU memory from a GPU memory idle pool;

[0109] modify the physical address in the hot page to the newly allocated GPU memory address, and reset an access flag in the hot page to a first identifier; the first identifier is used to indicate that the GPU page table is a cold page.

[0110] In a possible implementation, the device further includes:

[0111] The saving module is configured to save an original system memory page to a preset system memory page pool; the original system memory page is a system memory page that originally stores memory data corresponding to the hot page before the memory data is replaced into the GPU memory.

[0112] The releasing module is configured to release excess memory pages when a capacity of the system memory page pool exceeds a set capacity threshold.

[0113] In a possible implementation, the replacement module is configured to:

[0114] obtain an idle system memory address from a preset system memory page pool;

[0115] modify a physical address corresponding to a virtual address in the cold page to the obtained system memory address, and maintain an access flag in the cold page as the first identifier;

[0116] release GPU memory corresponding to the cold page.

[0117] In a possible implementation, the apparatus further includes:

[0118] The identification changing module is configured to set an access flag in the corresponding page table entry to a second identification when the GPU accesses the GPU page table; the second identification is used to indicate that the GPU page table is a hot page.

[0119] In a possible implementation, the cold page identification module 21 is configured to preferentially traverse the GPU page table of the process that is not active in the preset time interval to identify the cold page in the GPU page table.

[0120] In a possible implementation, the apparatus further includes:

[0121] The initialization module is configured to initialize the GPU driver in response to a GPU driver loading instruction of the virtual machine.

[0122] The memory creating module is configured to create a periodically running scanning thread to perform the cold page and hot page identification operation in the case of successful initialization.

[0123] The method has specific technical association with the internal structure of the computer system, and can solve the technical problem of how to improve the hardware operation efficiency or execution effect (including reducing data storage amount, reducing data transmission amount, improving hardware processing speed, etc.), so as to obtain the technical effect of improving the internal performance of the computer system in line with the natural law.

[0124] In some embodiments, the apparatus provided by the embodiments of the present disclosure has functions or includes modules that can be used to execute the methods described in the above method embodiments, and the specific implementation can refer to the description of the above method embodiments. For brevity, they will not be repeated here.

[0125] The embodiments of the present disclosure also propose a computer-readable storage medium having computer program instructions stored thereon, wherein the computer program instructions are executed by a processor to implement the above method. The computer-readable storage medium can be a volatile or non-volatile computer-readable storage medium.

[0126] The embodiments of the present disclosure also propose an electronic device, including a processor, a memory for storing processor-executable instructions, wherein the processor is configured to invoke the instructions stored in the memory to execute the above method.

[0127] The embodiments of the present disclosure also provide a computer program product, including computer readable code or a non-volatile computer readable storage medium carrying computer readable code, when the computer readable code is run in the processor of the electronic device, the processor in the electronic device executes the above method.

[0128] An electronic device can be provided as a terminal, a server, or other forms of devices.

[0129] Figure 5 A block diagram of an electronic device according to an embodiment of the present disclosure is shown. For example, the electronic device 1900 can be provided as a server or a terminal device. Referring to FIG. 19, Figure 5 The electronic device 1900 includes a processing component 1922, further including one or more processors, and a memory resource represented by a memory 1932, for storing instructions executable by the processing component 1922, such as an application program. The application program stored in the memory 1932 can include one or more than one module each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute the instructions to perform the above-mentioned method.

[0130] The electronic device 1900 can further include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output (I / O) interface 1958. The electronic device 1900 can operate based on an operating system stored in the memory 1932, such as Microsoft Windows Server TM ), Apple's graphical user interface-based operating system (Mac OSX TM ), multi-user multi-process computer operating system (Unix TM ), free and open source Unix-like operating system (Linux TM ), open source Unix-like operating system (FreeBSD TM ) or the like.

[0131] In an exemplary embodiment, a non-transitory computer readable storage medium, such as the memory 1932 including computer program instructions, is also provided, which can be executed by the processing component 1922 of the electronic device 1900 to complete the above-mentioned method.

[0132] The present disclosure can be a system, a method, and / or a computer program product. The computer program product can include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present disclosure.

[0133] Computer readable storage media can be tangible storage media which can retain and store instructions for use by an instruction execution device. Computer readable storage media can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer readable storage media include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.

[0134] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.

[0135] Computer readable program instructions for carrying out operations of the present disclosure can be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate array (FPGA), or programmable logic array (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.

[0136] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0137] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0138] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0139] The flow diagrams and the block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow diagrams and the block diagrams can represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logic functions. In some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flow diagrams, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and

[0140] The computer program product can be embodied in a tangible medium of

[0141] The above description of the various embodiments is intended to be illustrative in all aspects, rather than being restrictive. Those skilled in the art can refer to the description of the various embodiments to make modifications and / or improvements.

[0142] Those skilled in the art can understand that, in the above-described method of the specific embodiments, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible inherent logic.

[0143] If the technical solutions of the present application involve personal information, the product applying the technical solutions of the present application has been explicitly informed of the personal information processing rules before processing the personal information, and has obtained the personal independent consent. If the technical solutions of the present application involve sensitive personal information, the product applying the technical solutions of the present application has obtained the personal independent consent before processing the sensitive personal information, and at the same time meets the requirement of "explicit consent". For example, at the personal information collection device such as camera, a clear and prominent mark is set to inform that it has entered the personal information collection range and will collect personal information. If the individual voluntarily enters the collection range, it is considered to agree to collect personal information. Or on the device for processing personal information, through the pop-up information or by asking the individual to upload his personal information, the individual's authorization is obtained under the condition that the device uses obvious mark / information to inform the individual of the personal information processing rules. The personal information processing rules can include personal information processor, personal information processing purpose, processing method and personal information type, etc.

[0144] The above has described various embodiments of the present disclosure, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes are obvious to those skilled in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles, practical application or improvement of technology in the market of the embodiments, or to enable other ordinary skilled in the art to understand the embodiments disclosed herein.

Claims

1. A method for managing a video memory, the method comprising: A GPU driver applied to a virtual machine, comprising: when it is detected that the remaining amount of GPU display memory is lower than a set threshold, identifying a cold page in a GPU page table; wherein the cold page is a page table entry that has not been accessed in a preset time interval; in the case that the cold page is display memory, replacing the physical address corresponding to the virtual address in the cold page with a physical address in system memory, so as to replace the display memory data corresponding to the cold page into system memory.

2. The method of claim 1, wherein, The method further comprises: identifying a hot page in the GPU page table; wherein the hot page is a page table entry that has been accessed in a preset time interval; replacing the physical address corresponding to the virtual address in the hot page with a physical address in GPU display memory, so as to replace the memory data corresponding to the hot page into GPU display memory.

3. The method of claim 2, wherein, The replacing the physical address corresponding to the virtual address in the hot page with a physical address in GPU display memory comprises: allocating display memory from a display memory idle pool; modifying the physical address in the hot page to the newly allocated display memory address, and resetting the access flag in the hot page to a first identifier; the first identifier is used to indicate that the GPU page table is a cold page.

4. The method of claim 3, wherein, After modifying the physical address in the hot page to the newly allocated display memory address, the method further comprises: saving the original system memory page to a preset system memory page pool; the original system memory page is the system memory page that originally stores the memory data corresponding to the hot page before the memory data corresponding to the hot page is replaced into GPU display memory; when the capacity of the system memory page pool exceeds a set capacity threshold, releasing the excess memory pages.

5. The method of claim 1, wherein, The replacing the physical address corresponding to the virtual address in the cold page with a physical address in system memory comprises: obtaining an idle system memory address from a preset system memory page pool; modifying the physical address corresponding to the virtual address in the cold page to the obtained system memory address, and maintaining the access flag in the cold page as the first identifier; releasing the display memory corresponding to the cold page.

6. The method of claim 1, wherein, The method further comprises: when the GPU accesses the GPU page table, setting the access flag in the corresponding page table entry to a second identifier; the second identifier is used to indicate that the GPU page table is a hot page.

7. The method of claim 1, wherein, The identifying a cold page in a GPU page table comprises: preferentially traversing the GPU page table of a process that is not active in a preset time interval, to identify a cold page in the GPU page table.

8. The method of claim 1, wherein, The method further comprises: in response to a GPU driver loading instruction of the virtual machine, initializing the GPU driver; in the case of successful initialization, creating a periodically running scan thread to perform cold page and hot page identification operations.

9. A video memory management device, comprising: A GPU driver applied to a virtual machine, comprising: a cold page identification module, configured to identify a cold page in a GPU page table when it is detected that the remaining amount of GPU display memory is lower than a set threshold; wherein the cold page is a page table entry that has not been accessed in a preset time interval; a replacement module, configured to, in the case that the cold page is display memory, replace the physical address corresponding to the virtual address in the cold page with a physical address in system memory, so as to replace the display memory data corresponding to the cold page into system memory.

10. An electronic device, comprising: comprising: a processor; a memory for storing processor-executable instructions; The processor is configured to invoke the instructions stored in the memory to implement the method in any one of claims 1 to 8.

11. A computer-readable storage medium having stored thereon computer program instructions, wherein, The computer program instructions, when executed by the processor, implement the method in any one of claims 1 to 8.

Citation Information

Cited By

  • Video memory control method, device, equipment and system and computer storage medium

    CN121255472A

  • Display card storage resource management method and device, electronic equipment and readable medium

    CN121501520A

  • Method and device for managing display card storage resources, electronic equipment and readable medium

    CN121501520B