GPU video memory channel allocation method, apparatus and device, and storage medium
By pre-applying and allocating physical video memory pages in cloud services, the resource conflict problems in GPU resource virtualization and allocation are solved, and efficient GPU resource utilization and task performance guarantee are achieved.
Patent Information
- Application Number
- CN202510033200.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In cloud services, there are resource conflicts in virtualization and allocation of GPU resources, resulting in performance interference and affecting the service quality of the task.
By pre-applying multiple physical memory pages and allocating memory channels according to the mapping table of colored partitions and memory colors, ensuring that each task gets dedicated memory resources.
It effectively avoids competition in video memory channels between different tasks, improves the utilization rate of GPU resources, and ensures the performance of high-priority tasks.
Smart Images

Figure CN119938331A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present disclosure relate to the fields of cloud computing and GPU virtualization, and in particular to a method, apparatus, device and storage medium for allocating GPU memory channels. Background Art
[0002] In recent years, with the vigorous development of industries such as artificial intelligence, big data, and scientific computing, the demand for computer computing power in all walks of life has continued to increase. As the CPU architecture is difficult to cope with large-scale parallel computing, the contradiction between hardware resources and computing power requirements has gradually emerged, and the widespread use of GPUs has effectively alleviated this contradiction. As GPUs perform better and better in simple repetitive calculations on large-scale data, they have been widely used in computing-intensive applications, and GPU parallel computing technology has become a hot research direction.
[0003] In order to reduce operation and maintenance costs, more and more companies are deploying their GPU applications on cloud services. Cloud service providers can virtualize a GPU into multiple sub-GPUs, thereby deploying multiple GPU applications and improving GPU resource utilization. However, resource conflicts will arise between different tasks on a GPU, which will cause performance interference and affect the service quality of each task. Therefore, it is particularly important to virtualize a GPU and allocate resources on the GPU to different tasks. Summary of the invention
[0004] The embodiments of the present disclosure provide a method, apparatus, device and storage medium for allocating GPU memory channels.
[0005] In a first aspect, an embodiment of the present disclosure provides a method for allocating GPU video memory channels, comprising: pre-applying for multiple physical video memory pages from a physical video memory space; mapping the video memory colors to the multiple physical video memory pages according to a mapping table between shading partitions of the physical video memory pages and video memory colors, and obtaining video memory channels corresponding to each shading partition, wherein a shading partition refers to a segment of equal-length and continuous physical address space that is mapped to the same video memory channel and stored in a physical video memory page; in response to detecting a target task that performs video memory operations according to a required color and a required space, querying an unallocated shading partition of the video memory channel corresponding to the required color as a target shading partition; allocating a target physical video memory page to which the target shading partition belongs from the multiple physical video memory pages to the target task according to the required space, and returning the virtual address space corresponding to the target physical video memory page to the target task for use.
[0006] In some embodiments, a physical video memory page includes at least two shading partitions, each shading partition has a number representing the offset size of the shading partition relative to the starting address of the physical video memory page, and each video memory channel includes the same number of shading partitions; the target physical video memory page to which the target shading partition belongs is allocated from the multiple physical video memory pages to the target task according to the required space, comprising: counting the number of unallocated shading partitions in the shading partitions corresponding to the shading partition numbers of the required colors; and allocating the physical video memory page to which the target shading partition that meets the required space belongs from the unallocated shading partitions to the target task according to the shading partition number with the largest total number.
[0007] In some embodiments, the method further includes: hijacking a request to call an API of a native dynamic link library through a custom dynamic link library, wherein each API in the native dynamic link library has an API declaration with the same name in the custom dynamic link library; in response to detecting that the hijacked request involves a video memory operation, forwarding the hijacked request to a custom API, wherein the custom API converts video memory operations for a continuous physical address space into video memory operations for a physical address space of a specified color, wherein the video memory operations include at least one of the following: allocating video memory space, releasing video memory space, copying data in the video memory space, and resetting data in the video memory space.
[0008] In some embodiments, the method further includes: in response to detecting that the GPU is disconnected from the host, releasing the plurality of physical video memory pages.
[0009] In some embodiments, the method further includes: in response to a request to hijack a GPU-side function, forwarding the hijacked request to a custom GPU-side function, wherein the custom GPU-side function adds an instruction to translate a memory access offset, wherein the instruction to translate a memory access offset maps a memory access address of the GPU-side function from a continuous virtual address space to a discrete video memory shading partition.
[0010] In some embodiments, the custom GPU-side function is generated by the following steps: detecting whether there is source code for the GPU-side function; in response to detecting the existence of source code for the GPU-side function, adding instructions for translating memory access offsets to the source code to obtain a custom GPU-side function; in response to detecting the absence of source code for the GPU-side function, obtaining a binary file of the GPU-side function, disassembling the binary file to obtain source code, and adding instructions for translating memory access offsets to the source code to obtain a custom GPU-side function.
[0011] In some embodiments, the querying of the unallocated shading partitions of the video memory channel corresponding to the required color includes: querying the number of unallocated shading partitions of the video memory channel corresponding to the required color in a structure recording the multiple physical video memory pages, and updating the variables in the structure according to the allocated target shading partitions. The structure includes the following variables: a starting physical address, an ending physical address, the number of unallocated shading partitions, the total number of shading partitions, an allocated color number, an allocated partition number, and a linked list storing unallocated physical video memory pages; and the method further includes: updating the variables in the structure according to the allocated target physical video memory page.
[0012] In a second aspect, an embodiment of the present disclosure provides a device for allocating GPU video memory channels, comprising: an initialization unit, configured to apply for multiple physical video memory pages in advance; a shading unit, configured to map video memory colors to the multiple physical video memory pages according to a mapping table between shading partitions of the physical video memory pages and video memory colors, and obtain video memory channels corresponding to each shading partition; a query unit, configured to, in response to detecting a target task that performs video memory operations according to a required color and a required space, query an unallocated shading partition of the video memory channel corresponding to the required color as a target shading partition; an allocation unit, configured to allocate a target physical video memory page to which the target shading partition belongs from the multiple physical video memory pages to the target task according to the required space, and return the virtual address space corresponding to the target physical video memory page to the target task for use.
[0013] In some embodiments, a physical video memory page includes at least two shading partitions, each shading partition has a number, which represents the offset size of the shading partition relative to the starting address of the physical video memory page, and each video memory channel includes the same number of shading partitions; the allocation unit is further configured to: count the number of unallocated shading partitions in the shading partitions corresponding to each shading partition number of the required color; and allocate the physical video memory page belonging to the target shading partition that meets the required space from the unallocated shading partitions to the target task according to the shading partition number with the largest total number.
[0014] In some embodiments, the device also includes a first hijacking unit, which is configured to: hijack a request to call an API of a native dynamic link library through a custom dynamic link library, wherein each API in the native dynamic link library has an API declaration with the same name in the custom dynamic link library; in response to detecting that the hijacked request involves a video memory operation, forward the hijacked request to a custom API, wherein the custom API converts the video memory operation for a continuous physical address space into a video memory operation for a specified shaded physical address space, wherein the video memory operation includes at least one of the following: allocating video memory space, releasing video memory space, copying data in the video memory space, and resetting data in the video memory space.
[0015] In some embodiments, the initialization unit is further configured to: in response to detecting that the GPU is disconnected from the host, release the plurality of physical video memory pages.
[0016] In some embodiments, the device further comprises a second hijacking unit configured to: in response to hijacking a request to call a GPU side function, forward the hijacked request to a custom GPU side function, wherein the custom GPU side function adds an instruction to translate the memory access offset. The instruction to translate the memory access offset maps the memory access address of the GPU side function from a continuous virtual address space to a discrete video memory shading partition.
[0017] In some embodiments, the device also includes a function generation unit, which is configured to: detect whether there is source code for a GPU side function; in response to detecting the existence of source code for the GPU side function, add instructions for translating memory access offsets to the source code to obtain a custom GPU side function; in response to detecting the absence of source code for the GPU side function, obtain a binary file of the GPU side function, disassemble the binary file to obtain source code, and add instructions for translating memory access offsets to the source code to obtain a custom GPU side function.
[0018] In some embodiments, the query unit is further configured to: query the number of unallocated shading partitions of the video memory channel corresponding to the required color in a structure recording the multiple physical video memory pages, wherein the structure includes the following variables: a starting physical address, an ending physical address, a number of unallocated shading partitions, a total number of shading partitions, an allocated color number, an allocated partition number, and a linked list storing unallocated physical video memory pages; and the method further includes: updating the variables in the structure according to the allocated target shading partition.
[0019] In a third aspect, an embodiment of the present disclosure provides an electronic device, comprising: one or more processors; a storage device on which one or more computer programs are stored, and when the one or more computer programs are executed by the one or more processors, the one or more processors implement a method as described in any one of the first aspects.
[0020] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method as described in any one of the first aspects.
[0021] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, including a computer program, which implements the method as described in any one of the first aspects when executed by a processor.
[0022] The embodiments of the present invention provide a method and device for allocating GPU video memory channels, which, by modifying the GPU driver code and implementing video memory shading in the GPU driver, hijacks and modifies the API calls of the GPU application involving video memory operations, modifies the GPU-side functions of the GPU application, and hijacks the GPU application's request to call the GPU-side functions, thereby allocating the GPU video memory channels to different tasks for use, thereby preventing low-priority tasks from occupying too much GPU video memory bandwidth and affecting the performance of high-priority tasks.
[0023] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Other features, objects and advantages of the present disclosure will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:
[0025] Figure 1 is an exemplary system architecture diagram in which an embodiment of the present disclosure may be applied;
[0026] Figure 2 A schematic diagram of the overall process of a method for allocating GPU memory channels provided in an embodiment of the present disclosure;
[0027] Figure 3a-3b A schematic diagram of a process of initializing video memory shading provided in an embodiment of the present disclosure;
[0028] Figure 4a-4b A flowchart of adding a user-defined video memory color set API and a video memory operation API supporting shading in a GPU driver provided by an embodiment of the present disclosure;
[0029] Figure 5a-5b A schematic diagram of a process of hijacking a GPU application program's request to operate video memory provided by an embodiment of the present disclosure;
[0030] Figure 6 A schematic diagram of a process for modifying a GPU application program GPU-side function code according to an embodiment of the present disclosure;
[0031] Figure 7 A schematic diagram of a process of hijacking a GPU application to call a GPU side function provided by an embodiment of the present disclosure;
[0032] Figure 8 It is a structural schematic diagram of an embodiment of a device for allocating GPU memory channels disclosed in the present invention;
[0033] Fig. 9 It is a schematic diagram of the structure of a computer system of an electronic device suitable for implementing an embodiment of the present disclosure. DETAILED DESCRIPTION
[0034] The present disclosure is further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the relevant invention, rather than to limit the invention. It is also necessary to explain that, for ease of description, only the parts related to the relevant invention are shown in the accompanying drawings.
[0035] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other. The present disclosure will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0036] Figure 1 An exemplary system architecture is shown to which an embodiment of the method for allocating GPU memory channels of the present disclosure can be applied.
[0037] like Figure 1 As shown, the system architecture includes a host 101 (i.e., a CPU) and a device 102 (i.e., a GPU). The software architecture on the CPU side includes: a GPU application 1011 (also called a user application), a dynamic link library 1012, and a GPU driver 1013 (which may include multiple GPU driver modules). The hardware structure on the GPU side includes: a computing unit 1021 (running GPU side functions) and a storage unit 1022.
[0038] The user application interacts with the GPU by calling an API (e.g., cuLaunchKernel) from a dynamic link library 1012 (e.g., CUDA). These dynamic link libraries forward the request to the GPU driver module. Some of the GPU driver modules are open source, for example, nvidia-uvm is a kernel module in the NVIDIA GPU driver, which stands for "Unified Virtual Memory". nvidia-uvm is a module for managing the unified memory of the GPU.
[0039] The computing unit 1021 includes multiple TPCs (texture processing clusters), and one TPC includes multiple SMs (streaming multiprocessors). Each SM includes multiple SMPs (SM partitions). The threads of the kernel (GPU side functions) are scheduled to different SMPs.
[0040] The storage unit 1022 includes L1 cache, L2 cache and video memory (shared by all SMs).
[0041] It should be noted that the method for allocating GPU memory channels provided in the embodiments of the present disclosure is generally executed by a CPU, and accordingly, the device for allocating GPU memory channels is generally disposed in the CPU.
[0042] It should be understood that Figure 1 The number of CPUs and GPUs in the embodiment is only for illustration. Any number of CPUs and GPUs may be provided as required.
[0043] Figure 2 A process 200 of an embodiment of a method for allocating GPU memory channels according to the present disclosure is shown. The method for allocating GPU memory channels comprises the following steps:
[0044] Step 201, pre-apply for a plurality of physical video memory pages from the physical video memory space.
[0045] In this embodiment, coloring refers to allocating a video memory space mapped to a specific video memory channel set for a task. Coloring can make the video memory channels used by different tasks independent of each other, thus avoiding video memory channel conflicts.
[0046] Add a structure describing the virtual address space after shading in the GPU driver code. The structure includes but is not limited to the following variables: starting physical address, ending physical address, number of unallocated shading partitions, total number of shading partitions, allocated color number, allocated partition number, and linked list for storing unallocated physical video memory pages. Through this structure, the unallocated shading partition of the video memory channel corresponding to the required color can be queried.
[0047] A plurality of physical video memory pages may be pre-applied from the physical video memory space at a predetermined ratio (eg, 70%) as a reserved video memory pool for subsequent video memory allocation.
[0048] Step 202 , according to a mapping table between the shading partitions of the physical memory pages and the memory colors, the memory colors are mapped to a plurality of physical memory pages to obtain the memory channels corresponding to the shading partitions.
[0049] In this embodiment, the video memory channel to which each video memory shading partition belongs is obtained according to a mapping table between the shading partitions of the physical video memory page and the video memory colors.
[0050] For each video memory color, the assigned color number can be filled into the structure that records the pre-applied physical video memory pages. The total number of physical video memory pages corresponding to the video memory color, the starting physical address of the physical address space, and the ending physical address can be filled into the structure. At initialization, all physical video memory pages are unallocated. After some physical video memory pages are allocated to a task, the number of unallocated physical video memory pages will decrease accordingly.
[0051] Step 203 , in response to detecting a target task that performs a video memory operation according to a required color and a required space, querying an unallocated shading partition of a video memory channel corresponding to the required color as a target shading partition.
[0052] In this embodiment, if it is detected that the API in the dynamic link library is called by a target task initiated by a user program, the relevant information of the target task that calls the API can be obtained, and the target task specifies the required color and required space for performing video memory operations. According to the required color, the structure of the physical video memory page previously applied for can be queried, and the unallocated coloring partition of the video memory channel corresponding to the required color can be obtained as the target coloring partition. The specific operation process can be seen in Figure 5b Steps S3301-S3305 shown.
[0053] Step 204 , allocating the target physical video memory page to which the target shading partition belongs from a plurality of physical video memory pages to the target task according to the required space, and returning the virtual address space corresponding to the target physical video memory page to the target task for use.
[0054] In this embodiment, according to the size of the required space, the target physical video memory page belonging to the target shading partition of the corresponding size is allocated from the unallocated physical video memory page to the target task. For example, if the required space is 1 MiB, the physical video memory page in the shading partition number with a remaining space greater than 1 MiB can be selected from the unallocated physical video memory page as the target physical video memory page. The specific operation process can be seen in Figure 4b Steps S2101-S2105 shown.
[0055] At present, virtualization technologies for GPUs are mainly based on time division multiplexing and space division multiplexing. GPU virtualization technology based on time division multiplexing allocates different time slices of a GPU to different tasks to avoid resource competition caused by running different tasks in parallel. However, due to the different resource consumption of different GPU tasks, the GPU resource utilization rate of GPU virtualization methods based on time division multiplexing is low. GPU virtualization technology based on space division multiplexing uses the multi-streaming or context merging features supported by GPU hardware to enable different tasks to be executed in parallel, solving the problem of low GPU resource utilization rate of methods based on time division multiplexing, but it cannot isolate the competition for video memory channels between different tasks. Although some products can implement resource segmentation on hardware, allocate computing resources, video memory space and video memory channels to different GPU instances, so that there will be no competition for video memory channels between multiple applications on the GPU. However, this solution requires modifying the GPU hardware architecture, so it cannot be applied to most GPUs.
[0056] Only memory shading is not limited by hardware conditions and can implement memory channel segmentation at the software level, that is, allocate the memory space mapped by different memory channels to different GPU applications. However, existing memory shading technologies are either limited by the page size of the GPU's memory management unit and cannot achieve fine-grained memory shading, or require the use of additional shadow page tables to achieve mapping, which brings additional memory access overhead.
[0057] The method provided by the above-mentioned embodiment of the present disclosure does not need to use an additional shadow page table to implement mapping, and does not bring additional memory access overhead. It optimizes the GPU virtualization technology of the prior art, solves the problem of low GPU resource utilization based on time division multiplexing technology, and can also solve the problem that the GPU virtualization technology based on space division multiplexing cannot isolate the competition of video memory channels between different tasks. It realizes universal, low-overhead, software-defined video memory channel segmentation in GPU virtualization.
[0058] In some optional implementations of the present embodiment, a physical video memory page includes at least two shading partitions, each shading partition has a number, and the number represents the offset size of the shading partition relative to the starting address of the physical video memory page, and each video memory channel includes the same number of shading partitions, wherein a shading partition refers to a continuous physical address space that is mapped to the same video memory channel and stored in a physical video memory page; the allocating a target physical video memory page from the unallocated physical video memory page to the target task according to the required space includes: counting the total number of unallocated physical video memory pages in the shading partitions corresponding to each shading partition number; and allocating a target physical video memory page belonging to the target shading partition that meets the required space from the unallocated physical video memory page to the target task according to the shading partition number with the largest total number. The target shading partition is the unallocated shading partition of the video memory channel corresponding to the required color.
[0059] For example, a physical video memory page can be divided into two shading partitions, and each time a physical video memory page is allocated, the number of the remaining blocks is selected. Through shading partitions, video memory shading is not limited by the page size of the GPU's memory management unit, thereby achieving fine-grained video memory shading, which can be applied to GPUs of new architectures.
[0060] In some optional implementations of this embodiment, the method further includes: hijacking a request to call an API of a native dynamic link library through a custom dynamic link library, wherein each API in the native dynamic link library has an API declaration with the same name in the custom dynamic link library; in response to detecting that the hijacked request involves a video memory operation, forwarding the hijacked request to a custom API, wherein the custom API converts the video memory operation for a continuous physical address space into a video memory operation for a specified shading physical address space, wherein the video memory operation includes at least one of the following: allocating video memory space, releasing video memory space, copying data in the video memory space, and resetting data in the video memory space.
[0061] Modify the code of the CPU-side dynamic link library by function hijacking. For the specific process, see Figure 5a Steps S3100-S3600 are shown.
[0062] In some optional implementations of this embodiment, the method further includes: in response to detecting that the GPU is disconnected from the host, releasing the multiple physical video memory pages.
[0063] If the GPU is disconnected from the host (CPU), the memory shading needs to be destroyed and the pre-allocated physical memory pages need to be restored to the state before the memory shading.
[0064] In some optional implementations of this embodiment, the method further includes: in response to a request to hijack a GPU-side function, forwarding the hijacked request to a custom GPU-side function, wherein the custom GPU-side function adds an instruction to translate the memory access offset. The instruction to translate the memory access offset maps the memory access address of the GPU-side function from a continuous virtual address space to a discrete video memory shading partition. This step is performed by hijacking the request to call the GPU-side function from the CPU side and forwarding it to the GPU for execution. The specific process can be seen in Figure 7 Steps S5100-S5300 are shown.
[0065] In some optional implementations of the present embodiment, a custom GPU-side function is generated by the following steps: detecting whether there is source code of the GPU-side function; in response to detecting the existence of the source code of the GPU-side function, adding instructions for translating memory access offsets to the source code to obtain a custom GPU-side function; in response to detecting the absence of the source code of the GPU-side function, obtaining a binary file of the GPU-side function, disassembling the binary file to obtain source code, and adding instructions for translating memory access offsets to the source code to obtain a custom GPU-side function.
[0066] The native GPU side function of the system needs to be modified to obtain a custom GPU side function. There are two processing methods: with source code and without source code. For the specific process, see Figure 6 Steps S4100-S4700 are shown.
[0067] In some optional implementations of the present embodiment, the querying of the unallocated shading partition of the video memory channel corresponding to the required color as the target shading partition includes: querying the number of unallocated shading partitions of the video memory channel corresponding to the required color in a structure recording the multiple physical video memory pages, wherein the structure includes the following variables: a starting physical address, an ending physical address, a number of unallocated shading partitions, a total number of shading partitions, an allocated color number, an allocated partition number, and a linked list storing unallocated physical video memory pages; and the method also includes: updating the variables in the structure according to the allocated target shading partition.
[0068] Predefine some structures to save the parameters required for the shading process. The specific process can be seen in Figure 3b The process of querying the unallocated physical video memory page of the video memory channel corresponding to the required color can be referred to in Figure 4b Steps S2101-S2105 shown.
[0069] The embodiments of the present disclosure are based on the NVIDIA RTX A2000 12GB GPU with Ampere architecture and 6 video memory channels, supporting 2KiB video memory shading granularity and 6 video memory colors. Each 4KiB physical video memory page contains two 2KiB shading partitions. In actual applications, the number of shading partitions for a physical video memory page can be more. The following is only an exemplary description, and the code implementation process on the CPU side and the GPU side in actual applications is not limited to the following examples.
[0070] In this embodiment, the function of initializing video memory coloring can be added by modifying the code of the driver. The operation process of initializing video memory coloring is as follows: Figure 3a As shown, including:
[0071] Step S1100: modify the structure and header file definition of the GPU driver.
[0072] In this embodiment, the operation process of step S1100 includes:
[0073] Step S1101, add attribute information to the structure used to describe the GPU device attributes, including but not limited to: 1) the number of video memory colors allocated to the GPU, 2) the number of shading partitions allocated to the GPU, 3) a quick lookup table that stores the page color number of each shading page, 4) the number of entries in the page color number quick lookup table, 5) the block size of the video memory page allocated by the GPU.
[0074] For example, in the nvidia-uvm module of the GPU driver, five attributes are added to the structure uvm_gpu_struct that describes the GPU device attributes in uvm_gpu.h: 1) the number of video memory colors allocated to the GPU num_allocation_mem_colors, 2) the number of shading partitions allocated to the GPU num_mem_color_partition_ids, 3) the quick lookup table color_lookup_table[] that stores the page color number of each shading page, 4) the number of entries in the page color number quick lookup table color_lookup_table[] num_lookup_table_entries, 5) the chunk size chunk_size of the video memory page allocated by the GPU.
[0075] Step S1102, modify the function of initialization attribute, in which a quick lookup table of color numbers corresponding to each physical video memory page is read. The quick lookup table can be used to quickly query the mapping relationship between the video memory physical address prepared for video memory coloring and the video memory channel number.
[0076] For example, in the nvidia-uvm module of the GPU driver, modify the uvm_hal_ampere_arch_init_properties() function in uvm_ampere.c. In the function, read the quick lookup table of the page color number of each physical video memory address of RTX A2000
[0077] color_lookup_table[]. Set num_mem_color_partition_ids=2,
[0078] num_allocation_mem_colors=6,
[0079] num_lookup_table_entries=4*1024*1024=4194304. Set chunk_size=4096.
[0080] Step S1103, add a structure describing the virtual address space after coloring in the header file. The content of the structure may include but is not limited to: 1) describing the starting physical address of the page coloring physical address space; 2) describing the ending physical address of the page coloring physical address space; 3) describing the number of unallocated coloring pages; 4) describing the total number of coloring pages in the page coloring physical address space; 5) the allocated color number; 6) the allocated partition number; 7) the linked list for storing unallocated blocks;
[0081] For example, in the uvm_pmm_gpu.h of the nvidia-uvm module, add the structure uvm_gpu_color_range_struct that describes the virtual address space after coloring. It includes 1) describing the starting physical address start_phys_addr of the page coloring physical address space; 2) describing the ending physical address end_phys_addr of the page coloring physical address space; 3) describing the number of unallocated coloring pages left_num_chunks; 4) describing the total number of coloring pages total_num_chunks of the page coloring physical address space; 5) the allocated color number allocation_color; 6) the allocated partition number allocation_partition_id; 7) the linked list free_chunks that stores the unallocated blocks;
[0082] Step S1104: Add attribute information to the structure used to describe the GPU physical page block, including but not limited to: a block linked list pre-allocated for page coloring, and a virtual address space after coloring to which the block belongs.
[0083] For example, in the structure uvm_gpu_chunk_struct describing the GPU physical page chunks of the nvidia-uvm module, a chunk list reserved_list saved for pre-allocation of page coloring is added, and the virtual address space color_range after coloring to which the chunk belongs is added.
[0084] Step S1105, add attribute information related to the shading partition to the structure used to describe the GPU physical memory manager, including but not limited to: 1) a linked list of blocks pre-allocated for page shading; 2) the number of GPU video memory colors; 3) the number of GPU video memory shading partitions; 4) a linked list of blocks reserved for page shading; 5) the shaded virtual address space corresponding to each video memory color and video memory shading partition. 6) a user-specified set of video memory colors that can be used for the next video memory allocation; 7) the video memory shading partition automatically selected by the GPU driver for the next video memory allocation; 8) the block number when the next block is allocated; 9) whether the next video memory allocation will pull the block from the block pool pre-allocated for video memory shading.
[0085] For example, in the structure uvm_pmm_gpu_struct describing the GPU physical memory manager (PMM) of the nvidia-uvm module, add: 1) pre_alloc_chunk_list, a linked list of chunks pre-allocated for page shading; 2) num_colors, the number of GPU video memory colors; 3) num_partition_ids, the number of GPU video memory shading partitions; 4) reserved_chunks, a linked list of chunks reserved for page shading; 5) color_ranges[][], a colored virtual address space corresponding to each video memory color and video memory shading partition. 6) color_bitset, a user-specified set of video memory colors that can be used for the next video memory allocation; 7) selected_partition_id, a video memory shading partition automatically selected by the GPU driver for the next video memory allocation; 8) chunk_index, the chunk number for the next chunk allocation; 9) is_alloc_from_reserved, whether the next video memory allocation will pull chunks from the pre-allocated chunk pool for video memory shading.
[0086] Step S1200: Add a function for querying the color of the physical page to which the specified physical address belongs.
[0087] For example, in uvm_ampere_mmu.c of the nvidia-uvm module, a function uvm_hal_ampere_mmu_phys_addr_to_allocation_color(uvm_gpu_t*gpu,NvU32 partition_id,NvU64 phys_addr) is added to query the color of the partition_idth partition of the physical page to which the specified physical address phys_addr belongs.
[0088] In this embodiment, the operation process of step S1200 includes:
[0089] Step S1201, obtain the physical page number entry_id to which the physical address belongs.
[0090] Step S1202: Return the value corresponding to the physical page number entry_id in the page color quick lookup table as the color of the physical page to which the specified physical address belongs.
[0091] You can also query by color partition, for example, return the entry_id value of the page color quick lookup table color_lookup_table[partition_id] of the partition_id partition, that is, color_lookup_table[partition_id][entry_id], as the color of the partition_id partition of the physical page numbered entry_id belonging to the specified physical address phys_addr.
[0092] Step S1300: Add an API for initializing video memory shading. The API is implemented based on a structure for describing GPU hardware and a structure for describing a GPU physical memory manager.
[0093] For example, add the API reserve_color_memory(gpu,pmm) to initialize the video memory coloring in uvm_pmm_gpu.c, where gpu is an instance of the uvm_gpu_t structure that describes a GPU hardware, and pmm is an instance of the uvm_pmm_gpu_t structure that describes a GPU physical memory manager.
[0094] In this embodiment, the operation process of step S1300 includes:
[0095] S1301. Initialize the shaded virtual address space.
[0096] For example, in pmm, initialize the colored virtual address space pmm->color_ranges[][] for num_allocation_mem_colors colors and num_partition_ids colored partitions, and set allocation_color = i, allocation_partition_id = j for each pmm->color_ranges[i][j];
[0097] S1302: Set the size of the video memory space to be reserved according to a predetermined ratio of the GPU video memory size.
[0098] For example, the size of the reserved video memory space resv_mem is set. In this embodiment, it is set to 70% of the GPU video memory size, that is, 0.7*12GB=8.4GB.
[0099] S1303: Apply for a physical video memory page of a specified size from a set of physical video memory pages (the physical video memory page is also called a block).
[0100] For example, call the alloc_chunk() function of the nvidia-uvm module to apply for a chunk of size chunk_size from the pmm object.
[0101] S1304. Add the linked list pointer of the requested block to the end of the linked list of the blocks reserved for page coloring.
[0102] For example, the chunk->reserved_list pointer of the chunk applied for in S1303 is added to the end of the linked list pmm->reserved_chunks.
[0103] S1305. Initialize the coloring partition number, for example, partition_id=0.
[0104] S1306: Initialize the shadow block of the block, and copy the data describing the block to the shadow block.
[0105] For example, initialize the shadow chunk shadow_chunk[partition_id] of the partition_id chunk, and copy the data describing the chunk to shadow_chunk[partition_id] through the memcpy() function.
[0106] S1307. Query the color corresponding to the shadow block through the function defined in step S1200.
[0107] For example, call the gpu->parent->arch_hal->phys_addr_to_allocation_color() function to get the corresponding color of shadow_chunk[partition_id].
[0108] S1308. Add the linked list pointer of the shadow block to the end of the linked list of the virtual address space.
[0109] For example, add the linked list pointer shadow_chunk[partition_id]->list of shadow_chunk[partition_id] to the end of the linked list pmm->color_ranges[color][partition_id].
[0110] S1309: Increase the total number of blocks and the number of remaining free blocks in the virtual address space.
[0111] For example, increase the total number of chunks pmm->color_ranges[color][partition_id]->total_num_chunks and the number of remaining free chunks pmm->color_ranges[color][partition_id]->num_left_chunks.
[0112] S1310, the coloring partition number is incremented, that is, partition_id=partition_id+1.
[0113] S1311 . Repeat S1306 to S1310 until the coloring partition number reaches the number of coloring partitions, that is, partition_id>=num_partition_id.
[0114] S1312, repeat S1303 to S1311 for a predetermined number of times, wherein the ratio of the reserved video memory space size to the chunk size is rounded up to obtain a predetermined number of times, namely (resv_mem / chunk_size) times.
[0115] S1313. Return status to the caller.
[0116] Step S1400: Add an API for destroying video memory shading when the GPU is disconnected from the host.
[0117] For example, in the uvm_pmm_gpu.c of the nvidia-uvm module, add the API free_reserved_color_memory(pmm) to destroy the video memory coloring when the GPU is disconnected from the host. Among them, pmm is an instance of the uvm_pmm_gpu_t structure that describes a GPU physical memory manager.
[0118] In this embodiment, the operation process of step S1400 includes:
[0119] Step S1401: Set the color number of the currently traversed color, for example, color=0.
[0120] Step S1402: Set the partition number of the current traversal, for example, partition_id=uvm_pmm_to_gpu(pmm)->num_mem_color_partition_ids-1.
[0121] Step S1403: traverse the linked list of unallocated blocks in the virtual address space and delete each block.
[0122] For example, traverse the free_chunks list of pmm->color_ranges[color][partition_id], call the list_del_init() function of the nvidia-uvm module, and delete each block in it.
[0123] Step S1404: Delete the partition corresponding to the current partition number of the virtual address space.
[0124] For example, call the uvm_kvfree() function of the nvidia-uvm module to delete pmm->color_ranges[color][partition_id].
[0125] Step S1405: the partition number is decremented, that is, partition_id=partition_id-1.
[0126] Step S1406: Repeat steps S1402 to S1405 until the minimum partition number is reached, partition_id<0.
[0127] Step S1407: The color number is incremented, that is, color=color+1.
[0128] Step S1408, repeat steps S1402 to S1407 until the maximum color number is reached, that is, color>=num_colors.
[0129] Step S1409, traverse the linked list consisting of the reserved blocks in the virtual address space, and delete each block therein.
[0130] For example, traverse the linked list reserved_chunks composed of the blocks reserved by pmm, and call the free_chunk() function of the nvidia-uvm module to delete each block.
[0131] Step S1410: Return status to the caller.
[0132] In this embodiment, the operation process of adding an API for user-defined video memory color set in the GPU driver and adding an API for supporting the allocation, recycling, copying and resetting of video memory data to support the coloring is as follows: Figure 4a As shown, including:
[0133] Step S2100, add a function for setting color, which has two input parameters: 1) a 64-bit unsigned integer value describing the color set to be assigned, where the i-th binary bit is 1, indicating that the color set contains color i, otherwise it does not contain color i; 2) the size of the video memory space to be allocated next time. When the GPU has enabled video memory shading, this function is used to set the video memory color set to be allocated by the next user program and determine the partition number of the video memory to be allocated by the next user program.
[0134] For example, add the uvm_api_set_color_bitset() API to the nvidia-uvm module. This API passes in two parameters. 1) request_color_bitset is a 64-bit unsigned integer value describing the color set to be allocated, where the i-th binary bit is 1, indicating that the color set contains color i, otherwise it does not contain it; 2) request_allocated_size_bytes is the size of the video memory space to be allocated next time. When the GPU has enabled video memory shading, this function is used to set the video memory color set to be allocated by the next user program and to determine the partition number of the video memory to be allocated by the next user program. The video memory color set to be allocated by the next user program and the size of the video memory space to be allocated next time are the parameters passed in by the user's GPU application when calling the API to request video memory allocation. If the video memory page is 4KiB in size and the shading partition size is 2KiB, then the number of chunks to be applied for is (4KiB / 2KiB)*(request_allocated_size_bytes / 4KiB).
[0135] In this embodiment, the operation process of step S2100 includes:
[0136] The number with the most remaining blocks is selected from all the coloring partition numbers as the partition number to be selected next time, and the color requested by the user program is used as the color for the next application for video memory allocation.
[0137] For example, step S2101, obtain a structure instance of the physical memory manager of the current GPU, for example, pmm in uvm_pmm_gpu_t.
[0138] Step S2102: Select the number request_partition_id with the most remaining blocks from all the colored partition numbers.
[0139] Step S2103, set the partition number to be selected next time pmm->selected_partition_id=request_partition_id.
[0140] Step S2104: Set the color set pmm->color_bitset=request_color_bitset for the next request for video memory allocation.
[0141] Step S2105: Return the partition number selected next time (eg, pmm->selected_partition_id) to the caller.
[0142] Step S2200: Add a function for allocating video memory space, where the function is used to allocate video memory space from a reserved video memory pool.
[0143] For example, a try_alloc_user_color_chunk() function is added to the nvidia-uvm module, which is used to allocate video memory space from the blocks reserved for video memory coloring in step S1000.
[0144] In this embodiment, the operation process of step S2200 includes: taking the block number modulo the number of colors allowed to be allocated when allocating blocks next time as the color number of the currently allocated block, and then using the color number and the partition number obtained in step S2100 as the index of the virtual address space to obtain the head block from the linked list of free coloring blocks.
[0145] For example, in step S2201, the number of colors in the color set currently allowed to be allocated, num_available_colors, is calculated.
[0146] Step S2202, calculate the color number of the currently allocated chunk: curr_color=pmm->chunk_index%num_available_colors.
[0147] Step S2203: Obtain the shading partition number selected_partition_id calculated in advance in step S2100 from the GPU physical memory manager pmm.
[0148] Step S2204:
[0149] The head block of the chain is obtained from the free_chunks linked list of free coloring blocks stored in pmm->color_ranges[curr_color][pmm->selected_partition_id] and returned to the caller.
[0150] Step S2300: Add a function for releasing the video memory space after rendering, wherein the released video memory space is returned to the reserved video memory pool.
[0151] For example, a try_user_color_free_chunk() function is added to the nvidia-uvm module. This function is used to release the video memory space after coloring from the user program and return the reserved blocks for initializing the video memory coloring in step S1000.
[0152] In this embodiment, the operation process of step S2300 includes:
[0153] Step S2301, obtain the linked list range to which the colored block originally belongs.
[0154] Step S2302: put the block at the end of the linked list range.
[0155] Step S2303: The number of free chunks in the linked list range is left_num_chunks=left_num_chunks+1.
[0156] Step S2400: Add a function for copying data in the video memory space, which is used to copy a piece of data of a specified length from the virtual video memory space of the source address to the target address.
[0157] For example, add the color_memcpy_chunk() function to the nvidia-uvm module. This function is used to copy a piece of data of length size from the virtual video memory space with the first address src to the virtual video memory space with the first address dest in the GPU application.
[0158] In this embodiment, the operation process of step S2400 includes:
[0159] Step S2401, initialize the variables copy_from_addr and copy_to_addr to store the source address and target address of the video memory copy respectively. copy_from_addr is set to src, and copy_to_addr is set to dest.
[0160] Step S2402: Write a LAUNCH_DMA instruction to the GPU and set the data transmission mode to PIPELINED (pipeline non-blocking mode).
[0161] Step S2403, write the LAUNCH_DMA instruction to the GPU to copy the 2048 bytes of data with the first address copy_from_addr to the 2048 bytes of data with the first address copy_to_addr.
[0162] Step S2404, assign copy_from_addr to copy_from_addr+4096, and copy_to_addr to copy_to_addr+4096.
[0163] Step S2405: If size bytes of data have been copied, jump to step S2406; otherwise, jump to step S2403.
[0164] Step S2406, calling the uvm_tracker_wait_for_entry() function of the nvidia-uvm module to wait for the DMA operation to complete.
[0165] Step S2500: Add a function for resetting the data in the video memory space.
[0166] For example, add the color_memset_chunk() function in the nvidia-uvm module. This function is used in the GPU application to reset a segment of data of length size in the virtual video memory space with the first address dest, and reassign the value of each byte therein to flush_value.
[0167] In this embodiment, the operation process of step S2500 includes:
[0168] Step S2501, initialize the variable reset_to_addr to store the first address of the target video memory space for resetting video memory data, and set reset_to_addr to dest.
[0169] Step S2502: Write a LAUNCH_DMA instruction to the GPU and set the data transmission mode to PIPELINED.
[0170] Step S2503, write the LAUNCH_DMA instruction to the GPU, and write flush_value into 2048 bytes of data with the first address being reset_to_addr.
[0171] Step S2504, assign reset_to_addr to reset_to_addr+4096.
[0172] Step S2505: If size bytes of data have been assigned, jump to step S2506; otherwise, jump to step S2503.
[0173] Step S2506, calling the uvm_tracker_wait_for_entry() function of the nvidia-uvm module to wait for the DMA operation to complete.
[0174] In this embodiment, the operation process of hijacking the GPU application to call the native dynamic link library (for example, NVIDIA CUDA library) to operate the video memory and forwarding it to the corresponding equivalent API supporting video memory shading is as follows: Figure 5a As shown, including:
[0175] Step S3100: intercept the user program's API call request to the native dynamic link library.
[0176] Step S3200: hijack the user program's call request for an API for allocating video memory blocks, for example, alloc_chunk().
[0177] Step S3300: hijack the user program's call request for APIs for allocating video memory and allocating managed video memory, for example, calls to NVIDIA CUDA library cuMemAlloc() and cuMemAllocManaged() APIs.
[0178] Step S3400: hijack the user program's call request for an API for reclaiming video memory, for example, a free_chunk() function.
[0179] Step S3500: hijack the user program's API call request for copying data in the video memory space, for example, cuMemcpyHtoD() and cuMemcpyDtoH().
[0180] Step S3600: hijack the user program's call request for an API that resets data in the video memory space, for example, calls to cuMemsetD8(), cuMemsetD16(), and cuMemsetD32() APIs of the NVIDIA CUDA library.
[0181] In this embodiment, the operation process of step S3100 can be implemented by function hijacking, for example,
[0182] Step S3101: Create a libcuda_new dynamic link library to disguise the original libcuda.so dynamic link library file of the NVIDIA CUDA library.
[0183] Step S3102: When the libcuda_new.so dynamic link library is called for the first time by the user application, the libcuda_new.so dynamic link library loads the symbols of all functions in the libcuda.so dynamic link library.
[0184] Step S3103: For each original API in the original libcuda.so dynamic link library of the NVIDIA CUDA library, add an API declaration with the same name in the libcuda_new.so dynamic link library.
[0185] Step S3104: For each API that does not need to be hijacked in the libcuda.so dynamic link library, the API with the same name in the libcuda_new dynamic link library forwards the API call request to the original libcuda.so dynamic link library.
[0186] Step S3105: compile the libcuda_new dynamic link library to generate a libcuda_new.so binary file.
[0187] Step S3106: modify the / etc / ld.so.preload file of the operating system, add the path of the libcuda_new.so file, so that the process of the user application preloads the libcuda_new.so dynamic link library when it starts.
[0188] In this embodiment, the operation process of step S3200 can be divided into two cases. If the video memory shading has been completed, the custom API is called through function hijacking. If the video memory shading has not been completed, the API in the native dynamic link library is called.
[0189] For example,
[0190] Step S3201, check whether the video memory coloring has been initialized. If the video memory coloring has been initialized, execute step S3202, otherwise execute step S3203.
[0191] Step S3202: Call the try_alloc_user_color_chunk() function to allocate a chunk from the video memory pool reserved for video memory coloring in step S1000. Jump to step S3204.
[0192] Step S3203, call alloc_chunk() function to allocate a chunk from the free uncolored video memory space.
[0193] Step S3204: Return the result of the try_alloc_user_color_chunk() or alloc_chunk() function to the caller.
[0194] In this embodiment, the operation process of step S3300 can be divided into two cases. If the video memory shading has been completed, the custom API is called through function hijacking. If the video memory shading has not been completed, the API in the native dynamic link library is called.
[0195] For example,
[0196] Step S3301, check whether the video memory coloring has been initialized. If the video memory coloring has been initialized, execute steps S3302 to S3303, otherwise execute step S3304.
[0197] Step S3302: Set the video memory color set to be allocated next time, and obtain the partition number of the virtual video memory space to be allocated next time.
[0198] For example, call the uvm_api_set_color_bitset() API to set the next allocated video memory color set and obtain the partition number partition_id of the next allocated virtual video memory space.
[0199] Step S3303, apply for the video memory space for rendering, obtain the first address of the allocated virtual video memory space, and jump to step S3305.
[0200] For example, the alloc_chunk() function described in step S3200 is called to apply for the video memory for shading, the first address original_addr of the allocated virtual video memory space is obtained, and the process jumps to step S3305.
[0201] Step S3304: calling the native API for allocating video memory and allocating managed video memory, and returning the call result to the caller.
[0202] For example, calling the CUDA library API (cuMemAlloc() or cuMemAllocManaged()) requested by the user program, and returning the CUDA library API call result to the caller.
[0203] Step S3305: Add the offset calculated according to the partition number to the starting address of step S3303, and return it to the caller as the starting address of the allocated rendering virtual video memory space.
[0204] For example, original_addr+2048*partition_id is used as the first address of the allocated virtual video memory space for shading, and is returned to the caller as the result of the API call.
[0205] Step S3400: hijack the user program's call request for a function to reclaim video memory, for example, free_chunk().
[0206] In this embodiment, the operation process of step S3400 can be divided into two cases. If the video memory shading has been completed, the custom API is called by function hijacking to recycle the blocks from the user program and return them to the reserved video memory pool. If the video memory shading has not been completed, the API in the native dynamic link library is called.
[0207] For example,
[0208] Step S3401, check whether the video memory coloring has been initialized. If the video memory coloring has been initialized, execute step S3402, otherwise execute step S3403.
[0209] Step S3402, call the try_free_user_color_chunk() function to recycle the chunks from the user program and return them to the video memory pool reserved for video memory coloring in step S1000. Jump to step S3304.
[0210] Step S3403, call the free_chunk() function to return the chunk to the uncolored video memory space.
[0211] Step S3404: Return the running result of the try_free_user_color_chunk() or free_chunk() function to the caller.
[0212] Step S3500: hijack the user program's API call request for copying data in the video memory space, for example, NVIDIA CUDA library cuMemcpyHtoD() and cuMemcpyDtoH() APIs.
[0213] In this embodiment, the operation process of step S3500 includes forwarding the API call request to the custom API for copying data in the video memory space in the custom dynamic link library:
[0214] For example, in step S3501, in the cuMemcpyHtoD() and cuMemcpyDtoH() APIs of the libcuda_new dynamic link library in step S3100, the API request is forwarded to the color_memcpy_chunk() function added to the nvidia-uvm module in step S2400.
[0215] Step S3502: Return the execution result of the color_memcpy_chunk() function to the caller.
[0216] Step S3600: hijack the user program's call request for an API that resets data in the video memory space, for example, calls to the NVIDIA CUDA library cuMemsetD8(), cuMemsetD16(), and cuMemsetD32() APIs.
[0217] In this embodiment, the operation process of step S3600 includes forwarding the API call request to the custom API for resetting the data in the video memory space in the custom dynamic link library:
[0218] For example, in step S3601, in the cuMemsetD8(), cuMemsetD16() and cuMemsetD32() APIs of the libcuda_new dynamic link library described in step S3100, the API request is forwarded to the color_memset_chunk() function added to the nvidia-uvm module described in step S2500.
[0219] Step S3602: Return the execution result of the color_memset_chunk() function to the caller.
[0220] In this embodiment, the GPU side function code of the GPU application is modified to add address translation for each memory access instruction, thereby supporting the operation process of video memory shading. Figure 6 As shown, including:
[0221] Step S4100: Check whether there is source code of the GPU side function.
[0222] If it exists, add instructions for translating memory access offsets to the source code to obtain a custom GPU side function. Otherwise, obtain the binary file of the GPU side function; disassemble the binary file to obtain the source code; add instructions for translating memory access offsets to the source code to obtain a custom GPU side function. Then compile the custom GPU side function to generate a binary file, and load the newly generated binary file.
[0223] Steps S4200-S4400 are operations when the source code of the GPU side function exists, and steps S4500-4700 are operations when the source code of the GPU side function does not exist.
[0224] Step S4200: Add a macro definition for translating the address in the source code, for example, translate(offset)=(offset)+((offset)&0xFFFFFE00). This macro definition operation is equivalent to (offset+(offset / 2048)*2048).
[0225] Step S4300: traverse the source code and map all indexes accessing the array to the shading partitions. For example, replace the array index i of all operations accessing the array with translate(i).
[0226] Step S4400, jump to step S4800.
[0227] Step S4500: hijack the user program's call request to the module loading API to obtain the path of the binary file to be loaded, for example, cuModuleLoad() in the NVIDIA CUDA library to obtain the path cubin_path of the binary CUDA file to be loaded.
[0228] Step S4600: Use a disassembly tool (eg, NVIDIA's cuobjdump) to disassemble the binary file in the path and obtain its PTX assembly code.
[0229] Step S4700: before each assembly command of the PTX assembly code for reading video memory (for example, ld.global) and writing video memory (for example, st.global), add an operation of translating the video memory address.
[0230] In this embodiment, the operation process of step S4700 includes:
[0231] Step S4701, declare a new register %temp.
[0232] Step S4702: read a line of PTX assembly commands.
[0233] Step S4703: If the line of assembly command does not involve ld.global and st.global operations, skip it.
[0234] Step S4704, obtain the offset register %offset of the row memory access command.
[0235] Step S4705, insert two instructions and.b32%temp,%offset,0xFFFFFE00 and add.u32%offset,%offset,%temp for translating the memory access offset before the line.
[0236] Step S4706: Repeat steps S4702 to S4705 until the entire binary CUDA file is traversed.
[0237] Compile the modified source code or assembly code, generate a binary file and save it to the path, and load the modified binary file from the path.
[0238] For example,
[0239] Step S4800: Use NVIDIA's nvcc compiler to compile the modified PTX or CUDA source code and save it to cubin_path.
[0240] Step S4900: Call the cuModuleLoad() API in the NVIDIA CUDA library to load the modified binary CUDA file from cubin_path.
[0241] In this embodiment, the operation process of hijacking the GPU application to call the GPU side function API is as follows: Figure 7 As shown, including:
[0242] Step S5100: Obtain the GPU function to be started and the parameter list passed into the GPU function from the custom dynamic link library by function hijacking.
[0243] For example, in step S3100, in the libcuda_new dynamic link library, the user application program calls the cuLaunchKernel() API in the NVIDIA CUDA library, and obtains the GPU function gpu_func to be started and the parameter list args_list passed into the GPU function.
[0244] Step S5200: Replace the GPU function to be started with a custom GPU function, wherein the custom GPU function is a GPU function with the same name as the GPU function to be started, and an instruction for translating a memory access offset is added to the custom GPU function.
[0245] For example, replace the GPU function gpu_func that needs to be started with the revised GPU function gpu_func_revised with the same name.
[0246] Step S5300: start the custom GPU function and pass the parameter list into the custom GPU function.
[0247] For example, the cuLaunchKernel() API in the NVIDIA CUDA library is called to start the GPU function gpu_func_revised of the same name described in step S5200, and pass the parameter list args_list described in step S5100 into the GPU function.
[0248] Further references Figure 8 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a GPU memory channel allocation device. Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.
[0249] like Figure 8 As shown, the GPU video memory channel allocation device 800 of this embodiment includes: an initialization unit 801, a shading unit 802, a query unit 803 and an allocation unit 804. The initialization unit 801 is configured to pre-apply for multiple physical video memory pages from the physical video memory space; the shading unit 802 is configured to map the video memory colors to the multiple physical video memory pages according to a mapping table between the shading partitions of the physical video memory pages and the video memory colors, and obtain the video memory channels corresponding to each shading partition; the query unit 803 is configured to query the unallocated shading partition of the video memory channel corresponding to the required color as the target shading partition in response to detecting a target task that performs video memory operations according to the required color and required space; the allocation unit 804 is configured to allocate the target physical video memory page to which the target shading partition belongs from the multiple physical video memory pages to the target task according to the required space, and return the virtual address space corresponding to the target physical video memory page to the target task for use.
[0250] In this embodiment, the specific processing of the initialization unit 801, the shading unit 802, the query unit 803 and the allocation unit 804 of the GPU memory channel allocation device 800 can be referred to. Figure 2 Corresponding to step 201, step 202, step 203, and step 204 in the embodiment.
[0251] In some optional implementations of the present embodiment, a physical video memory page includes at least two shading partitions, each shading partition has a number, which represents the offset size of the shading partition relative to the starting address of the physical video memory page, and each video memory channel includes the same number of shading partitions; the allocation unit is further configured to: count the number of unallocated shading partitions in the shading partitions corresponding to each shading partition number of the required color; and allocate the physical video memory page belonging to the target shading partition that meets the required space from the unallocated shading partitions to the target task according to the shading partition number with the largest total number.
[0252] In some optional implementations of this embodiment, the device also includes a first hijacking unit (not shown in the drawings), which is configured to: hijack a request to call an API of a native dynamic link library through a custom dynamic link library, wherein each API in the native dynamic link library has an API declaration with the same name in the custom dynamic link library; in response to detecting that the hijacked request involves a video memory operation, forward the hijacked request to a custom API, wherein the custom API converts the video memory operation for a continuous physical address space into a video memory operation for a specified shading physical address space, wherein the video memory operation includes at least one of the following: allocating video memory space, releasing video memory space, copying data in the video memory space, and resetting data in the video memory space.
[0253] In some optional implementations of this embodiment, the initialization unit 801 is further configured to: in response to detecting that the GPU is disconnected from the host, release the multiple physical video memory pages.
[0254] In some optional implementations of this embodiment, the device further includes a second hijacking unit (not shown in the drawings), configured to: in response to a request to hijack a GPU-side function, forward the hijacked request to a custom GPU-side function, wherein the custom GPU-side function adds an instruction to translate the memory access offset. The instruction to translate the memory access offset maps the memory access address of the GPU-side function from a continuous virtual address space to a discrete video memory shading partition.
[0255] In some optional implementations of the present embodiment, the device further includes a function generation unit (not shown in the drawings), which is configured to: detect whether there is source code for a GPU side function; in response to detecting the existence of source code for the GPU side function, add instructions for translating a memory access offset to the source code to obtain a custom GPU side function; in response to detecting the absence of source code for the GPU side function, obtain a binary file of the GPU side function, disassemble the binary file to obtain source code, and add instructions for translating a memory access offset to the source code to obtain a custom GPU side function.
[0256] In some optional implementations of this embodiment, the query unit 803 is further configured to: query the number of unallocated shading partitions of the video memory channel corresponding to the required color in a structure recording the multiple physical video memory pages, wherein the structure includes the following variables: a starting physical address, an ending physical address, a number of unallocated shading partitions, a total number of shading partitions, an allocated color number, an allocated partition number, and a linked list storing unallocated physical video memory pages; and the method further includes: updating the variables in the structure according to the allocated target shading partition.
[0257] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device and a readable storage medium.
[0258] An electronic device comprises: one or more processors; a storage device on which one or more computer programs are stored, and when the one or more computer programs are executed by the one or more processors, the one or more processors implement the method described in process 200.
[0259] A computer-readable medium stores a computer program, wherein the computer program implements the method described in process 200 when executed by a processor.
[0260] Fig. 9 A schematic block diagram of an example electronic device 900 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0261] like Fig. 9 As shown, the device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0262] A number of components in the device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0263] The computing unit 901 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 901 performs the various methods and processes described above, such as the road area planning method. For example, in some embodiments, the road area planning method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 908. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the road area planning method described above may be performed. Alternatively, in other embodiments, the computing unit 901 may be configured to perform the road area planning method in any other appropriate manner (e.g., by means of firmware).
[0264] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0265] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0266] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0267] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0268] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0269] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server may be a server of a distributed system, or a server combined with a blockchain. The server may also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. The server may be a server of a distributed system, or a server combined with a blockchain. The server may also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.
[0270] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.
[0271] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A method for allocating GPU memory channels, comprising: Pre-apply multiple physical video memory pages from the physical video memory space; According to a mapping table between shading partitions of a physical memory page and memory colors, the memory colors are mapped to the multiple physical memory pages to obtain memory channels corresponding to each shading partition, wherein a shading partition refers to a segment of equal-length and continuous physical address space that is mapped to the same memory channel and stored in a physical memory page; In response to detecting a target task that performs a video memory operation according to a required color and a required space, querying an unallocated shading partition of a video memory channel corresponding to the required color as a target shading partition; The target physical video memory page to which the target shading partition belongs is allocated from the multiple physical video memory pages to the target task according to the demand space, and the virtual address space corresponding to the target physical video memory page is returned to the target task for use.
2. The method according to claim 1, wherein: A physical video memory page includes at least two shading partitions, each shading partition has a number, the number represents the offset size of the shading partition relative to the starting address of the physical video memory page, and each video memory channel includes the same number of shading partitions; The allocating the target physical video memory page to which the target shading partition belongs from the plurality of physical video memory pages to the target task according to the required space includes: Count the number of unassigned coloring partitions in the coloring partitions corresponding to each coloring partition number of the required color; According to the shading partition number with the largest total number, a physical video memory page to which a target shading partition that meets the required space belongs is allocated from the unallocated shading partitions to the target task.
3. The method according to claim 1, wherein: The method further comprises: Hijacking a request to call an API of a native dynamic link library through a custom dynamic link library, wherein each API in the native dynamic link library has an API declaration with the same name in the custom dynamic link library; In response to detecting that the hijacked request involves a video memory operation, the hijacked request is forwarded to a custom API, wherein the custom API converts the video memory operation for the continuous physical address space into the video memory operation for the specified shading physical address space, wherein the video memory operation includes at least one of the following: allocating video memory space, releasing video memory space, copying data in the video memory space, and resetting data in the video memory space.
4. The method according to claim 1, wherein: The method further comprises: In response to detecting that the GPU is disconnected from the host, the plurality of physical video memory pages are released.
5. The method according to claim 1, wherein: The method further comprises: In response to a hijacked request to call a GPU side function, the hijacked request is forwarded to a custom GPU side function, wherein the custom GPU side function adds an instruction to translate a memory access offset, wherein the instruction to translate the memory access offset maps the memory access address of the GPU side function from a continuous virtual address space to a discrete video memory shading partition.
6. The method according to claim 5, wherein: The custom GPU side function is generated by the following steps: Check whether there is source code of GPU-side function; In response to detecting the existence of the source code of the GPU side function, adding an instruction for translating the memory access offset to the source code to obtain a custom GPU side function; In response to detecting that the source code of the GPU side function does not exist, a binary file of the GPU side function is obtained, the binary file is disassembled to obtain the source code, and instructions for translating the memory access offset are added to the source code to obtain a custom GPU side function.
7. The method according to claim 1, wherein: The querying of the unallocated shading partition of the video memory channel corresponding to the required color as the target shading partition includes: querying the number of unallocated shading partitions of the video memory channel corresponding to the required color in a structure recording the plurality of physical video memory pages; and The method further comprises: Updates the variables in the structure according to the assigned target shading partition.
8. The method according to claim 7, wherein: The structure Includes the following variables: starting physical address, ending physical address, number of unallocated shading partitions, total number of shading partitions, allocated color numbers, allocated partition numbers, and a linked list for storing unallocated physical video memory pages.
9. A GPU memory channel allocation device, comprising: The initialization unit is configured to pre-apply for a plurality of physical video memory pages; A shading unit is configured to map the video memory colors to the plurality of physical video memory pages according to a mapping table between shading partitions of the physical video memory pages and video memory colors, so as to obtain video memory channels corresponding to each shading partition, wherein a shading partition refers to a segment of equal-length and continuous physical address space that is mapped to the same video memory channel and stored in one physical video memory page; A query unit is configured to query an unallocated shading partition of a video memory channel corresponding to the required color as a target shading partition in response to detecting a target task that performs a video memory operation according to a required color and a required space; The allocation unit is configured to allocate the target physical video memory page belonging to the target shading partition from the multiple physical video memory pages to the target task according to the required space, and return the virtual address space corresponding to the target physical video memory page to the target task for use.
10. An electronic device comprising: one or more processors; a storage device having one or more computer programs stored thereon, When the one or more computer programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 8.
11. A computer readable medium having a computer program stored thereon, wherein: When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
12. A computer program product, comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 8.
Citation Information
Cited By
Equipment fault processing method, electronic equipment, storage medium and program product
CN120256187A
Video memory management method and device, storage medium and program product
CN120510022A
Video memory management method and device, storage medium and program product
CN120510022B
GPU memory channel allocation method and apparatus, device, and storage medium
WO2026148822A1