A register resource management method, device and medium for GPU

By building a computed parameter mapping binary tree in the GPU and optimizing the storage of GPU initialization parameters, the problem of thread parallelism reduction caused by GPU register resource limitation is solved, and efficient memory utilization and register resource management are achieved.

CN119477664BActive Publication Date: 2025-06-17SHANDONG INSPUR SCI RES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510067082.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-06-17
Estimated Expiration
2045-01-16

AI Technical Summary

Technical Problem

When the GPU processes computing tasks, due to register resource limitations, the thread parallelism decreases, which affects the computing performance.

Method used

By applying for memory space by a pre-compiled first executable file, and building a computed parameter mapping binary tree, optimizing the storage and access of GPU initialization parameters, reducing direct dependence on register resources.

Benefits of technology

It improves memory utilization efficiency, reduces memory fragmentation, enhances register usage flexibility, maintains high thread parallelism, and meets the needs of complex computing tasks for register resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119477664B_ABST
    Figure CN119477664B_ABST
Patent Text Reader

Abstract

The present application provides a register resource management method, device, and medium for a GPU, belonging to the technical field of electrical digital data processing. The method can apply to the GPU device for a first memory space corresponding to an actual computing task based on a pre-compiled first executable file; the first memory space is used to store corresponding computing parameters and a second executable file; the second executable file is used for the GPU device to execute an actual computing task based on the computing parameters and preset registers; generate a computing parameter mapping binary tree according to the memory address of the computing parameters in the first memory space, and write the root node pointer of the computing parameter mapping binary tree into the GPU initialization parameter structure of the GPU device to obtain corresponding GPU initialization parameters; write the memory start address where the GPU device stores the GPU initialization parameters into the preset register so that the GPU device executes the second executable file.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of electronic digital data processing, and in particular, to a method, device, and medium for managing register resources of a Graphics Processing Unit (GPU). Background Art

[0002] With the continuous improvement of computing requirements, the GPU plays an important role in many fields by virtue of the characteristic of using a large number of threads to improve the parallelism of operations. However, its performance is restricted by storage resources.

[0003] Specifically, taking the NVIDIA V100 GPU as an example, its hardware specifications limit that each Streaming Multiprocessor (SM) can control a maximum register size of 256KB, and each SM allows a maximum of 2048 threads to run simultaneously. After conversion, it is equivalent that each thread can occupy at most 32 32-bit register resources.

[0004] The inventors found that in actual operations, if the computing tasks processed by each thread require more registers, the number of threads running simultaneously has to be reduced, which in turn leads to a reduction in the number of schedulable threads, and ultimately reduces the parallelism of threads, affecting the full play of the overall computing performance of the GPU. Summary of the Invention

[0005] Embodiments of this application provide a method, device, and medium for managing register resources of a GPU, which are used to solve the technical problem that threads unreasonably consume register resources, resulting in a reduction in the thread parallelism of the system, in an operating environment where the GPU is restricted by storage resources.

[0006] On the one hand, embodiments of this application provide a method for managing register resources of a GPU, and the method includes:

[0007] Based on a pre-compiled first executable file, apply to the GPU device for a first memory space corresponding to an actual computing task; the first memory space is used to store corresponding computing parameters and a second executable file; the second executable file is used for the GPU device to execute the actual computing task based on the computing parameters and preset registers;

[0008] Generate a computing parameter mapping binary tree according to the memory address of the computing parameters in the first memory space, and write the root node pointer of the computing parameter mapping binary tree into the GPU initialization parameter structure of the GPU device to obtain corresponding GPU initialization parameters;

[0009] Write the memory start address where the GPU device stores the GPU initialization parameters into the preset register, so that the GPU device executes the second executable file; wherein, the preset register is used to transfer pre-stored parameters when the GPU device executes a kernel function corresponding to the actual computing task.

[0010] In one implementation manner of the present application, applying for a first memory space corresponding to an actual computing task from a GPU device based on a pre-compiled first executable file specifically includes:

[0011] Determine each computing parameter corresponding to the actual computing task and the second executable file through the first executable file; wherein, the parameter types of the computing parameters include one or more of the following: source operation data, destination operation data;

[0012] Generate a memory application request according to the number of parameters of the computing parameters and the second executable file, and send the memory application request to the GPU device, so that the GPU device allocates video memory according to the memory application request to generate the first memory space;

[0013] After applying for a first memory space corresponding to an actual computing task from a GPU device based on a pre-compiled first executable file, the method further includes:

[0014] Store the source operation data and the second executable file into the first memory space, and record the memory address corresponding to the data stored in the first memory space.

[0015] In one implementation manner of the present application, generating a computing parameter mapping binary tree according to the memory address of the computing parameters in the first memory space, and writing the root node pointer of the computing parameter mapping binary tree into a GPU initialization parameter structure body of the GPU device to obtain corresponding GPU initialization parameters specifically includes:

[0016] Construct a binary tree data structure according to the parameter type and the number of parameters corresponding to the computing parameters;

[0017] Store each memory address of the computing parameters in the first memory space into corresponding nodes of the binary tree data structure to generate a computing parameter mapping binary tree;

[0018] Write the root node pointer of the computing parameter mapping binary tree into a preset position of the GPU initialization parameter structure body to obtain the GPU initialization parameters.

[0019] In one implementation manner of the present application, writing the memory start address where the GPU device stores the GPU initialization parameters into the preset register specifically includes:

[0020] After obtaining the GPU initialization parameters, apply for a second memory space from the GPU device;

[0021] Write the GPU initialization parameters into the second memory space, and obtain the memory start address of the second memory space;

[0022] Write the memory start address into the preset register.

[0023] In an implementation manner of the present application, the GPU device executing the second executable file specifically includes:

[0024] The GPU device reads the parameters pre-stored in the preset register through a pre-written startup program file to determine the memory start address of the GPU initialization parameters;

[0025] The GPU device uses the memory start address as the base address to read the corresponding preset offset address to determine the root node pointer; wherein, the preset offset address read with the memory start address as the base address is the writing position of the GPU initialization parameters;

[0026] The GPU device traverses the calculation parameter mapping binary tree according to the root node pointer and the parameter type and number of the actual calculation task, obtains the node data values corresponding to each binary tree node, and based on the correspondence between the binary tree node and the calculation parameter, determines the calculation parameters corresponding to each node data value, and then executes the actual calculation task.

[0027] In an implementation manner of the present application, generating a calculation parameter mapping binary tree according to the memory address of the calculation parameter in the first memory space specifically includes:

[0028] Determine the memory addresses corresponding to the source operation data and the second executable file in the first memory space respectively, and add the corresponding memory addresses to the first linked list in the memory application order;

[0029] Add the memory address corresponding to the destination operation data in the first memory space to the second linked list according to the memory application order;

[0030] Based on the spatial locality principle of the program and the second executable file in the first linked list, determine the root node data of the calculation parameter mapping binary tree, and use the memory addresses corresponding to the remaining source operation data in the first linked list as the left subtree node data and add them to the binary tree in sequence;

[0031] Use the memory address of the target operation data in the second linked list as the right subtree node data, and add it to the binary tree in sequence to obtain the calculation parameter mapping binary tree; wherein, the number of binary tree nodes of the calculation parameter mapping binary tree is equal to the number of parameters of the calculation parameter.

[0032] In an implementation manner of the present application, based on the principle of spatial locality of the program and the second executable file in the first linked list, determine the root node data of the calculation parameter mapping binary tree, and use the memory addresses corresponding to the remaining source operation data in the first linked list as the left subtree node data, and add them to the binary tree in sequence, specifically including:

[0033] Based on the principle of spatial locality of the program, determine the memory address of the second executable file in the first linked list;

[0034] Determine the memory address with the smallest storage address distance from the memory address of the second executable file in the first linked list as the root node data, and remove the memory address of the second executable file and the root node data from the first linked list; wherein, the storage address distance is the difference between the two memory addresses;

[0035] Determine the memory address with the smallest storage address distance from the root node data in the first linked list as the first left subtree node data, and remove the first left subtree node data from the first linked list until the first linked list is empty, to obtain the root node of the calculation parameter mapping binary tree and each left subtree node; wherein, the first left subtree node data is the node data of the left subtree node of the root node.

[0036] In an implementation manner of the present application, when the GPU device executes the second executable file, traverse the calculation parameter mapping binary tree through the preorder traversal algorithm.

[0037] On the other hand, an embodiment of the present application further provides a register resource management device for a GPU, and the device includes:

[0038] At least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a register resource management method for a GPU as described above.

[0039] On yet another aspect, an embodiment of the present application further provides a non-volatile computer storage medium storing computer executable instructions that can execute a register resource management method for a GPU as described above.

[0040] Compared with the prior art, the remarkable effects of this application are as follows:

[0041] (1) This application applies for memory space through a pre-compiled first executable file and constructs a computational parameter mapping binary tree, achieving fast access and efficient management of computational parameters. At the same time, the memory start address of the GPU initialization parameters is written into a preset register, enabling the GPU device to flexibly access the required parameters when executing the kernel function. This design not only improves memory utilization efficiency, reduces memory fragmentation, but also significantly enhances the flexibility of register usage, reduces the direct dependence of computational tasks on register resources, and enables the GPU to meet the requirements of complex computational tasks for register resources while maintaining a high degree of thread parallelism.

[0042] (2) The first memory space of the GPU device stores computational parameters and a second executable file, which helps to ensure that the GPU device has sufficient video memory resources before executing computational tasks, avoiding problems such as resource shortages or improper allocation. And the binary tree structure helps to quickly search for and access computational parameters, improving computational efficiency. Furthermore, in a computational environment where the GPU is restricted by storage resources, it solves the technical problem that threads unreasonably consume register resources, resulting in a reduction in the thread parallelism of the system. Description of the Drawings

[0043] The drawings described herein are used to provide a further understanding of this application, form a part of this application, and the illustrative embodiments of this application and their descriptions are used to explain this application, and do not constitute an improper limitation to this application. In the drawings:

[0044] Figure 1 is a schematic flowchart of a method for managing register resources for a GPU in an embodiment of this application;

[0045] Figure 2 is a schematic system structure diagram corresponding to a method for managing register resources for a GPU in an embodiment of this application;

[0046] Figure 3 is a schematic diagram of the GPU memory in a method for managing register resources for a GPU in an embodiment of this application;

[0047] Figure 4 is a schematic diagram of a second linked list in a method for managing register resources for a GPU in an embodiment of this application;

[0048] Figure 5 is a schematic diagram of a computational parameter mapping binary tree in a method for managing register resources for a GPU in an embodiment of this application;

[0049] Figure 6Schematic diagram of the program design process of an executable file in a register resource management method for a GPU in an embodiment of the present application;

[0050] Figure 7 Schematic diagram of the structure of a register resource management device for a GPU in an embodiment of the present application. Detailed implementation manners

[0051] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0052] The heterogeneous programming model divides the executable program into a device-side (GPU) program and a host-side program, and the host side is the Central Processing Unit (CPU). In currently mainstream heterogeneous programming frameworks such as Compute Unified Device Architecture (CUDA), Open Computing Language (OpenCL), etc., the designed programming model requires the host-side program to explicitly apply for device-side memory and write the data to be processed and the bin file executed on the device side into the applied device-side memory; then write the addresses and data lengths of each parameter in the device-side memory into the GPU registers, and each thread completes the execution of the calculation tasks in the bin file according to the data passed into the registers. When the number of parameters required for the device-side calculation tasks is too large, a large amount of register resources need to be occupied.

[0053] Embodiments of the present application provide a register resource management method, device and medium for a GPU, which are used to solve the technical problem that threads unreasonably consume register resources, resulting in a reduction in the thread parallelism of the system, in an operation environment where the GPU is restricted by storage resources.

[0054] The following will describe each embodiment of the present application in detail with reference to the drawings.

[0055] Embodiments of the present application provide a register resource management method for a GPU, which is applied to an electronic device including a GPU and a CPU, such as Figure 1 As shown, the method may include steps S101-S103:

[0056] S101. Apply for a first memory space corresponding to an actual computing task from the GPU device based on a pre-compiled first executable file.

[0057] Among them, the above-mentioned first memory space is used to store corresponding computing parameters and a second executable file; the second executable file is used for the GPU device to execute the actual computing task based on the computing parameters and preset registers.

[0058] It should be noted that the register resource management method for the GPU in this application is described with the execution entity being the CPU. The execution entity is not limited to the CPU and can also be other types of devices. This application does not make specific limitations in this regard.

[0059] The first executable file is generated by compiling the host program using a compiler for the CPU host side. When GPU participation is required for an actual computing task, the CPU will execute the first executable file to complete steps S101 - S103. When the CPU executes the first executable file, it will configure GPU parameters. The CPU can use the GPU driver application programming interface (Application Programming Interface, API), such as (CUDA Runtime API or OpenCL API), to apply for memory on the GPU device and generate a first memory space in the DDR of the video memory of the GPU device. Among them, the schematic diagram of the system structure composed of the host side (CPU) and the device side (GPU device) is as Figure 2 shown. The device side includes a GPU and DDR, and the host side is connected to the device side through a PCIe bus.

[0060] In the embodiment of this application, the above-mentioned step of applying for a first memory space corresponding to an actual computing task from the GPU device based on a pre-compiled first executable file specifically includes:

[0061] Determine each computing parameter and the second executable file corresponding to the actual computing task through the first executable file. Among them, the parameter types of the computing parameters include one or more of the following: source operation data, destination operation data. Generate a memory application request according to the number of parameters of the computing parameters and the second executable file, and send it to the GPU device, so that the GPU device allocates video memory according to the memory application request and generates a first memory space.

[0062] In other words, the CPU can determine the calculation parameters involved in the actual calculation task. The calculation parameters include source operation data, destination operation data, and the parameter types of their combination. For example, for the actual calculation task of the vector addition kernel function: the vector addition kernel function vecadd(__global const uint32_t *a, __global const uin32_t *b, __global uint32_t *c), this function includes two source operation data and one destination operation data, with 2 parameter types and 3 parameter quantities. At the same time, the CPU can also determine the second executable file of the pre-compiled GPU, which can be understood as the GPU device-side bin file and can enable the GPU to execute the actual calculation task in a way that reasonably consumes register resources. Subsequently, the CPU will apply for memory in the GPU device according to the number of calculation parameters and the second executable file, so that the applied memory can store the calculation parameters and the second executable file.

[0063] Generally understood, the source operation data is the data required for the GPU to execute the actual calculation task, and the destination operation data is the data obtained or generated after the actual calculation task is executed.

[0064] In another embodiment of the present application, after applying for the first memory space corresponding to the actual calculation task from the GPU device based on the pre-compiled first executable file, the method further includes:

[0065] Store the source operation data and the second executable file into the first memory space, and record the memory address corresponding to the stored data in the first memory space.

[0066] That is to say, when the CPU enables the GPU to execute the actual calculation task, it needs to first write the source operation data and the second executable file into the first memory space. As Figure 3 shown in the GPU memory, the code segment from start to kernel_func to end corresponds to the memory address of the second executable file, the kernel parameter node data corresponds to the binary tree node of the calculation parameter mapping binary tree of the present application, and the kernel parameter root node data corresponds to the binary tree root node. Combinations of numbers and letters such as 0x9000_3000, 0x9000_3004, 0x9000_3008 in the figure are memory addresses.

[0067] In addition, the destination operation data also has a memory address on the device side (GPU). For example, for the addition calculation task, which includes the destination operation data and can be understood as the calculation result of the addition, a memory address needs to be allocated in advance for storage.

[0068] For example, the above vector addition kernel function requires a total of three operands. The CPU applies for memory from the device side, and the storage addresses obtained for the three parameters are arg1_addr, arg2_addr, and arg3_addr respectively.

[0069] S102. Generate a calculation parameter mapping binary tree according to the memory address of the calculation parameter in the first memory space, and write the root node pointer of the calculation parameter mapping binary tree into the GPU initialization parameter structure of the GPU device to obtain the corresponding GPU initialization parameter.

[0070] In the embodiment of the present application, generating a calculation parameter mapping binary tree according to the memory address of the calculation parameter in the first memory space and writing the root node pointer of the calculation parameter mapping binary tree into the GPU initialization parameter structure of the GPU device to obtain the corresponding GPU initialization parameter specifically includes:

[0071] Construct a binary tree data structure according to the parameter type and parameter quantity corresponding to the calculation parameter. Store each memory address of the calculation parameter in the first memory space into the corresponding node of the binary tree data structure to generate a calculation parameter mapping binary tree. Write the root node pointer of the calculation parameter mapping binary tree into a preset position of the GPU initialization parameter structure to obtain the GPU initialization parameter.

[0072] In other words, the CPU can determine the memory addresses allocated for all calculation parameters in the first memory space, and then construct a binary tree according to the parameter type and parameter quantity, and generate binary tree nodes for the above source operation data and destination operation data respectively. The node data in the binary tree node is the memory address allocated for the calculation parameter. For example, the CPU can preset a correspondence list between the parameter type and parameter quantity and the binary tree node, such as the parameter type being the source operation data as the left subtree node, the destination operation data as the right subtree node, and the number of binary tree nodes being equal to the correspondence list of the parameter quantity of the calculation parameter. Construct a binary tree through this correspondence list to obtain a calculation parameter mapping binary tree. Subsequently, write the root node pointer of this calculation parameter mapping binary tree to a preset position of the GPU initialization parameter structure. For Figure 3 example, the root node pointer of 0x9000_3000, that is, the storage address of the root node, is written into the GPU initialization parameter. Among them, the preset position can be set during actual use, and the present application does not make specific limitations on this.

[0073] In another embodiment of the present application, generating a calculation parameter mapping binary tree according to the memory address of the calculation parameter in the first memory space specifically includes:

[0074] Determine the memory addresses corresponding to the source operation data and the second executable file in the first memory space, and add the corresponding memory addresses to the first linked list in the order of memory application. Add the memory address corresponding to the destination operation data in the first memory space to the second linked list in the order of memory application. Based on the principle of spatial locality of the program and the second executable file in the first linked list, determine the root node data of the calculation parameter mapping binary tree, and use the memory addresses corresponding to the remaining source operation data in the first linked list as the left subtree node data, and add them to the binary tree in sequence. Use the memory addresses of the destination operation data in the second linked list as the right subtree node data, and add them to the binary tree in sequence to obtain the calculation parameter mapping binary tree. Among them, the number of binary tree nodes of the calculation parameter mapping binary tree is equal to the number of parameters of the calculation parameters.

[0075] That is to say, the present application can group the memory addresses corresponding to the above-mentioned source operation data and the second executable file to generate a first linked list, and group the memory addresses corresponding to the destination operation data to generate a second linked list (as Figure 4 shown). Among them, when generating the linked list, it can be arranged in the order of memory address application, or in other orders, and the present application does not make specific limitations on this.

[0076] The present application searches for the root node data according to the principle of spatial locality of the program, then generates the left subtree nodes of the binary tree through the first linked list, obtains the right subtree nodes through the second linked list, and further constructs the calculation parameter mapping binary tree.

[0077] More specifically, the above-mentioned method for determining the root node data of the calculation parameter mapping binary tree based on the principle of spatial locality of the program and the second executable file in the first linked list, and using the memory addresses corresponding to the remaining source operation data in the first linked list as the left subtree node data and adding them to the binary tree in sequence specifically includes:

[0078] Based on the principle of spatial locality of the program, determine the memory address of the second executable file in the first linked list. Determine the memory address with the smallest storage address distance from the memory address of the second executable file in the first linked list as the root node data, and remove the memory address of the second executable file and the root node data from the first linked list. Among them, the storage address distance is the difference between two memory addresses. Determine the memory address with the smallest storage address distance from the root node data in the first linked list as the first left subtree node data, and remove the first left subtree node data from the first linked list until the first linked list is empty, to obtain the root node and each left subtree node of the calculation parameter mapping binary tree. Among them, the first left subtree node data is the node data of the left subtree node of the root node.

[0079] That is to say, according to the principle of spatial locality of the program, this application takes the memory address of the second executable file as a reference, first calculates the memory address in the first linked list that has the smallest storage address distance from this reference memory address, that is, searches for the source operation data with the closest storage address to the storage address of the second executable file. For example, if the memory address of the second executable file is addr_bin in the first linked list, and the memory address of the source operation data with the smallest storage address distance from it is addr_src_7 in the first linked list, then the memory address corresponding to this addr_src_7 is used as the root node data to create a binary tree. Then, addr_bin and addr_src_7 in the first linked list are deleted from the first linked list List_order_SRC. Immediately afterwards, search for the remaining source operation data in the first linked list, and find the memory address with the smallest storage address distance from addr_src_7, such as addr_src_5 in the first linked list, to obtain the left subtree node of the root node, and remove addr_src_5 from the first linked list; then find the memory address with the smallest storage address distance from addr_src_5... until the first linked list is empty, obtaining the root node and the left subtree of the binary tree.

[0080] According to the above execution process, search the second linked list, so as to use each memory address in the second linked list as the right subtree node. For example, read the second linked list List_order_DST, and insert addr_dst_1 corresponding to the memory address of the destination operation data in the linked list as the right subtree node of the binary tree root node into the binary tree. And delete addr_dst_1 from List_order_DST, and repeat this step until the List_order_DST linked list is empty. At this time, all the memory addresses of the destination operation data have been stored in the binary tree. The generated binary tree is as Figure 5 shown.

[0081] In the embodiment of this application, after constructing the calculation parameter mapping binary tree according to the above binary tree generation rule, in order to improve the efficiency and calculation speed of the GPU traversing the binary tree, when the GPU device of this application executes the second executable file, it traverses the above calculation parameter mapping binary tree through the preorder traversal algorithm. Thus, the source operation data and the destination operation data can be obtained efficiently and orderly.

[0082] S103, write the memory start address where the GPU device stores the GPU initialization parameters into a preset register, so that the GPU device executes the second executable file.

[0083] Among them, the preset register is used to transfer the pre-stored parameters when the GPU device executes the kernel function corresponding to the actual calculation task. For example, the preset memory transfers the memory start address of the GPU initialization parameters.

[0084] In the embodiment of the present application, writing the memory start address storing the GPU initialization parameters of the GPU device into a preset register specifically includes:

[0085] After obtaining the GPU initialization parameters, apply for a second memory space from the GPU device. Write the GPU initialization parameters into the second memory space, and obtain the memory start address of the second memory space. Write the memory start address into the preset register.

[0086] In other words, after the present application obtains the GPU initialization parameters with the root node pointer written therein, the CPU can further apply for a second memory space from the GPU device to store the GPU initialization parameters. In addition, the present application also configures other conventional startup parameters, such as grid and block sizes, to start the GPU kernel function. The present application writes the memory start address of the second memory space storing the GPU initialization parameters into a preset register. The memory start address is, for example, Figure 3 0x9002_4000 as in the above. Write the GPU initialization parameters to the position of 0x9002_4000, and write 0x9002_4000 into the gpu_args register. Among them, when the GPU executes the assembly startup program in the second executable file, it will first read the preset gpu_args register to obtain the content of the GPU initialization parameter structure.

[0087] Further, the GPU device executing the second executable file specifically includes:

[0088] The GPU device reads the parameters pre-stored in the preset register through a pre-written startup program file to determine the memory start address of the GPU initialization parameters. The GPU device uses the memory start address as the base address to read the corresponding preset offset address to determine the root node pointer. Among them, the preset offset address read with the memory start address as the base address is the writing position of the GPU initialization parameters. The GPU device traverses the computational parameter mapping binary tree according to the root node pointer and the parameter type and number of the actual computational task to obtain the node data values corresponding to each binary tree node. After determining the computational parameters corresponding to each node data value based on the correspondence between the binary tree node and the computational parameter, the GPU device executes the actual computational task.

[0089] That is to say, the GPU device determines the calculation parameters corresponding to the actual calculation task through the second executable file, including the parameter type and the number of parameters. Execute the startup program file such as start.S. By reading the data in the gpu_agrs register, the storage address of the GPU initialization parameters is obtained as 0x9002_4000. Then, taking this storage address as the base address, read the preset offset address 0x9002_4004. This preset offset address corresponds to the root node pointer of the calculation parameter mapping binary tree, and the memory address of the root node pointer is obtained as 0x9000_3000. Then, the GPU device performs a traversal operation on the calculation parameter mapping binary tree through the preorder traversal algorithm to obtain the data values of each node. At the same time, the GPU of this application presets the corresponding relationship between the binary tree nodes and the calculation parameters. For example, the root node and the left subtree node are source operation data, and the right subtree node is destination operation data, so as to obtain each calculation parameter. Furthermore, the GPU kernel function executes the actual calculation task according to the obtained calculation parameters.

[0090] Further, after waiting for the GPU to complete the calculation task and read the output data, the CPU can release the memory resources applied on the GPU device, so as to reasonably utilize the GPU register resources for subsequent calculation tasks.

[0091] Through the above solution, when the CPU host program runs, completes the GPU initialization parameter configuration and starts the GPU kernel function to run, during the execution of the assembly startup program of the GPU kernel function, it can address the root node of the binary tree storing the calculation parameters on the storage device through the base address of the GPU initialization parameter storage space, and obtain the storage addresses of all calculation parameters through traversal, realizing the acquisition of all calculation parameter data and passing it to the device-side program for calculation. Since the parameter addresses in this application are obtained through indirect addressing (i.e., binary tree traversal) instead of being directly stored in registers, it reduces the consumption of register resources by the calculation parameters. And more register resources in this application can be used for other purposes (such as temporary variables and index calculation), thereby improving the calculation efficiency of the GPU and enhancing the parallelism of threads.

[0092] This application applies for memory space through the pre-compiled first executable file and constructs a calculation parameter mapping binary tree, realizing the fast access and efficient management of calculation parameters. At the same time, writing the memory start address of the GPU initialization parameters into the preset register enables the GPU device to flexibly access the required parameters when executing the kernel function. This design not only improves the memory utilization efficiency, reduces memory fragmentation, but also significantly enhances the flexibility of register use, reduces the direct dependence of calculation tasks on register resources, and enables the GPU to meet the requirements of complex calculation tasks for register resources while maintaining a high thread parallelism.

[0093] In addition, the first memory space of the GPU device stores computing parameters and the second executable file, which helps ensure that the GPU device has sufficient video memory resources before executing computing tasks, avoiding problems such as resource shortages or improper allocation. Moreover, the binary tree structure helps to quickly search for and access computing parameters, improving computing efficiency. Furthermore, in an operating environment where the GPU is restricted by storage resources, it solves the technical problem that threads unreasonably consume register resources, resulting in a reduction in the thread parallelism of the system.

[0094] Figure 6 The flowchart of the programming method for separately generating the first executable file and the second executable file for the host side (CPU) and the device side (GPU device) of the embodiments of the present application is as Figure 6 shown:

[0095] Host side: S601, write the host side program to implement device side memory application, and write source operands and device side bin file data into the applied device side memory; S602, store the device side memory addresses of all computing parameters obtained by the application into a binary tree data structure. Write the pointer of the root node of the binary tree into the preset position of the GPU initialization parameter structure. S603, the host program applies for a memory space for storing GPU initialization parameters from the device side, writes the GPU initialization parameters into the applied device side memory, writes the first address of the applied memory into the register gpu_args specified by the device side. Then configure other conventional startup parameters and start the execution of the GPU kernel function. S604, use the compiler for the CPU host side to compile the host side program to generate the second executable file.

[0096] Device side: S701, write the GPU device side program, determine the parameter types and quantities of the source operands and destination operands participating in the calculation, and design the calculation task. S702, write the start.S assembly startup program, and implement the preorder traversal of the binary tree in the assembly program. According to the parameter types, parameter quantities determined in S701 and the pointer of the root node of the binary tree provided by the host side, obtain the corresponding number of source operation data and destination operation data from the parameter storage area. S703, use the compiler for the GPU device to compile the source files designed in S701 and S702 to generate the second executable file.

[0097] Figure 7 The structural schematic diagram of a register resource management device for GPU provided by the embodiments of the present application is as Figure 7 shown, and the device includes:

[0098] At least one processor; and a memory communicatively connected to the at least one processor. Wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can:

[0099] Based on a pre-compiled first executable file, apply to the GPU device for a first memory space corresponding to an actual computing task. The first memory space is used to store corresponding computing parameters and a second executable file. The second executable file is used for the GPU device to execute the actual computing task based on the computing parameters and preset registers. Generate a computing parameter mapping binary tree according to the memory address of the computing parameters in the first memory space, and write the root node pointer of the computing parameter mapping binary tree into the GPU initialization parameter structure of the GPU device to obtain corresponding GPU initialization parameters. Write the memory start address where the GPU device stores the GPU initialization parameters into the preset register so that the GPU device can execute the second executable file. Among them, the preset register is used to transfer pre-stored parameters when the GPU device executes a kernel function corresponding to the actual computing task.

[0100] The embodiment of the present application also provides a non-volatile computer storage medium storing computer-executable instructions, and the computer-executable instructions are set as:

[0101] Based on a pre-compiled first executable file, apply to the GPU device for a first memory space corresponding to an actual computing task. The first memory space is used to store corresponding computing parameters and a second executable file. The second executable file is used for the GPU device to execute the actual computing task based on the computing parameters and preset registers. Generate a computing parameter mapping binary tree according to the memory address of the computing parameters in the first memory space, and write the root node pointer of the computing parameter mapping binary tree into the GPU initialization parameter structure of the GPU device to obtain corresponding GPU initialization parameters. Write the memory start address where the GPU device stores the GPU initialization parameters into the preset register so that the GPU device can execute the second executable file. Among them, the preset register is used to transfer pre-stored parameters when the GPU device executes a kernel function corresponding to the actual computing task.

[0102] The various embodiments in the present application are all described in a progressive manner. For the same or similar parts among the various embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device and medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the partial description of the method embodiments for the relevant parts.

[0103] The device and medium provided by the embodiment of the present application correspond one-to-one with the method. Therefore, the device and medium also have beneficial technical effects similar to those of the corresponding method. Since the beneficial technical effects of the method have been described in detail above, the beneficial technical effects of the device and medium will not be elaborated here.

[0104] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising said element.

[0105] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various modifications and changes can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A register resource management method for a GPU, characterized in that: The method comprises: Based on the pre-compiled first executable file, apply to the GPU device for a first memory space corresponding to the actual computing task; the first memory space is used to store corresponding computing parameters and a second executable file; the second executable file is used by the GPU device to execute the actual computing task based on the computing parameters and preset registers; Generate a computing parameter mapping binary tree according to the memory address of the computing parameter in the first memory space, and write the root node pointer of the computing parameter mapping binary tree into the GPU initialization parameter structure of the GPU device to obtain the corresponding GPU initialization parameter; The first memory address of the GPU device storing the GPU initialization parameters is written into the preset register so that the GPU device executes the second executable file; wherein the preset register is used to pass the pre-stored parameters when the GPU device executes the kernel function corresponding to the actual computing task.

2. A register resource management method for a GPU according to claim 1, characterized in that: Based on the pre-compiled first executable file, applying to the GPU device for a first memory space corresponding to the actual computing task specifically includes: Determine, through the first executable file, various computing parameters corresponding to the actual computing task and the second executable file; wherein the parameter type of the computing parameter includes one or more of the following: source operation data, destination operation data; Generate a memory request according to the number of parameters of the calculation parameter and the second executable file, and send the request to the GPU device, so that the GPU device allocates video memory according to the memory request to generate the first memory space; After applying for a first memory space corresponding to an actual computing task from a GPU device based on a pre-compiled first executable file, the method further includes: The source operation data and the second executable file are stored in the first memory space, and the memory address corresponding to the data stored in the first memory space is recorded.

3. The register resource management method for a GPU according to claim 1, characterized in that: Generate a computing parameter mapping binary tree according to the memory address of the computing parameter in the first memory space, and write the root node pointer of the computing parameter mapping binary tree into the GPU initialization parameter structure of the GPU device to obtain the corresponding GPU initialization parameter, specifically including: Constructing a binary tree data structure according to the parameter type and parameter quantity corresponding to the calculation parameters; Storing each of the memory addresses of the calculation parameters in the first memory space into corresponding nodes of the binary tree data structure to generate a calculation parameter mapping binary tree; The root node pointer of the computing parameter mapping binary tree is written into a preset position of the GPU initialization parameter structure to obtain the GPU initialization parameter.

4. The method for managing register resources for a GPU according to claim 1, characterized in that: Writing the first memory address of the GPU device storing the GPU initialization parameters into the preset register specifically includes: After obtaining the GPU initialization parameters, applying for a second memory space from the GPU device; Writing the GPU initialization parameters into the second memory space, and obtaining the memory first address of the second memory space; The memory first address is written into the preset register.

5. The method for managing register resources for a GPU according to claim 1, characterized in that: The GPU device executes the second executable file, specifically including: The GPU device reads the parameters pre-stored in the preset register through a pre-written startup program file to determine the memory first address of the GPU initialization parameter; The GPU device uses the memory first address as a base address and reads a corresponding preset offset address to determine the root node pointer; wherein the preset offset address read using the memory first address as a base address is a write position of the GPU initialization parameter; The GPU device traverses the computing parameter mapping binary tree according to the root node pointer and the parameter type and parameter quantity of the actual computing task, obtains the node data value corresponding to each binary tree node, and determines the computing parameter corresponding to each node data value based on the correspondence between the binary tree nodes and the computing parameters, and then executes the actual computing task.

6. The method for managing register resources for a GPU according to claim 2, characterized in that: Generating a computing parameter mapping binary tree according to the memory address of the computing parameter in the first memory space specifically includes: Determine the memory addresses corresponding to the source operation data and the second executable file in the first memory space, and add the corresponding memory addresses to the first linked list according to the memory application order; Adding the memory address corresponding to the target operation data in the first memory space to the second linked list according to the memory application order; Based on the spatial locality principle of the program and the second executable file in the first linked list, the root node data of the calculation parameter mapping binary tree is determined, and the memory addresses corresponding to the remaining source operation data in the first linked list are used as left subtree node data and added to the binary tree in sequence; The memory address of the target operation data in the second linked list is added to the binary tree in sequence as the right subtree node data to obtain the calculation parameter mapping binary tree; wherein the number of binary tree nodes of the calculation parameter mapping binary tree is equal to the number of parameters of the calculation parameters.

7. A register resource management method for a GPU according to claim 6, characterized in that: Based on the spatial locality principle of the program and the second executable file in the first linked list, determining the root node data of the calculation parameter mapping binary tree, and using the memory addresses corresponding to the remaining source operation data in the first linked list as left subtree node data, and sequentially adding them to the binary tree, specifically includes: Based on the spatial locality principle of the program, determining the memory address of the second executable file in the first linked list; Determine, in the first linked list, a memory address with the shortest storage address distance from the memory address of the second executable file as the root node data, and remove the memory address of the second executable file and the root node data from the first linked list; wherein the storage address distance is the difference between the two memory addresses; Determine the memory address in the first linked list that has the shortest distance from the storage address of the root node data, which is the first left subtree node data, and remove the first left subtree node data from the first linked list until the first linked list is empty, thereby obtaining the root node and each left subtree node of the calculation parameter mapping binary tree; wherein the first left subtree node data is the node data of the left subtree node of the root node.

8. The method for managing register resources for a GPU according to claim 7, characterized in that: When the GPU device executes the second executable file, the calculation parameter mapping binary tree is traversed by a pre-order traversal algorithm.

9. A register resource management device for a GPU, characterized in that: The device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a register resource management method for a GPU as described in any one of claims 1 to 8.

10. A non-volatile computer storage medium storing computer executable instructions, characterized in that: The computer executable instructions can execute a register resource management method for a GPU as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Resource configuration method and device, electronic equipment and computer readable storage medium

    CN114398172A

  • Image classification method based on improved OfficientNetV2 model, electronic equipment and readable storage medium

    CN118941840A