Memory space allocation method and device for thread, equipment and medium
By splitting and alternating the thread beam memory address offset, the problem of low GPU prefetching accuracy is solved, and more efficient memory management and performance improvement is achieved.
Patent Information
- Application Number
- CN202510528954.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-07-22
AI Technical Summary
The existing GPU prefetch strategy has the problem of low prefetch accuracy, especially the prefetched memory in the neighborhood tree algorithm is independent of the thread being used, resulting in performance impact.
By sending the memory request amount of thread bundle to the host side, splitting the memory request amount of threads, and alternately determining the memory address offset according to the offset arrangement order between threads and thread fragment request amount, ensuring that the memory space of threads with similar memory fetch modes is alternately arranged, improving the prefetch accuracy.
It improves the accuracy of the GPU's memory prefetching, reduces the waiting time and overhead of memory management, and improves the memory sharing performance between the host and the GPU.
Smart Images

Figure CN120353598A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of memory allocation, and particularly to a method, device, equipment and medium for allocating memory space for threads. Background Art
[0002] The unified virtual memory space provides a transparent operation for memory management between the GPU side and the host side for users through the page replacement mechanism, simplifying the resource sharing and access between the host side and the GPU side. However, this memory management mechanism affects performance due to the huge overhead brought by page replacement during cross-device access. To alleviate this problem, current GPUs generally adopt a prefetching strategy. When handling page faults, the data to be accessed is predicted and pre-loaded into the local memory in advance, thereby reducing the frequency of page faults.
[0003] The prefetching strategy adopted by current GPUs is based on the Tree-based Neighborhood algorithm. In the Tree-based Neighborhood algorithm, each node represents the usage of a memory range; the unused memory part in the nodes where the prefetch memory usage rate exceeds a certain threshold is prefetched. However, the memory prefetched in this way is often irrelevant to the threads being used, and there is a defect of low prefetch accuracy. Summary of the Invention
[0004] The present invention provides a method, device, equipment and medium for allocating memory space for threads to improve the prefetch accuracy of GPU memory.
[0005] In a first aspect, an embodiment of the present invention provides a method for allocating memory space for threads, which is executed by the GPU side and includes:
[0006] Sending the memory request amounts of at least one warp to the host side, so that the host side allocates the starting addresses of memory spaces for each warp according to the memory request amounts of each warp and feeds back the starting addresses of memory spaces of each warp to the GPU side;
[0007] For each warp, splitting the memory request amount of each thread in the warp according to the minimum scheduling memory amount to obtain at least one thread fragment request amount for each thread;
[0008] Determining at least one memory address offset of each thread in the warp alternately according to the offset order between the threads in the warp and the thread fragment request amounts of the threads in the warp;
[0009] For each thread in the warp, allocating the starting address of the memory space of the warp and each memory address offset of the thread to the thread.
[0010] Second aspect, an embodiment of the present invention further provides a method for allocating memory space for threads, which is executed by the host side and includes:
[0011] Receiving the memory request amounts of at least one warp sent by the GPU side;
[0012] For each warp, splitting the memory request amount of the warp according to the pre-partitioned space memory amount to obtain at least one warp fragment request amount;
[0013] According to the warp fragment request amounts of each warp, allocating free pre-partitioned memory space for each warp fragment request amount of each warp;
[0014] For each warp, determining the starting address of the memory space of the warp according to the starting addresses of the pre-partitioned memory spaces of the warp and the warp fragment request amounts in the warp;
[0015] Sending the starting address of the memory space of the warp to the GPU side, so that the GPU side allocates the starting address of the memory space and the memory address offset for at least one thread in the warp according to the starting address of the memory space of the warp.
[0016] Third aspect, an embodiment of the present invention further provides a device for allocating memory space for threads, which is configured on the GPU side and includes:
[0017] A request sending module, configured to send the memory request amounts of at least one warp to the host side, so that the host side allocates the starting address of the memory space for each warp according to the memory request amounts of each warp and feeds back the starting address of the memory space of each warp to the GPU side;
[0018] A memory splitting module, configured to, for each warp, split the memory request amount of each thread in the warp according to the minimum scheduling memory amount to obtain at least one thread fragment request amount for each thread;
[0019] An offset determining module, configured to alternately determine at least one memory address offset for each thread in the warp according to the offset arrangement order between the threads in the warp and the thread fragment request amounts of the threads in the warp;
[0020] A memory space determining module, configured to, for each thread in the warp, allocate the starting address of the memory space of the warp and the memory address offsets of the thread to the thread.
[0021] Fourth aspect, an embodiment of the present invention further provides a device for allocating memory space for threads, which is configured on the host side and includes:
[0022] A receiving module, configured to receive the memory request amounts of at least one warp sent by the GPU side;
[0023] A splitting module, configured to split the memory request amount of each warp according to the pre-partitioned space memory amount for each warp, so as to obtain at least one warp fragment request amount;
[0024] An allocation module, configured to allocate free pre-partitioned memory spaces for each warp fragment request amount of each warp according to the warp fragment request amounts of each warp;
[0025] A determination module, configured to determine the starting address of the memory space of each warp according to the starting addresses of the pre-partitioned memory spaces of each warp and the warp fragment request amounts in each warp;
[0026] A sending module, configured to send the starting address of the memory space of the warp to the GPU side, so that the GPU side allocates the starting address of the memory space and the memory address offset for at least one thread in the warp according to the starting address of the memory space of the warp.
[0027] In a fifth aspect, an embodiment of the present invention further provides a memory space allocation device for threads, including:
[0028] At least one processor; and
[0029] A memory communicatively connected to the at least one processor; wherein
[0030] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor, so that the at least one processor can execute the memory space allocation method for threads provided in any embodiment of the present invention.
[0031] In a sixth aspect, an embodiment of the present invention further provides a computer-readable storage medium, and the computer-readable storage medium stores computer instructions, and the computer instructions are used to implement the memory space allocation method for threads provided in any embodiment of the present invention when executed by a processor.
[0032] The technical solution of the embodiment of the present invention splits the memory request amounts of the threads in a warp according to the minimum scheduling memory amount, so that the warp fragment request amounts are aligned with the minimum scheduling memory amount of the GPU; by alternately determining at least one memory address offset of each thread in the warp according to the offset order, the memory spaces of different threads with highly similar memory access patterns are alternately arranged, so that when the GPU processes a certain thread, it can completely prefetch the memory spaces of other threads with similar memory access patterns in the same warp through the minimum scheduling memory amount, improving the prefetch accuracy of the GPU.
[0033] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become readily apparent from the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0035] Figure 1 is a flowchart of a method for allocating memory space for a thread according to Embodiment 1 of the present invention;
[0036] Figure 2A is a flowchart of a method for allocating memory space for a thread according to Embodiment 2 of the present invention;
[0037] Figure 2B is a schematic structural diagram of an alternating arrangement of memory spaces for a thread according to Embodiment 2 of the present invention
[0038] Figure 3 is a flowchart of a method for allocating memory space for a thread according to Embodiment 3 of the present invention;
[0039] Figure 4A is a flowchart of a method for allocating memory space for a thread according to Embodiment 4 of the present invention;
[0040] Figure 4B is a schematic diagram showing the correspondence between the amount of second thread bundle fragments and full binary subtrees in a full binary tree according to Embodiment 4 of the present invention;
[0041] Figure 4C is a schematic diagram of the processing flow of a thread memory allocation framework according to Embodiment 4 of the present invention;
[0042] Figure 5 is a schematic structural diagram of a device for allocating memory space for a thread according to Embodiment 5 of the present invention;
[0043] Figure 6 is a schematic structural diagram of a device for allocating memory space for a thread according to Embodiment 6 of the present invention;
[0044] Figure 7 is a structural diagram of an electronic device for implementing the method for allocating memory space for a thread in the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] To enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solution in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0046] It should be noted that the terms "first" and "second" in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0047] In the technical solution of the embodiments of the present invention, the acquisition, storage, application, etc. of the information to be displayed and the like all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0048] Embodiment 1
[0049] Figure 1 The flowchart of a method for allocating memory space for a thread provided in Embodiment 1 of the present invention is applicable to the situation of allocating memory space for a thread. This method can be executed by a device for allocating memory space for a thread, and the device for allocating memory space for a thread can be implemented in the form of hardware and / or software and is specifically configured in an electronic device.
[0050] See Figure 1 The method for allocating memory space for a thread shown in the figure is executed by the GPU side and includes:
[0051] S101. Send the memory request amounts of at least one warp to the host side, so that the host side allocates the starting addresses of memory spaces for each warp according to the memory request amounts of each warp and feeds back the starting addresses of memory spaces for each warp to the GPU side.
[0052] S102. For each warp, split the memory request amount of each thread in the warp according to the minimum scheduling memory amount to obtain at least one thread fragment request amount for each thread.
[0053] S103. Determine, alternately, at least one memory address offset for each thread in the warp according to the arrangement order of the offsets between the threads in the warp and the thread fragment request amounts of the threads in the warp.
[0054] S104. For each thread in the warp, allocate the starting address of the memory space of the warp and the memory address offsets of the thread to the thread.
[0055] In this embodiment, a warp may include at least one thread; the threads in the same warp have similar memory access patterns. The host side may be the side including a CPU (Central Processing Unit) and the memory managed thereby. The memory request amount may be the amount of memory requested to be allocated. The memory space may be a Unified Virtual Memory Space (UVM Space). The minimum scheduled memory amount may be the minimum amount of memory scheduled by the GPU when performing a prefetch operation. The thread fragment request amount is the result of splitting the memory request amount of a thread.
[0056] The offset arrangement order may be the order of the alternating arrangement of the memory address offsets of the threads. For example, if there are 3 threads in a warp, namely thread A, thread B, and thread C; and the offset arrangement order is thread B, thread A, and thread C, then the memory offsets should be determined alternately for thread B, thread A, and thread C. Exemplarily, the memory request amount of thread A may be split into 3 thread fragment request amounts; the memory request amount of thread B may be split into 2 thread fragment request amounts; the memory request amount of thread C may be split into 2 thread fragment request amounts. Then, the finally determined memory address offsets of the threads, arranged in ascending order according to the size of the offsets, can be expressed as B m -A m -C m -B m -A m -C m -A m ; where B m represents the memory address offset of thread B; A m represents the memory address offset of thread A; C m represents the memory address offset of thread C.
[0057] Specifically, on the GPU side, when a thread requests memory space, through the set library function, the warp scheduler can use the underlying warp shared register method to reduce and merge the thread allocation requests of each thread in the warp into a unified warp memory allocation request, or merge the thread allocation requests of each thread in the warp into a unified warp memory allocation request in a serial manner; where the thread allocation request includes the memory request amount of the thread. By merging the thread allocation requests of each thread in the warp into a unified warp memory allocation request, a warp memory allocation request including the memory request amount of the warp is obtained. Among them, the memory request amount of the warp is the sum of the memory request amounts of each thread in the warp. Exemplarily, if there are 3 threads in a warp, thread A, thread B, and thread C, the memory request amount of thread A is 3M, the memory request amount of thread B is 4M, and the memory request amount of thread C is 4M, then the memory request amount of the merged warp is 11M.
[0058] Through the warp scheduler, by means of page replacement in the unified virtual memory space, at least one way including the warp memory request is sent to the host side to send the memory request amount of each warp to the host side; the host side receives the memory request amount of at least one warp sent by the GPU side; for each warp, according to the pre-partitioned space memory amount, the memory request amount of this warp is split to obtain at least one warp fragment request amount; according to the number and the warp fragment request amount of the warp fragment request amounts of each warp, free pre-partitioned memory space is allocated for each warp fragment request amount of each warp; for each warp, according to the starting address of each pre-partitioned memory space of this warp and each warp fragment request amount in this warp, the starting address of the memory space of this warp is determined; the starting address of the memory space of this warp is sent to the GPU side.
[0059] The GPU side receives the starting address of the memory space of the warp. For each thread in the warp, the memory request amount of this thread is split according to the minimum scheduling memory amount; if there is a split result less than the minimum scheduling memory amount, then the split result can be amplified, or the split result can be directly used as the thread fragment request amount; using a certain algorithm, according to the offset order between each thread in this warp and the thread fragment request amounts of each thread in this warp, at least one memory address offset of each thread in this warp is alternately determined.
[0060] For each thread in this warp, the memory address offset of this thread is registered in the address converter, and the starting address of the memory space of this thread, as well as each memory address offset of this thread, are allocated to this thread, so as to be able to accurately address and ensure that the thread can efficiently access the allocated memory space.
[0061] Optionally, during the period when the GPU side is waiting for the host side to feedback the starting address of the memory space, the warp scheduler can register the data structures required for memory space allocation for each warp, determine that after receiving the starting address of the memory space, quickly and accurately associate the starting address of the memory space with the threads, and effectively overlap the overhead of memory registration, thereby reducing the waiting time and improving the efficiency of memory management.
[0062] According to the domain prefetching rule, when a part of a memory space is accessed, there is a high probability that the nearby memory spaces will be accessed. Through the technical solution of the embodiment of the present invention, the memory request amounts of the threads in the warp are split according to the minimum scheduling memory amount, so that the thread fragment request amount is aligned with the minimum scheduling memory amount of the GPU; according to the offset arrangement order, at least one memory address offset of each thread in the warp is alternately determined, and the memory spaces of different threads with highly similar memory access patterns are alternately arranged, which can enable the GPU side to completely prefetch the memory spaces of other threads with similar memory access patterns in the same warp through the minimum scheduling memory amount, improving the memory prefetch accuracy of the GPU.
[0063] Embodiment 2
[0064] Figure 2A It is a flowchart of a method for allocating memory space for threads provided by Embodiment 2 of the present invention. On the basis of the technical solution of the above embodiment, the operation of determining the memory address offset is optimized and improved.
[0065] Further, "alternately determine at least one first address offset of each first thread in the warp according to the offset arrangement order between the threads and the first thread fragment amount of each first thread in the warp; wherein, the first thread is a thread with a first thread fragment amount; according to the offset arrangement order between the threads, the second thread fragment amount of each second thread in the warp and the splitting quantity of the first thread fragment amount, alternately determine the second address offset of each second thread in the warp" to improve the operation of determining the memory offset address amount.
[0066] It should be noted that for the parts not detailed in the embodiments of the present invention, reference can be made to the descriptions of the foregoing embodiments.
[0067] See Figure 2A The method for allocating memory space for threads shown in the following, includes:
[0068] S201. Send the memory request amounts of at least one warp to the host side, so that the host side allocates the starting address of the memory space for each warp according to the memory request amounts of the warps, and feedbacks the starting address of the memory space of each warp to the GPU side.
[0069] S202. For each thread bundle, split the memory request amount of each thread in the thread bundle according to the minimum scheduling memory amount, to obtain at least one thread fragment request amount for each thread.
[0070] In this embodiment, the thread fragment request amount includes a first thread fragment amount and / or a second thread fragment amount; the first thread fragment amount is equal to the minimum scheduling memory amount; the second thread fragment amount is less than the minimum scheduling memory amount. Taking the minimum scheduling memory amount being 64KB as an example, if the memory request amount of a thread is 90KB, then a first thread fragment amount of 64KB and a second thread fragment amount of 26KB are split; if the memory request amount of a thread is 32KB, then a second thread fragment amount of 32KB is obtained.
[0071] S203. According to the offset arrangement order between each thread and the first thread fragment amounts of each first thread in the thread bundle, alternately determine at least one first address offset for each first thread in the thread bundle; where a first thread is a thread with a first thread fragment amount.
[0072] In this embodiment, the first address offset can be the memory address offset corresponding to the first thread fragment amount. According to the offset arrangement order, split the number of the first thread fragment amounts of each first thread in the thread bundle, and alternately sort the first thread fragment amounts of each first thread; according to the first thread fragment amount, that is, according to the minimum scheduling memory amount, determine the first address offset corresponding to each first thread fragment amount.
[0073] Exemplarily, if the offset arrangement order is thread B, thread A, and thread C; there are 3 first threads, namely thread A, thread B, and thread C; thread A includes 3 first thread fragment amounts; thread B includes 2 first thread fragment amounts; thread C includes 2 first thread fragment amounts, then the final determined sorting of the first thread fragment amounts is B1 - A1 - C1 - B2 - A2 - C2 - A3; it should be noted that in this exemplary embodiment, the subscripts 1, 2, and 3 of the first thread fragment amounts are only used to distinguish different first thread fragment amounts.
[0074] Exemplarily, after sorting, the first address offset of the first thread fragment amount with the order of 1 is 0; the first address offset of the first thread fragment amount with the order of 2 is the minimum scheduling memory amount; the first address offset of the first thread fragment amount with the order of 3 is twice the minimum scheduling memory amount; the first address offset of the first thread fragment amount with the order of 4 is three times the minimum scheduling memory amount ···.
[0075] S204. Determine the second address offsets of the second threads in the warp alternately according to the sorting order of the offsets between threads, the amounts of second thread fragments of the second threads in the warp, and the splitting quantity of the amount of first thread fragments; wherein, the second threads are the threads with the amounts of second thread fragments.
[0076] In this embodiment, the second address offset may be the memory address offset corresponding to the amount of second thread fragments. Sort the amounts of second thread fragments of the second threads alternately according to the sorting order of the offsets; determine the second address offset corresponding to the amount of second thread fragments with the order of 1 after sorting according to the splitting quantity and the amount of first thread fragments; after sorting, the second address offsets corresponding to the amounts of second thread fragments with the order above 2 are determined according to at least one amount of second thread fragments before it in the order and the second address offset corresponding to the amount of second thread fragments with the order of 1.
[0077] Exemplarily, when the sorting order of the offsets is thread B, thread A, and thread C; there are 3 first threads, namely thread A, thread B, and thread C; thread A includes 3 amounts of first thread fragments; thread B includes 2 amounts of first thread fragments; thread C includes 2 amounts of first thread fragments; and the sorting of the amounts of first thread fragments is B1 - A1 - C1 - B2 - A2 - C2 - A3; if there are 2 second threads, namely thread B and thread C, then the final sorting of the amounts of second thread fragments determined is B - C; after sorting, the second address offset corresponding to the amount of second thread fragments with the order of 1, that is, the amount of second thread fragments of thread B, is seven times the minimum scheduling memory amount; the second address offset of the amount of second thread fragments with the order of 2 is the sum of the amount of second thread fragments of thread B and the second address offset corresponding to the amount of second thread fragments of thread B.
[0078] An example of the overall process of determining the address offset is as follows. There are threads A, B, and C in a warp. The memory request volume of thread A is 192 KB; the memory request volume of thread B is 130 KB; the memory request volume of thread C is 135 KB. Then, the memory request volume of thread A can be split into 3 first thread fragment volumes of 64 KB each; the memory request volume of thread B can be split into 2 first thread fragment volumes of 64 KB each and 1 second thread fragment volume of 2 KB; the memory request volume of thread C can be split into 2 first thread fragment volumes of 64 KB each and 1 second thread fragment volume of 7 KB. The sorting of the first thread fragment volumes is B1 - A1 - C1 - B2 - A2 - C2 - A3. Then, the first address offsets of thread A are 64 KB, 256 KB, and 384 KB respectively; the first address offsets of thread B are 0 and 192 KB respectively; the first address offsets of thread C are 128 KB and 320 KB respectively. The sorting of the second thread fragment volumes is B - C. The second address offset of thread B is 448 KB; the second address offset of thread C is 450 KB. It should be noted that in this example, the second address offset is only an example represented in decimal, and it can also be represented in hexadecimal or binary. The present invention does not limit this.
[0079] S205. For each thread in the warp, allocate the starting address of the memory space of the warp and the memory address offsets of the thread to the thread.
[0080] In a specific embodiment, the starting address of the memory space of the warp and the memory address offsets of the thread can be allocated to the thread, so that the thread accesses the corresponding memory space according to the starting address of the memory space and the memory address offsets.
[0081] In another specific embodiment, the memory space can be determined according to the starting address of the memory space of the warp and the memory address offsets of the thread, and the determined memory space is allocated to the thread.
[0082] Exemplarily, Figure 2B is a schematic diagram of the alternating arrangement of the memory spaces of threads. As Figure 2B shown, the warp includes threads T0, T1, ···, and T n-1 . The memory request volume of thread T0 is split into the first thread fragment volume TB 00 and the first thread fragment volume TB 01 , as well as the thread fragment request volume not marked in the figure; the memory request volume of thread T1 is split into the first thread fragment volume TB 10 and the first thread fragment volume TB 11 , as well as the thread fragment request volume not marked in the figure; thread T n-1The memory request volume is split into the first thread fragment volume TB n-1,0 and the first thread fragment volume TB n-1,1 , and the thread fragment request volume not labeled in the figure. The UVM Space is the unified virtual memory space. In this exemplary embodiment, the offset arrangement order is determined according to the identification number of the thread, that is, the offset arrangement order is T0, T1,..., T n-1 . According to the offset arrangement order, the sorting of the first thread fragment volume is determined as TB 00 , TB 10 , TB 20 , ···, T n-1,0 , TB 01 , TB 11 , ···, TB n-1,1 , ···, then the first address offset corresponding to each first thread fragment volume can be determined. According to the first address offset and the starting address of the memory space, the memory space of each thread is determined. The memory space of thread T0 in the unified virtual memory space includes but is not limited to VB0 and VB n ; the memory space of thread T1 in the unified virtual memory space includes but is not limited to VB1; the memory space occupied by thread T n-1 in the unified virtual memory space but is not limited to VB n-1 .
[0083] Optionally, before alternately determining the second address offset of each second thread in the thread bundle according to the offset arrangement order between the threads, the second thread fragment volume of each second thread in the thread bundle, and the maximum first offset, it further includes:
[0084] For each second thread, determine the memory multiple between the second thread fragment volume of the second thread and the memory volume of the memory page; check whether the memory multiple is an integer; if the memory multiple is not an integer, then determine the target memory volume that is an integer between the memory volume of the memory page, has the smallest difference from the second thread fragment volume of the second thread, and is greater than the second thread fragment volume of the second thread; update the second thread fragment volume of the second thread to the target memory volume.
[0085] Wherein, the memory page is a basic unit of memory management. The memory volume of the memory page usually has 4KB, 8KB, etc. Taking the memory volume of the memory page as 4KB and the second thread fragment volume as 7KB as an example, the memory multiple between the second thread fragment volume of 9KB and the memory volume of the memory page of 4KB is not an integer, then 12KB, which is an integer between the memory volume of the memory page, has the smallest difference from the second thread fragment volume of the second thread, and is greater than the second thread fragment volume of the second thread, is used as the target memory volume, and the second thread fragment volume is updated to 12KB.
[0086] It can be understood that by adopting the above technical solution, the amount of the second thread fragments can be aligned with the memory amount of the memory page, so that when allocating a memory space for the amount of the second thread fragments, an integer number of memory page spaces can be allocated; furthermore, the situation of allocating the same memory page for different amounts of the second thread fragments can be avoided, ensuring that when prefetching the memory page corresponding to the amount of the second thread fragments, the memory of other threads will not be prefetched, improving the accuracy of prefetching, and at the same time avoiding the generation of memory space fragmentation.
[0087] Optionally, the integer includes non - negative integer powers of 2, that is, the integer can only be non - negative integer powers of 2, and can be expressed as 2 i , where i is a non - negative integer. Taking the memory amount of a memory page as 4KB and the amount of the second thread fragments as 9KB as an example, the determined target memory amount is 16KB, and the amount of the second thread fragments is updated to 12KB. In the specific implementation, the minimum scheduling memory amount is 64KB. By adopting the above technical solution, updating the amount of the second thread fragments to a non - negative integer power multiple of the memory amount of the memory page can facilitate the integration of multiple updated amounts of the second memory fragments into a minimum - scheduled memory block, thereby reducing the number of memory space fragments, improving the utilization rate of the memory space and the memory access efficiency.
[0088] The technical solution of the embodiment of the present invention can make the second address offset allocated according to the amount of the second thread fragments all greater than the first address offset allocated according to the amount of the first thread fragments, so that the memory spaces of the minimum scheduling memory amounts of each thread are alternately arranged in the front, and the memory spaces smaller than the minimum scheduling memory amount are arranged in the back. When the GPU prefetches the memory space of a certain minimum scheduling memory amount, it can completely prefetch the adjacent memory spaces of the minimum scheduling memory amount, avoiding the situation that due to the existence of the memory spaces smaller than the minimum scheduling memory amount, the memory spaces of the minimum scheduling memory amounts of other threads cannot be completely obtained, and improving the accuracy of prefetching.
[0089] Embodiment III
[0090] Figure 3 As shown in the flowchart of a method for allocating memory space for a thread provided by Embodiment III of the present invention, this embodiment is applicable to the situation of allocating memory space for a thread. This method can be executed by a device for allocating memory space for a thread, and the device for allocating memory space for a thread can be implemented in the form of hardware and / or software and is specifically configured in an electronic device.
[0091] It should be noted that for the parts not detailed in the embodiments of the present invention, reference can be made to the descriptions of the foregoing embodiments.
[0092] See Figure 3 The method for allocating memory space for a thread shown, which is executed by the host side, includes:
[0093] S301. Receive the memory request amounts of at least one warp sent by the GPU side.
[0094] S302. For each warp, split the memory request amount of this warp according to the pre-partitioned space memory amount to obtain at least one warp fragment request amount.
[0095] S303. According to the warp fragment request amounts of each warp, allocate free pre-partitioned memory spaces for each warp fragment request amount of each warp.
[0096] S304. For each warp, determine the starting address of the memory space of this warp according to the starting addresses of the pre-partitioned memory spaces of this warp and the warp fragment request amounts in this warp.
[0097] S305. Send the starting address of the memory space of this warp to the GPU side, so that the GPU side can, according to the starting address of the memory space of this warp, allocate the starting address of the memory space and the memory address offset for at least one thread in this warp.
[0098] In this embodiment, a warp may include at least one thread. The memory request amount of a warp is the amount of memory that the warp requests to be allocated for it. The host side may be the side including the CPU (Central Processing Unit) and the memory managed by it. The memory space may be a Unified Virtual Memory Space (UVM Space). The pre-partitioned space memory amount may be the memory amount of the pre-partitioned memory spaces pre-partitioned in the unified virtual memory space. The pre-partitioned space memory amount may be set independently by those skilled in the art according to actual needs or practical experience, and the present invention does not limit this. In a specific implementation manner, the unified virtual memory space may be divided into 2048 pre-partitioned memory spaces, and the memory amount of each pre-partitioned memory space is 2MB. The warp fragment request amount is the result of splitting the memory request amount of the warp.
[0099] Specifically, receive the warp memory allocation requests of at least one warp sent by the GPU side; wherein, the warp memory allocation requests may be used to instruct the host side to allocate memory for the warp, and the allocated memory amount is the memory request amount. For each warp, split the memory request amount of this warp according to the pre-partitioned space memory amount, and use the splitting result as the warp fragment request amount; further, if there is a splitting result less than the pre-partitioned space memory amount, then this splitting result may be amplified, or this splitting result may be directly used as the warp fragment request amount.
[0100] Using a certain algorithm, allocate idle pre-partitioned memory space for each warp fragment request amount of each warp according to the warp fragment request amounts of each warp; using a certain algorithm, for each warp, determine the starting address of the memory space of the warp according to the starting addresses of the pre-partitioned memory spaces of the warp and the warp fragment request amounts in the warp.
[0101] Send the starting address of the memory space of the warp to the GPU side; for each warp, on the GPU side, split the memory request amount of each thread in the warp according to the minimum scheduling memory amount to obtain at least one thread fragment request amount for each thread; alternately determine at least one memory address offset for each thread in the warp according to the offset order between the threads in the warp and the thread fragment request amounts of the threads in the warp; for each thread in the warp, allocate the starting address of the memory space of the warp and the memory address offsets of the thread to the thread.
[0102] The technical solution of the embodiment of the present invention splits the memory request amount of the warp through the pre-partitioned space memory amount, and allocates idle pre-partitioned memory space for the split warp fragment request amounts, realizing the alignment of the memory space of the warp and the pre-partitioned memory space, ensuring that at least one complete pre-partitioned memory space is allocated to the warp, so that when the GPU prefetches memory, the memory boundary of the warp can be accurately defined, avoiding the interference of other warps, and improving the prefetch hit rate.
[0103] Embodiment 4
[0104] Figure 4A It is a flowchart of a method for allocating the memory space of a thread provided in Embodiment 4 of the present invention. On the basis of the technical solution of the above embodiment, the allocation operation of the pre-partitioned memory space is optimized and improved.
[0105] Further, "allocate idle pre-partitioned memory space for each warp fragment request amount of each warp according to the warp fragment request amounts of each warp" is refined to "allocate idle pre-partitioned memory space in the unified virtual memory space for each first warp fragment amount; the pre-partitioned memory spaces corresponding to each first warp fragment amount are different; according to each second warp fragment amount, determine at least one idle pre-partitioned memory space in the unified virtual memory space, and use the determined pre-partitioned memory space as the integrated memory space; determine the allocation association relationship between each integrated memory space and at least one second warp fragment amount" to improve the allocation operation of the pre-partitioned memory space.
[0106] It should be noted that for the parts not described in detail in the embodiments of the present invention, reference can be made to the descriptions of the foregoing embodiments.
[0107] SeeFigure 4A The memory space allocation method for the shown threads includes:
[0108] S401. Receive the memory request amounts of at least one warp sent from the GPU side.
[0109] S402. For each warp, split the memory request amount of the warp according to the pre-partitioned space memory amount to obtain at least one warp fragment request amount.
[0110] In this embodiment, the warp fragment request amount includes a first warp fragment amount and a second warp fragment amount; the first warp fragment amount is equal to the pre-partitioned space memory amount; the second warp fragment amount is less than the pre-partitioned space memory amount.
[0111] Exemplarily, if there are 4 warps, namely warp A, warp B, warp C, and warp D; the memory request amount of warp A is 6096KB; the memory request amount of warp B is 4096KB; the memory request amount of warp C is 256KB; the memory request amount of warp D is 64KB; the pre-partitioned space memory amount is 2MB, then warp A is split into 2 first warp fragment amounts of 2MB and 1 second warp fragment amount of 2000KB; warp B is split into 2 first warp fragment amounts of 2MB; warp C has 1 second warp fragment amount of 256KB after splitting; warp D is split into 1 second warp fragment amount of 64KB.
[0112] S403. Allocate the idle pre-partitioned memory space in the unified virtual memory space for each first warp fragment amount; the pre-partitioned memory spaces corresponding to each first warp fragment amount are different.
[0113] Specifically, for each first warp fragment amount, allocate the idle pre-partitioned memory space in the unified virtual memory space for the first warp fragment amount; among them, the pre-partitioned memory spaces allocated for each first warp fragment amount are different; preferably, allocate the idle and continuously addressed pre-partitioned memory spaces for the first warp fragment amounts belonging to the same warp.
[0114] S404. Determine at least one idle pre-partitioned memory space in the unified virtual memory space according to each second warp fragment amount, and use the determined pre-partitioned memory space as the integrated memory space.
[0115] In this embodiment, one integrated memory space can correspond to the second warp fragment amounts of at least one thread. Specifically, count the total memory amount among the second warp fragment amounts; determine the number of spaces of the pre-partitioned memory space to be allocated according to the total memory amount and the pre-partitioned space memory amount; select the idle pre-partitioned memory spaces with the number of spaces as the integrated memory space.
[0116] Exemplarily, if there are 4 warps, namely warp A, warp B, warp C, and warp D; warp A has a second warp fragment amount of 2000 KB; warp B has no second warp fragment amount; warp C has a second warp fragment amount of 256 KB; warp D has a second warp fragment amount of 64 KB, then the total memory amount of each second warp fragment amount is 2320 KB; the memory amount of a pre-partitioned memory space is 2048 KB, then the number of pre-partitioned memory spaces to be allocated is 2; select 2 idle pre-partitioned memory spaces as integrated memory spaces.
[0117] Optionally, before determining at least one idle pre-partitioned memory space in the unified virtual memory space according to each second warp fragment amount, it includes:
[0118] For each second warp fragment amount, determine the memory multiple between the second warp fragment amount and the minimum scheduling memory amount; check whether the memory multiple is a set multiple; the set multiple includes positive integer powers of 2; if the memory multiple is not a set multiple, then determine the memory multiple between the second warp fragment amount and the minimum scheduling memory amount as the set multiple, which has the smallest difference from the second warp fragment amount of this second thread and is greater than the target memory amount of this second warp fragment amount; update this second warp fragment amount to the target memory amount.
[0119] Among them, the minimum scheduling memory amount can be the minimum memory amount scheduled when the GPU performs a prefetch operation. Exemplarily, if the minimum scheduling memory amount is 64 KB and there is a second warp fragment amount of 63 KB, then the memory multiple between this second warp fragment amount and the minimum scheduling memory amount is not a set multiple; determine the memory multiple between this second warp fragment amount and the minimum scheduling memory amount as the set multiple, which has the smallest difference from the second warp fragment amount of this second thread and is greater than the target memory amount of this second warp fragment amount of 64 KB; update this second warp fragment amount to 64 KB.
[0120] It can be understood that when allocating memory space for the second warp fragment amount, it is possible to allocate memory spaces in integer multiples of the minimum scheduling memory amount; thereby avoiding different second warp fragment amounts from occupying the same memory space of the minimum scheduling memory amount, ensuring that the GPU will not be interfered by other threads when processing with the minimum scheduling memory amount, and also avoiding the generation of memory space fragmentation.
[0121] S405. Determine the allocation association relationship between each integrated memory space and at least one second warp fragment amount.
[0122] Specifically, according to the sizes of each second warp fragment amount, determine the allocation association relationship between each integrated memory space and at least one second warp fragment amount.
[0123] Exemplarily, if there are 4 warps, namely warp A, warp B, warp C, and warp D; warp A has a second warp fragment amount of 2000 KB; warp B has no second warp fragment amount; warp C has a second warp fragment amount of 256 KB; warp D has a second warp fragment amount of 64 KB; the integrated memory space includes integrated memory space A and integrated memory space B; the total memory amount of the second warp fragment amount of warp A and the second warp fragment amount of any other warp exceeds the memory amount of the integrated memory space; associate integrated memory space A with the second warp fragment amount of warp A; associate integrated memory space B with the second warp fragment amount of warp C and the second warp fragment amount of warp D.
[0124] S406. For each warp, determine the starting address of the memory space of the warp according to the starting addresses of the pre-partitioned memory spaces of the warp and the warp fragment request amounts in the warp.
[0125] Optionally, determining the starting address of the memory space of the warp according to the starting addresses of the pre-partitioned memory spaces of the warp and the warp fragment request amounts in the warp includes:
[0126] For each first warp fragment amount in the warp, determine the starting address of the pre-partitioned memory space corresponding to the first warp fragment amount as the starting address of the memory space of the warp; for each second warp fragment amount in the warp, according to the second warp fragment amount and the idle state of the full binary tree node corresponding to the integrated memory space of the second warp fragment amount, allocate at least one minimum prefetch subspace for the second warp fragment amount; the leaf nodes of the full binary tree correspond to the minimum prefetch subspaces in the integrated memory space; the minimum regional subspaces corresponding to the leaf nodes are different; according to the starting address of the allocated minimum prefetch subspace, determine the starting address of the memory space corresponding to the second warp fragment amount of the warp; determine the starting address of the memory space corresponding to the second warp fragment amount of the warp as the starting address of the memory space of the warp.
[0127] In this alternative embodiment, an integrated memory space can be further divided into minimum prefetch sub-spaces; the memory amount of the minimum prefetch sub-space is the minimum scheduling memory amount. The minimum prefetch sub-spaces of the integrated memory space are managed by a full binary tree. The leaf nodes of the full binary tree correspond to the minimum prefetch sub-spaces; among them, the addresses of the minimum prefetch sub-spaces corresponding to adjacent leaf nodes are consecutive. Exemplarily, if the integrated memory space is 2MB, that is, 2048KB; and the memory amount of the minimum prefetch sub-space is 64KB, then the integrated memory space includes 32 minimum prefetch sub-spaces, and the corresponding full binary tree has 32 leaf nodes. If the minimum prefetch sub-space corresponding to a leaf node in the full binary tree has been occupied, the idle state of this leaf node is occupied, otherwise the idle state of this leaf node is idle. For any node in the full binary tree other than the leaf nodes, if the child nodes of this node are all in the idle state, then this node is in the idle state, otherwise, this node is in the occupied state.
[0128] Specifically, according to the amount of the second thread bundle fragmentation, determine the height for finding the parent node; exemplarily, the height for finding the parent node can be expressed by the following formula:
[0129]
[0130] Among them, h represents the height for finding the parent node; M represents the second fragmentation amount; d represents the minimum scheduling memory amount; [] represents rounding up.
[0131] In the way that the height of the leaf nodes in the full binary tree is 0 and the height increases sequentially upwards, determine the first nodes at the positions of the height for finding the parent node; select an idle first node from the first nodes as the target node; determine the leaf nodes of the full binary subtree with this target node as the parent node as the target subtree; allocate the minimum prefetch sub-spaces corresponding to the leaf nodes of this target subtree to the amount of the second thread bundle fragmentation.
[0132] Among the starting addresses of the allocated minimum prefetch sub-spaces, determine the smallest starting address as the starting address of the memory space of the second thread bundle fragmentation amount of this thread bundle; determine the starting address of the memory space corresponding to the second thread bundle fragmentation amount of this thread bundle as the starting address of the memory space of this thread bundle.
[0133] In an alternative embodiment, the full binary tree includes 32 leaf nodes; this full binary tree has 6 layers of nodes; the height of the leaf nodes is 0, and the height of the root node of this full binary tree is 5; the memory amount of the integrated memory space corresponding to this full binary tree is 2048KB, and each leaf node corresponds to a 64KB minimum prefetch sub-space. Figure 4B It is a schematic diagram of the correspondence between the amount of the second thread bundle fragmentation in the full binary tree and the full binary subtree. As Figure 4BAs shown, there are red subtrees, blue subtrees, and green subtrees in the full binary tree. The minimum prefetch subspace corresponding to the 4 leaf nodes included in the red subtree is the minimum prefetch subspace allocated for the second warp fragment amount corresponding to the red subtree; the minimum prefetch subspace corresponding to the 2 leaf nodes included in the blue subtree is the minimum prefetch subspace allocated for the second warp fragment amount corresponding to the blue subtree; the minimum prefetch subspace corresponding to the 8 leaf nodes included in the green subtree is the minimum prefetch subspace allocated for the second warp fragment amount corresponding to the green subtree.
[0134] Take the second warp fragment amount corresponding to the red subtree as the second warp fragment amount a; take the second warp fragment amount corresponding to the blue subtree as the second warp fragment amount b; take the second warp fragment amount corresponding to the green subtree as the second warp fragment amount c. When the integrated memory space is free, first allocate a full binary subtree for the second warp fragment amount a. The second warp fragment amount a is 256KB. The determined parent node searches for a height of 3. Then, in the level with a height of 3, search for a free node and allocate the next full binary subtree under this node to the second warp fragment amount a, that is, allocate the minimum prefetch subspace corresponding to the leaf nodes of the red subtree to the second warp fragment amount a; update the free state of the leaf nodes in the red subtree from free to occupied, and starting from the leaf nodes, recursively update the free states of each node until the root node of the full binary tree is updated.
[0135] Allocate a full binary subtree for the second warp fragment amount b. The second warp fragment amount b is 128KB. The determined parent node searches for a height of 2. Then, in the level with a height of 2, search for a free node and allocate the next full binary subtree under this node to the second warp fragment amount a, that is, allocate the minimum prefetch subspace corresponding to the leaf nodes of the blue subtree to the second warp fragment amount a; update the free state of the leaf nodes in the blue subtree from free to occupied, and starting from the leaf nodes, recursively update the free states of each node until the root node of the full binary tree is updated.
[0136] Allocate a full binary subtree for the second warp fragment amount c. The second warp fragment amount c is 512KB. The determined parent node searches for a height of 4. Then, in the level with a height of 4, search for a free node and allocate the next full binary subtree under this node to the second warp fragment amount c, that is, allocate the minimum prefetch subspace corresponding to the leaf nodes of the green subtree to the second warp fragment amount c; update the free state of the leaf nodes in the green subtree from free to occupied, and starting from the leaf nodes, recursively update the free states of each node until the root node of the full binary tree is updated.
[0137] It can be understood that by adopting the above technical solution, the search height of the parent node is determined based on the amount of the second warp fragment, and a node is determined among the nodes with the height being the search height of the parent node, and the full binary tree of the determined node is allocated to the amount of the second warp fragment, which can avoid that the full binary tree with the root node being the parent node of the root node of an already allocated full binary tree covers the full binary tree corresponding to another warp fragment amount, and avoid memory overlap and interference across warps; meanwhile, the smallest prefetch subspace not occupied in the memory space is integrated and retained in the unified virtual memory space and will not be mapped to the physical memory space during the prefetch process, thereby effectively preventing the prefetch of these filled areas and improving the accuracy and efficiency of memory prefetch.
[0138] In a specific embodiment, before receiving the memory request amount of at least one warp sent by the GPU side, the unified virtual memory space can be pre-divided into 2048 pre-divided memory spaces in advance, and the full binary tree structure is initialized to manage the smallest prefetch subspace in each pre-divided memory space.
[0139] S407. Send the starting address of the memory space of the warp to the GPU side so that the GPU side can determine the memory space corresponding to at least one thread in the warp according to the starting address of the warp in the unified virtual memory space.
[0140] Optionally, Figure 4C is a schematic diagram of the processing flow of a thread memory allocation framework. As Figure 4C shown, in the stage of the GPU dynamically applying for memory, the thread allocation requests inside the warp are "reduced" and merged into a unified request through the underlying warp shared register method; the memory requests are processed by the host side as a whole in units of warps instead of processing each thread one by one, thereby reducing the overhead in memory management; after the requests are merged, the framework sends the memory requests of the warp to the host side through the page replacement method of the unified virtual memory space via the scheduler; wait for the host side to respond with the starting address of the domestic allocated memory space. At the same time, during the waiting for the host response, the scheduler registers the data structures required for memory allocation for each thread to ensure that the starting address of the memory space can be accurately associated with the thread after the starting address of the allocated memory space arrives; this process effectively overlaps the overhead of memory registration, thereby reducing the waiting time and improving the efficiency of memory management.
[0141] The host side pre-divides the continuous unified virtual memory space into pre-divided memory spaces of the same size in advance to align with the GPU's own neighborhood prefetch policy. After receiving the memory request of the GPU warp, locate the idle pre-divided memory spaces in the pre-divided memory spaces that have been pre-divided in advance, and then return the starting addresses of these pre-divided memory spaces to the GPU side. This framework no longer reallocates a new address segment for each warp separately, improving the efficiency of memory allocation.
[0142] When the GPU side receives the starting address of the allocated memory space from the host side, the scheduler alternately allocates memory space for each thread to align with the SIMT (single-instruction multiple-thread) operation scheduling mode; the warp scheduler will also register the memory address offset of each thread to the address converter to ensure that subsequent memory accesses can perform accurate addressing according to the new memory layout. Finally, the scheduler distributes the allocated memory addresses to each thread within the warp to ensure that the threads can efficiently access under the new memory layout.
[0143] The solution of this alternative embodiment realizes the precise matching of memory prefetching and GPU access mode by aligning the prefetching strategy of the unified virtual memory space and the operation scheduling strategy of the SIMT architecture of the GPU, enabling the migrated memory space to be efficiently accessed by GPU threads in a short time, thereby reducing the eviction of the migrated memory space due to physical memory oversubscription before it is accessed, improving the utilization rate of physical memory, and being able to significantly improve the performance of the prefetching strategy without changing the underlying architecture of the GPU, and having strong adaptability and generality, so that this technology can be effectively deployed in actual applications rather than in a simulation environment, improving the performance of the overall program. Fundamentally solves the problems of low memory utilization rate and poor transmission performance caused by the overly aggressive GPU prefetching strategy. Different from the traditional method of simply dynamically adjusting the prefetch intensity. By accurately prefetching the memory space that the subsequent GPU needs to access, it reduces the performance degradation caused by frequent GPU page faults and improves the collaborative memory sharing performance between the host side and the GPU side.
[0144] The embodiment of the present invention determines the integrated memory space and classifies the minimum prefetch subspaces for at least one second warp fragment amount in one integrated memory space, avoiding each second warp fragment amount occupying an integrated memory space alone, reducing the generation of memory fragments, and improving the memory utilization rate.
[0145] Embodiment 5
[0146] Figure 5 It is a schematic structural diagram of a device for allocating the memory space of a thread provided in Embodiment 5 of the present invention. The embodiment of the present invention is applicable to the situation of allocating the memory space of a thread. This device can execute the method for allocating the memory space of a thread. This device for allocating the memory space of a thread can be implemented in the form of hardware and / or software, and this device can be configured in an electronic device.
[0147] See Figure 5The memory space allocation device for the threads shown includes a request sending module 501, a memory splitting module 502, an offset determination module 503, and a memory space determination module 504. Among them,
[0148] The request sending module 501 is configured to send the memory request amounts of at least one warp of threads to the host side, so that the host side allocates the starting addresses of the memory spaces for each warp of threads according to the memory request amounts of each warp of threads, and feeds back the starting addresses of the memory spaces of each warp of threads to the GPU side;
[0149] The memory splitting module 502 is configured to, for each warp of threads, split the memory request amount of each thread in the warp of threads according to the minimum scheduling memory amount, to obtain at least one thread fragment request amount of each thread;
[0150] The offset determination module 503 is configured to alternately determine at least one memory address offset of each thread in the warp of threads according to the offset arrangement order among the threads in the warp of threads and the thread fragment request amounts of the threads in the warp of threads;
[0151] The memory space determination module 504 is configured to, for each thread in the warp of threads, allocate the starting address of the memory space of the warp of threads and each memory address offset of the thread to the thread.
[0152] According to the domain prefetching rule, when a part of a memory space is accessed, there is a high probability that the nearby memory spaces will be accessed. Through the technical solution of the embodiment of the present invention, the memory request amounts of the threads in the warp of threads are split according to the minimum scheduling memory amount, so that the thread fragment request amounts are aligned with the minimum scheduling memory amount of the GPU; according to the offset arrangement order, at least one memory address offset of each thread in the warp of threads is alternately determined, and the memory spaces of different threads with highly similar memory access patterns are alternately arranged, which can enable the GPU side to completely prefetch the memory spaces of other threads with similar memory access patterns in the same warp of threads through the minimum scheduling memory amount, and improve the memory prefetch accuracy of the GPU.
[0153] Optionally, in this device, the thread fragment request amount includes a first thread fragment amount and / or a second thread fragment amount; the first thread fragment amount is equal to the minimum scheduling memory amount; the second thread fragment amount is less than the minimum scheduling memory amount;
[0154] The offset determination module 503 includes:
[0155] The first determination unit is configured to alternately determine at least one first address offset of each first thread in the warp of threads according to the offset arrangement order among the threads and the first thread fragment amounts of the first threads in the warp of threads; where the first thread is a thread with a first thread fragment amount;
[0156] A second determination unit, configured to alternately determine second address offsets of each second thread in the warp according to the arrangement order of offsets between threads, the amount of second thread fragments of each second thread in the warp, and the splitting quantity of the amount of first thread fragments; wherein, the second thread is a thread with an amount of second thread fragments.
[0157] Optionally, the apparatus further includes:
[0158] A multiple determination module, configured to determine, for each second thread, a memory multiple between the amount of second thread fragments of the second thread and the memory amount of a memory page;
[0159] A verification module, configured to verify whether the memory multiple is an integer;
[0160] A target memory amount determination module, configured to, if the memory multiple is not an integer, determine a target memory amount that has an integer memory multiple with the memory amount of the memory page, has the smallest difference from the amount of second thread fragments of the second thread, and is greater than the amount of second thread fragments of the second thread;
[0161] An update module, configured to update the amount of second thread fragments of the second thread to the target memory amount.
[0162] Optionally, in the apparatus, the integer includes non - negative integer powers of 2.
[0163] The thread memory space allocation apparatus provided by an embodiment of the present invention can execute the thread memory space allocation method provided by any embodiment of the present invention, and has functional modules and beneficial effects corresponding to executing the thread memory space allocation method.
[0164] Embodiment Six
[0165] Figure 6 It is a schematic structural diagram of a thread memory space allocation apparatus provided by Embodiment Six of the present invention. Embodiments of the present invention are applicable to the situation of allocating memory space for threads. The apparatus can execute the thread memory space allocation method. The thread memory space allocation apparatus can be implemented in the form of hardware and / or software, and the apparatus can be configured in an electronic device.
[0166] Refer to Figure 6 the thread memory space allocation apparatus shown, including a receiving module 601, a splitting module 602, an allocation module 603, a determination module 604, and a sending module 605, wherein,
[0167] The receiving module 601 is configured to receive the memory request amount of at least one warp sent by the GPU side;
[0168] The splitting module 602 is used to split the memory request amount of each warp according to the pre-partitioned space memory amount for each warp, so as to obtain at least one warp fragment request amount;
[0169] The allocation module 603 is used to allocate free pre-partitioned memory spaces for each warp fragment request amount of each warp according to the number and the warp fragment request amount of each warp;
[0170] The determination module 604 is used to determine the starting address of the memory space of each warp according to the starting addresses of the pre-partitioned memory spaces of the warp and each warp fragment request amount in the warp;
[0171] The sending module 605 is used to send the starting address of the memory space of the warp to the GPU side, so that the GPU side can allocate the starting address of the memory space and the memory address offset for at least one thread in the warp according to the starting address of the memory space of the warp.
[0172] The technical solution of the embodiment of the present invention can make the second address offsets allocated according to the second thread fragment amounts all greater than the first address offsets allocated according to the first thread fragment amounts, so that the memory spaces of the minimum scheduling memory amounts of each thread are alternately arranged in the front, and the memory spaces smaller than the minimum scheduling memory amounts are arranged in the back, enabling the GPU to completely prefetch the adjacent memory spaces of the minimum scheduling memory amounts when prefetched a certain memory space of the minimum scheduling memory amount, avoiding the inability to completely obtain the memory spaces of the minimum scheduling memory amounts of other threads due to the existence of the memory spaces smaller than the minimum scheduling memory amounts, and improving the accuracy of prefetching.
[0173] Optionally, the warp fragment request amount includes a first warp fragment amount and a second warp fragment amount; the first warp fragment amount is equal to the pre-partitioned space memory amount; the second warp fragment amount is smaller than the pre-partitioned space memory amount;
[0174] The allocation module 603 includes:
[0175] The first allocation unit is used to allocate free pre-partitioned memory spaces in the unified virtual memory space for each first warp fragment amount; the pre-partitioned memory spaces corresponding to each first warp fragment amount are different;
[0176] The second allocation unit is used to determine at least one free pre-partitioned memory space in the unified virtual memory space according to each second warp fragment amount, and use the determined pre-partitioned memory space as the integrated memory space;
[0177] The association relationship determination unit is used to determine the allocation association relationship between each integrated memory space and at least one second warp fragment amount.
[0178] Optionally, the apparatus further includes:
[0179] A multiple determination module, configured to determine, for each second thread bundle fragment amount, a memory multiple between the second thread bundle fragment amount and the minimum scheduling memory amount;
[0180] A verification module, configured to verify whether the memory multiple is a set multiple; the set multiple includes a positive integer power of 2;
[0181] A target memory amount determination module, configured to, if the memory multiple is not the set multiple, determine a target memory amount that has a memory multiple of the set multiple with the minimum scheduling memory amount, has the smallest difference from the second thread fragment amount of the second thread, and is greater than the second thread bundle fragment amount;
[0182] An update module, configured to update the second thread bundle fragment amount to the target memory amount.
[0183] Optionally, the determination module 604 includes:
[0184] A first determination unit, configured to, for each first thread bundle fragment amount in the thread bundle, determine the start address of the pre-partitioned memory space corresponding to the first thread bundle fragment amount as the start address of the memory space of the thread bundle;
[0185] An allocation unit, configured to, for each second thread bundle fragment amount in the thread bundle, allocate at least one minimum prefetch subspace for the second thread bundle fragment amount according to the second thread bundle fragment amount and the idle state of the full binary tree nodes of the integrated memory space corresponding to the second thread bundle fragment amount; the leaf nodes of the full binary tree correspond to the minimum prefetch subspaces in the integrated memory space; the minimum region subspaces corresponding to each leaf node are different;
[0186] A second determination unit, configured to determine the start address of the memory space corresponding to the second thread bundle fragment amount of the thread bundle according to the start address of the allocated minimum prefetch subspace;
[0187] A third determination unit, configured to determine the start address of the memory space corresponding to the second thread bundle fragment amount of the thread bundle as the start address of the memory space of the thread bundle.
[0188] The memory space allocation apparatus for a thread provided by an embodiment of the present invention can execute the memory space allocation method for a thread provided by any embodiment of the present invention, and has function modules and beneficial effects corresponding to executing the memory space allocation method for a thread.
[0189] Embodiment Seven
[0190] Figure 7FIG. 0 shows a schematic structural diagram of a thread memory space allocation device 700 that can be used to implement embodiments of the present invention. The thread memory space allocation device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The thread memory space allocation device can also represent various forms of mobile devices, such as, for example, personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0191] As Figure 7 shown, the thread memory space allocation device 700 includes at least one processor 701, and a memory communicatively connected to the at least one processor 701, such as read-only memory (ROM) 702, random access memory (RAM) 703, etc. The memory stores a computer program executable by the at least one processor. The processor 701 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 702 or the computer program loaded from the storage unit 708 into the random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the thread memory space allocation device 700 can also be stored. The processor 701, ROM 702, and RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0192] Multiple components in the thread memory space allocation device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the thread memory space allocation device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0193] The processor 701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 701 executes the various methods and processes described above, such as the thread memory space allocation method.
[0194] In some embodiments, the method for allocating memory space for threads can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto the memory space allocation device 700 for threads via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by the processor 701, one or more steps of the method for allocating memory space for threads described above can be performed. Alternatively, in other embodiments, the processor 701 can be configured to execute the method for allocating memory space for threads by any other suitable means (e.g., by means of firmware).
[0195] The various implementations of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on a chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0196] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable memory space allocation devices for threads, such that when the computer programs are executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer programs can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0197] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0198] To provide for interaction with a user, the systems and techniques described herein can be implemented on a thread's memory space allocation device that has: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the thread's memory space allocation device. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0199] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0200] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The client-server relationship is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS (Virtual Private Server) services.
[0201] It should be understood that various forms of processes shown above can be used, steps can be reordered, added or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.
[0202] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for allocating memory space for threads, characterized in that Executed by a GPU (Graphics Processing Unit), the method includes: Sending the memory request amount of at least one warp to the host side, so that the host side allocates the starting address of the memory space for each warp according to the memory request amount of each warp, and feeds back the starting address of the memory space of each warp to the GPU side; For each warp, splitting the memory request amount of each thread in the warp according to the minimum scheduling memory amount to obtain at least one thread fragment request amount of each thread; According to the offset arrangement order among the threads in the warp and the thread fragment request amounts of the threads in the warp, alternately determining at least one memory address offset of the threads in the warp; For each thread in the warp, allocating the starting address of the memory space of the warp and the memory address offsets of the thread to the thread.
2. The method according to claim 1, characterized in that, The thread fragment request amount includes a first thread fragment amount and / or a second thread fragment amount; the first thread fragment amount is equal to the minimum scheduling memory amount; the second thread fragment amount is less than the minimum scheduling memory amount; The alternately determining at least one memory address offset of the threads in the warp according to the offset arrangement order among the threads in the warp and the thread fragment request amounts of the threads in the warp includes: According to the offset arrangement order among the threads and the first thread fragment amounts of the first threads in the warp, alternately determining at least one first address offset of the first threads in the warp; wherein, the first thread is a thread with a first thread fragment amount; According to the offset arrangement order among the threads, the second thread fragment amounts of the second threads in the warp and the splitting quantity of the first thread fragment amount, alternately determining a second address offset of each second thread in the warp; wherein, the second thread is a thread with a second thread fragment amount.
3. The method according to any one of claims 2, characterized in that, Before alternately determining the second address offset of each second thread in the warp according to the offset arrangement order among the threads, the second thread fragment amounts of the second threads in the warp and the maximum first offset, it further includes: For each second thread, determining the memory multiple between the second thread fragment amount of the second thread and the memory amount of the memory page; Verifying whether the memory multiple is an integer; If the memory multiple is not an integer, determining the target memory amount that is an integer multiple of the memory amount of the memory page, has the smallest difference from the second thread fragment amount of the second thread, and is greater than the second thread fragment amount of the second thread; Updating the second thread fragment amount of the second thread to the target memory amount.
4. The method according to claim 3, wherein The integer includes non - negative integer powers of 2.
5. A method for allocating memory space for a thread, characterized in that, Executed by the host side, the method includes: Receiving the memory request amount of at least one warp sent by the GPU side; For each warp, splitting the memory request amount of the warp according to the pre - divided space memory amount to obtain at least one warp fragment request amount; Allocate free pre-partitioned memory spaces for each thread bundle fragment request amount of each of the thread bundles; For each thread bundle, determine the starting address of the memory space of the thread bundle according to the starting addresses of the pre-partitioned memory spaces of the thread bundle and each thread bundle fragment request amount in the thread bundle; Send the starting address of the memory space of the thread bundle to the GPU side, so that the GPU side allocates the starting address of the memory space and the memory address offset for at least one thread in the thread bundle according to the starting address of the memory space of the thread bundle.
6. The method according to claim 5, wherein The thread bundle fragment request amount includes a first thread bundle fragment amount and a second thread bundle fragment amount; the first thread bundle fragment amount is equal to the pre-partitioned space memory amount; the second thread bundle fragment amount is less than the pre-partitioned space memory amount; The step of allocating free pre-partitioned memory spaces for each thread bundle fragment request amount of each of the thread bundles includes: Allocate free pre-partitioned memory spaces in the unified virtual memory space for each first thread bundle fragment amount; the pre-partitioned memory spaces corresponding to each of the first thread bundle fragment amounts are different; Determine at least one free pre-partitioned memory space in the unified virtual memory space according to each of the second thread bundle fragment amounts, and use the determined pre-partitioned memory space as an integrated memory space; Determine the allocation association relationship between each integrated memory space and at least one second thread bundle fragment amount.
7. The method according to claim 6, wherein Before determining at least one free pre-partitioned memory space in the unified virtual memory space according to each of the second thread bundle fragment amounts, it includes: For each second thread bundle fragment amount, determine the memory multiple between the second thread bundle fragment amount and the minimum scheduling memory amount; Verify whether the memory multiple is a set multiple; the set multiple includes positive integer powers of 2; If the memory multiple is not the set multiple, then determine the memory multiple between the second thread bundle fragment amount and the minimum scheduling memory amount as the set multiple, which has the smallest difference from the second thread fragment amount of the second thread and is greater than the second thread bundle fragment amount, as the target memory amount; Update the second thread bundle fragment amount to the target memory amount.
8. The method according to any one of claims 5-7, characterized in that, The step of determining the starting address of the memory space of the thread bundle according to the starting addresses of the pre-partitioned memory spaces of the thread bundle and each thread bundle fragment request amount in the thread bundle includes: For each first thread bundle fragment amount in the thread bundle, determine the starting address of the pre-partitioned memory space corresponding to the first thread bundle fragment amount as the starting address of the memory space of the thread bundle; For each second thread bundle fragment amount in the thread bundle, allocate at least one minimum prefetch subspace for the second thread bundle fragment amount according to the second thread bundle fragment amount and the free state of the full binary tree nodes corresponding to the integrated memory space of the second thread bundle fragment amount; the leaf nodes of the full binary tree correspond to the minimum prefetch subspaces in the integrated memory space; the minimum regional subspaces corresponding to each of the leaf nodes are different; Determine the starting address of the memory space corresponding to the second thread bundle fragment amount of the thread bundle according to the starting address of the allocated minimum prefetch subspace; Determine the starting address of the memory space corresponding to the second thread bundle fragment amount of the thread bundle as the starting address of the memory space of the thread bundle.
9. A memory space allocation device for a thread, characterized in that, Configured on the GPU side, the device includes: A request sending module, configured to send the memory request amounts of at least one thread bundle to the host side, so that the host side allocates starting addresses of memory spaces for each of the thread bundles according to the memory request amounts of the thread bundles, and feeds back the starting addresses of the memory spaces of each of the thread bundles to the GPU side; A memory splitting module, configured to, for each thread bundle, split the memory request amount of each thread in the thread bundle according to the minimum scheduling memory amount, to obtain at least one thread fragment request amount for each of the threads; An offset determining module, configured to alternately determine at least one memory address offset for each thread in the thread bundle according to the offset arrangement order among the threads in the thread bundle and the thread fragment request amounts of the threads in the thread bundle; A memory space determining module, configured to, for each thread in the thread bundle, allocate the starting address of the memory space of the thread bundle and each of the memory address offsets of the thread to the thread.
10. A memory space allocation device for a thread, characterized in that, Configured on the host side, the device includes: A receiving module, configured to receive the memory request amounts of at least one thread bundle sent by the GPU side; A splitting module, configured to, for each thread bundle, split the memory request amount of the thread bundle according to the pre-divided space memory amount, to obtain at least one thread bundle fragment request amount; An allocation module, configured to allocate free pre-divided memory spaces for each thread bundle fragment request amount of each of the thread bundles according to the thread bundle fragment request amounts of each of the thread bundles; A determining module, configured to, for each thread bundle, determine the starting address of the memory space of the thread bundle according to the starting addresses of the pre-divided memory spaces of the thread bundle and the thread bundle fragment request amounts in the thread bundle; A sending module, configured to send the starting address of the memory space of the thread bundle to the GPU side, so that the GPU side allocates the starting address of the memory space and the memory address offset for at least one thread in the thread bundle according to the starting address of the memory space of the thread bundle.
11. An electronic device, characterized in that, The device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to execute the method for allocating the memory space of the thread according to any one of claims 1-8.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and when the computer instructions are used to be executed by a processor, the method for allocating the memory space of the thread according to any one of claims 1-8 is implemented.
Citation Information
Cited By
Virtual memory management method and system, graphics processor and electronic equipment
CN121579188A