GPGPU thread bundle scheduling optimization method and device based on reordering, and medium
By improving the scheduling strategy of GPGPU thread bundles and using type registers and queue buffers for thread bundle classification and scheduling, the problem of low thread bundle scheduling efficiency in the existing technology is solved, and more efficient resource utilization and computing efficiency is achieved.
Patent Information
- Application Number
- CN202510030250.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-16
AI Technical Summary
The existing GPGPU thread bundle scheduling methods are difficult to make full use of the locality between thread bundles, resulting in cache misses and waste of computing resources, and it is difficult to effectively balance the allocation of computing resources, resulting in idle resources and reducing overall computing efficiency.
By introducing new scheduler design and scheduling strategies, the execution order of thread bundles is improved, and the type register and queue buffer are used to classify thread bundles into long thread bundles and short thread bundles, and the short thread bundles are preferentially scheduled through the queue controller, and the fixed insertion strategy is inserted into waiting thread bundles, and the queue order is dynamically adjusted to improve resource utilization and computing efficiency.
It realizes more efficient parallel task processing, improves the resource utilization and computing efficiency of GPGPU, reduces waiting time, retains the spatial locality between thread bundles, adapts to different workloads and access modes, and stabilizes the performance of GPGPU.
Smart Images

Figure CN120011011A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of parallel computing, and more particularly to a method, device and medium for optimizing GPGPU warp scheduling based on reordering. Background Art
[0002] Currently, general-purpose graphics processing units (GPGPUs), as an important platform for parallel computing, have been widely used in scientific computing, data analysis, and machine learning. They use graphics processing units for general computing tasks, achieving efficient support for large-scale parallel data processing. However, when GPGPUs perform complex computing tasks, their thread scheduling efficiency becomes one of the key factors that restrict performance.
[0003] Existing warp scheduling methods, such as Loose Round-Robin (LRR), may not fully utilize the locality between warps when processing diverse and complex computing tasks, resulting in cache misses and waste of computing resources. Due to the differences in memory access patterns of warps, when a large number of warps request memory resources at the same time, cache pollution and jitter problems may occur, which will further affect the performance of GPGPU.
[0004] In addition, different warps may have different execution times and access delays. For example, some warps may have short execution times and low access delays, while others may contain uncertain operations or require long access to external memory. This difference makes it difficult for existing scheduling methods to effectively balance the allocation of computing resources, causing some computing resources to be idle while waiting, thereby reducing the overall computing efficiency.
[0005] In order to overcome the above challenges, a smarter and more complete warp scheduling strategy is needed. This strategy should be able to effectively hide long-latency operations, reduce the loss of locality between warps, and adapt to different workloads and access patterns. Although some research works have been devoted to improving GPGPU thread scheduling, how to design a scheduling method that can effectively maintain data locality and reduce the impact of long latencies is still a technical problem that needs to be solved.
[0006] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present application, and therefore may include information that does not constitute the prior art known to ordinary technicians in the field. Summary of the invention
[0007] In order to provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. The summary is not an extensive review, nor is it intended to identify key / critical elements or delineate the scope of protection of these embodiments, but rather serves as a prelude to the detailed description that follows.
[0008] The disclosed embodiments provide a GPGPU warp scheduling optimization method, device, and medium based on reordering, which improves the execution order of warps to achieve more efficient parallel task processing. The method improves the resource utilization and computing efficiency of GPGPU by introducing a new scheduler design and scheduling strategy, thereby meeting the increasingly complex parallel computing needs.
[0009] In some embodiments, the method comprises: The scheduler is introduced with a type register for storing the warp type, a queue buffer for storing the warp scheduling order, and a queue controller for controlling the warp scheduling queue; Classify the warps into long warps and short warps according to their behaviors, and store the results in the type register, where the long warps include uncertain operation warps and long access delay warps; The queue controller accesses the type register, divides the classified thread warps into active thread warps and waiting thread warps, and puts the active thread warps into the front of the queue buffer first, and puts the waiting thread warps into the back of the queue buffer; The queue controller inserts a waiting thread warp into the active thread warp queue, wherein the insertion strategy is a fixed insertion strategy, and the fixed insertion strategy is to insert a waiting thread warp every other active thread warp; Schedule the thread warps for execution according to the adjusted queue order.
[0010] Preferably, the uncertain operation warp is a warp whose operation time is uncertain, the long access delay warp is a warp whose off-chip memory access operation time is long, and the short warp is a warp whose operation time is certain or whose operation time is short and access delay is small.
[0011] Preferably, the thread warp classification method is as follows: the instruction executed by the current thread warp is received by the decoding module, a type flag is issued according to the instruction type, and the type register receives the instruction, and the type register stores the thread warp number after matching the thread warp type.
[0012] Preferably, the short warp is used as the active warp, and the long warp is used as the waiting warp.
[0013] Preferably, the queue buffer adopts a priority queue mechanism, and for the thread warps that are also in the active thread warp queue or the waiting thread warp queue, the priorities are assigned according to the thread warp numbers.
[0014] The insertion strategy also includes a software insertion strategy that takes into account the current state of the pipeline, cache pressure, and expected execution time. Preferably, the specific strategy of the insertion strategy is as follows: the priority value is calculated by the following formula , insert the thread bundle with large priority value to the front of the queue. The formula is as follows: , in, Indicates whether to handle the stagnation state. A value of 1 indicates a stagnation state, and a value of 0 indicates a non-stagnation state. Indicates the cache hit rate, Indicates the estimated execution time; For thread warps with a priority value of 0, continue to calculate their stall consumption value , inserting the thread bundle with large stall consumption value to the front of the queue is as follows: .
[0015] Preferably, it also includes a backtracking and re-arrangement step, when it is found that the waiting time after scheduling remains unchanged or the spatial locality becomes worse, the queue controller reclassifies, arranges and inserts the thread warps.
[0016] In some embodiments, the apparatus includes: a processor and a memory storing program instructions, and the processor is configured to execute the reordering-based GPGPU warp scheduling optimization method when running the program instructions.
[0017] In some embodiments, the storage medium stores program instructions, and when the program instructions are run, the reordering-based GPGPU warp scheduling optimization method is executed.
[0018] The embodiments of the present disclosure provide a GPGPU warp scheduling optimization method, device, and medium based on reordering, which can achieve the following technical effects: By intelligently classifying and sorting thread warps, the present invention can preferentially schedule short thread warps with short operation time and small access delay, thereby reducing the overall waiting time of the system and improving the utilization rate of computing resources.
[0019] The long warps (including uncertain operation warps and long access delay warps) are placed in the waiting queue and inserted into the active queue at the appropriate time. This strategy effectively hides the long delay operations and avoids the problem of overall system performance degradation caused by a single long delay warp.
[0020] The present invention introduces a fixed insertion strategy, and can retain the spatial locality between thread warps as much as possible while adjusting the thread warp scheduling order, thereby reducing the pipeline stagnation time and improving the computing efficiency.
[0021] The scheduling method of the present invention not only considers the static characteristics of the thread warp (such as operation time and access delay), but also dynamically adapts to various factors in the execution process (such as dependencies between thread warps and cache pressure). Through the backtracking and reordering functions of the queue controller, the system can maintain efficient scheduling performance when facing complex and dynamically changing workloads.
[0022] Through intelligent scheduling strategies and dynamic adjustment mechanisms, the present invention can ensure that the performance of GPGPU remains stable under different workloads and operating conditions, and reduce performance fluctuations caused by improper scheduling strategies.
[0023] In summary, the reordering-based GPGPU warp scheduling optimization method of the present invention has achieved remarkable beneficial effects in improving GPGPU computing efficiency, reducing waiting time, preserving spatial locality, enhancing adaptability and flexibility, and improving system stability.
[0024] The above general description and the following description are exemplary and explanatory only and are not intended to limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] One or more embodiments are exemplarily described by corresponding drawings, which do not limit the embodiments. Elements with the same reference numerals in the drawings are shown as similar elements, and the drawings do not constitute a scale limitation, and wherein: Figure 1 Schematic diagram of the process of the present invention; Figure 2 is an overview diagram of the scheduler of this invention; Figure 3 is a schematic diagram of the fixed insertion strategy of this invention; Figure 4 It is a schematic diagram of a practical example of this invention; Figure 5 It is a logical schematic diagram of this invention; Figure 6 It is a schematic diagram of a device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0026] In order to be able to understand the features and technical contents of the embodiments of the present disclosure in more detail, the implementation of the embodiments of the present disclosure is described in detail below in conjunction with the accompanying drawings. The attached drawings are for reference only and are not used to limit the embodiments of the present disclosure. In the following technical description, for the convenience of explanation, a full understanding of the disclosed embodiments is provided through multiple details. However, one or more embodiments can still be implemented without these details. In other cases, to simplify the drawings, well-known structures and devices can be simplified for display.
[0027] The terms "first", "second", etc. in the specification and claims of the embodiments of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged where appropriate, so as to describe the embodiments of the embodiments of the present disclosure described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions.
[0028] like Figure 1 and Figure 5 As shown in the figure, a GPGPU warp scheduling optimization method based on reordering is presented. The present invention redesigns the scheduler and introduces multiple key components: a type register for storing warp types, a queue buffer for storing warp scheduling order, and a queue controller for controlling warp scheduling queues. These components work together to achieve efficient warp scheduling and management. Figure 1 As shown in the figure, the redesigned scheduler can classify thread warps before scheduling, arrange thread warps according to classification, and control the insertion of thread warps outside the queue. The following is a detailed explanation of each operation: S1: introducing a scheduler into a type register for storing a warp type, a queue buffer for storing a warp scheduling order, and a queue controller for controlling a warp scheduling queue; S2: Classify the warps into long warps and short warps according to their behaviors and store the results in the type register, where long warps include uncertain operation warps and long access delay warps; S3: The queue controller accesses the type register, divides the classified thread warps into active thread warps and waiting thread warps, and puts the active thread warps into the front of the queue buffer first, and puts the waiting thread warps into the back of the queue buffer; S4: the queue controller inserts a waiting thread warp into the active thread warp queue, wherein the insertion strategy is a fixed insertion strategy, and the fixed insertion strategy is to insert a waiting thread warp every other active thread warp; S5: Schedule the thread warps for execution according to the adjusted queue order.
[0029] As a refinement of the above embodiment, the uncertain operation warp is a warp whose operation time cannot be determined, the long access delay warp is a warp whose execution time of off-chip memory access operation is long, and the short warp is a warp whose operation time is determined or whose operation time is short and access delay is small.
[0030] As a refinement of the above embodiment, the thread warp classification method is as follows: the instruction executed by the current thread warp is received by the decoding module, and a type flag is issued according to the instruction type, which is received by the type register, and the type register stores the thread warp number after matching it with its type.
[0031] In order to distinguish the thread warps, the decoding module will issue a type flag according to the instruction type after receiving the instruction executed by the current thread warp, and the type flag will be received by the type register. The type register will store the thread warp number and its type in correspondence for the queue control module to access and perform subsequent operations.
[0032] As a refinement of the above embodiment, Figure 3 As shown in (a), all short thread warps will be treated as active thread warps, obtain priority scheduling qualifications, and will be placed in the front position of the queue buffer; while long thread warps will be treated as waiting thread warps and will be placed in the back position of the queue buffer, waiting for subsequent scheduling.
[0033] As a refinement of the above embodiment, the queue controller also considers the potential dependencies between the warps when classifying the warps. This means that even if some short warps have higher execution priorities, they may be temporarily placed in the waiting queue because they depend on the execution results of other warps. In addition, the queue buffer adopts a priority queue mechanism. For warps that are also in the active warp queue or the waiting warp queue, the priority is assigned according to the warp number size to support flexible warp scheduling and priority adjustment.
[0034] As a refinement of the above embodiment, the queue controller is responsible for inserting the thread warps in the waiting thread warp queue into the active thread warp queue. The insertion strategy is divided into a fixed insertion strategy and a software algorithm strategy. The fixed insertion strategy means inserting a waiting thread warp every other active thread warp, such as Figure 3 As shown in (b) of , the purpose of this is to ensure the spatial locality between thread warps as much as possible. After the queue insertion is completed, all thread warps are assigned to one queue for scheduling, see Figure 3 In (c), we can find that compared with the situation before insertion where only warp 3 and warp 2 have spatial locality, warp 2 and warp 1, and warp 4 and warp 3 all have spatial locality in the queue after insertion, indicating that this method can adjust the order of warp scheduling queues to reduce waiting time while retaining the spatial locality between warps.
[0035] The software insertion strategy needs to be combined with a software algorithm to evaluate the best location for the insertion point. The algorithm takes into account the current state of the pipeline, cache pressure, and expected execution time to maximize resource utilization and minimize latency.
[0036] The insertion strategy also includes a software insertion strategy that takes into account the current state of the pipeline, cache pressure, and expected execution time. The specific strategy is as follows: The priority value is calculated by the following formula , insert the thread bundle with large priority value to the front of the queue. The formula is as follows: , in, Indicates whether to handle the stagnation state. A value of 1 indicates a stagnation state, and a value of 0 indicates a non-stagnation state. Indicates the cache hit rate, Indicates the estimated execution time; For thread warps with a priority value of 0, continue to calculate their stall consumption value , inserting the thread bundle with large stall consumption value to the front of the queue is as follows: .
[0037] As a refinement of the above embodiment, it also includes backtracking and reordering steps. When it is found that the waiting time after scheduling remains unchanged or the spatial locality deteriorates, the queue controller reclassifies, arranges and inserts the thread warps.
[0038] During the insertion process, the queue controller may also make dynamic adjustments based on real-time feedback. For example, when it is found that the thread warps in the active thread warp queue are stalled due to waiting for memory access, some waiting thread warps may be scheduled preferentially to improve execution efficiency. In addition, the queue controller allows backtracking and rescheduling when necessary to correct suboptimal scheduling caused by incorrect predictions or changes in external conditions (such as unchanged waiting time after scheduling or deterioration of spatial locality). When rescheduling is required, the queue controller in the scheduler re-divides the remaining unscheduled thread warps into the initial active thread warp queue and waiting thread warp queue according to the thread warp type, and reselects the appropriate scheduling strategy. This adaptive capability ensures that the scheduling strategy remains efficient even in the face of complex and dynamically changing workloads.
[0039] The following is a detailed description of this method with specific examples: like Figure 4 As shown, attached Figure 4 (a) shows the computation time and waiting time of thread warps 0-4 during round-robin scheduling. It can be seen that there are different delays between thread warps, which may be caused by the difference in the types of instructions executed by the thread warps and the time to access memory. For example, thread warp 1 may have performed some shorter computational tasks, so its delay is shorter; while thread warp 0 may have performed some operations that require access to external memory, resulting in its longer delay. At the same time, the figure marks the spatial locality between thread warps (arrows on the left of the thread warps), and it can be seen that round-robin scheduling can make good use of the spatial locality of thread warps.
[0040] The decoding module classifies the thread bundles according to the instructions, and classifies thread bundles 1, 3, and 4 into short thread bundles, and thread bundles 0 and 2 into long thread bundles. According to the classification of thread bundles, short thread bundles are placed in the active thread bundle queue, and long thread bundles are placed in the waiting thread bundle queue. Figure 4 (b) It can be found that prioritizing the scheduling of short warps can reduce the waiting time of the system, but correspondingly, the spatial locality between warp scheduling is destroyed, and only warps 3 and 4 still have locality.
[0041] Then, a suitable insertion strategy is selected for insertion. Here, the fixed insertion strategy of hardware is taken as an example. Every other long warp in the waiting warp queue is inserted once. Figure 4 (c). As shown in the figure, it can be found that while the waiting time does not change, the spatial locality between the thread warps is enhanced (thread warps 1, 0, thread warps 3, 2). If the software algorithm strategy is selected, a more suitable scheduling solution can be obtained, which corresponds to the increase in computational complexity and real-time challenges. At the same time, it may face the disadvantages of increased tuning difficulty and reduced hardware dependence.
[0042] If the post-scheduling solution is not ideal, such as the waiting time remains unchanged or the spatial locality deteriorates, the queue controller will backtrack and rearrange when necessary. The specific process has been mentioned in the previous article and will not be repeated here.
[0043] It should be noted that: Through this comprehensive warp scheduling method, the present invention provides an efficient and flexible solution to optimize the performance of GPGPU when processing parallel computing tasks. The overall advantage is that the method not only considers the static characteristics of the warp, such as operation time and access latency, but also dynamically adapts to various factors in the execution process, including dependencies between warps and cache pressure. This adaptability significantly improves resource utilization, reduces waiting time, and can ensure high performance under different workloads and operating conditions.
[0044] In addition, the scheduling method of the present invention not only retains the spatial locality between thread bundles by intelligently rearranging thread bundles, but also ensures the flexibility and accuracy of thread bundle scheduling even in the face of complex thread dependencies and execution conditions through the priority queue mechanism and adaptive adjustment strategy. The backtracking and reordering functions of the queue controller further improve the robustness of the scheduling strategy, allowing the system to make quick adjustments when prediction errors or environmental changes occur, thereby ensuring the long-term effectiveness of scheduling decisions and the stability of the system.
[0045] Combination Figure 6As shown, the embodiment of the present disclosure provides a GPGPU warp scheduling optimization device 300 based on reordering, including a processor (processor) 304 and a memory (memory) 301. Optionally, the device may also include a communication interface (Communication Interface) 302 and a bus 303. Among them, the processor 304, the communication interface 302, and the memory 301 can communicate with each other through the bus 303. The communication interface 302 can be used for information transmission. The processor 304 can call the logic instructions in the memory 301 to execute the GPGPU warp scheduling optimization method based on reordering of the above embodiment.
[0046] In addition, the logic instructions in the memory 301 described above can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product.
[0047] The memory 301 is a computer-readable storage medium that can be used to store software programs and computer executable programs, such as program instructions / modules corresponding to the method in the embodiment of the present disclosure. The processor 304 executes the functional application and data processing by running the program instructions / modules stored in the memory 301, that is, the GPGPU thread warp scheduling optimization method based on reordering in the above embodiment is implemented.
[0048] The memory 301 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and an application required for at least one function; the data storage area may store data created according to the use of the terminal device, etc. In addition, the memory 301 may include a high-speed random access memory and may also include a non-volatile memory.
[0049] An embodiment of the present disclosure provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are configured to execute the above-mentioned GPGPU warp scheduling optimization method based on reordering.
[0050] The computer-readable storage medium mentioned above may be a transient computer-readable storage medium or a non-transitory computer-readable storage medium.
[0051] The technical solution of the embodiment of the present disclosure can be embodied in the form of a software product, which is stored in a storage medium and includes one or more instructions for enabling a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the embodiment of the present disclosure. The aforementioned storage medium may be a non-transient storage medium, including: a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a disk or an optical disk, and other media that can store program codes, or a transient storage medium.
[0052] The above description and the accompanying drawings fully illustrate the embodiments of the present disclosure so that those skilled in the art can practice them. Other embodiments may include structural, logical, electrical, process and other changes. The embodiments represent only possible changes. Unless explicitly required, separate components and functions are optional, and the order of operations may vary. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. Moreover, the words used in this application are only used to describe the embodiments and are not used to limit the claims. As used in the description of the embodiments and the claims, unless the context clearly indicates, the singular forms of "a", "an" and "the" are intended to include plural forms as well. Similarly, the term "and / or" as used in this application refers to any and all possible combinations of one or more associated listings. In addition, when used in this application, the term "comprise" and its variants "comprises" and / or comprising refer to the presence of stated features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or groups thereof. In the absence of further restrictions, the elements defined by the sentence "including one..." do not exclude the presence of other identical elements in the process, method or device including the elements. In this article, each embodiment may focus on the differences from other embodiments, and the same and similar parts between the embodiments may refer to each other. For the methods, products, etc. disclosed in the embodiments, if they correspond to the method part disclosed in the embodiments, then the relevant parts can refer to the description of the method part.
[0053] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software may depend on the specific application and design constraints of the technical solution. The technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the embodiments of the present disclosure. The technicians may clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here.
[0054] In the embodiments disclosed herein, the disclosed methods and products (including but not limited to devices, equipment, etc.) can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units can be only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between each other shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms. The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the units may be selected according to actual needs to implement this embodiment. In addition, each functional unit in the embodiment of the present disclosure may be integrated in a processing unit, or each unit may exist physically separately, or two or more units may be integrated in one unit.
[0055] The flowchart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to the embodiment of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. In some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, which can depend on the functions involved. In the description corresponding to the flowchart and the block diagram in the accompanying drawings, the operations or steps corresponding to different boxes can also occur in a different order from the order disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, which can depend on the functions involved. Each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified functions or actions, or may be implemented by a combination of dedicated hardware and computer instructions.
Claims
1. A GPGPU warp scheduling optimization method based on reordering, characterized in that: include: The scheduler is introduced with a type register for storing the warp type, a queue buffer for storing the warp scheduling order, and a queue controller for controlling the warp scheduling queue; Classify the warps into long warps and short warps according to their behaviors, and store the results in the type register, where the long warps include uncertain operation warps and long access delay warps; The queue controller accesses the type register, divides the classified thread warps into active thread warps and waiting thread warps, and puts the active thread warps into the front of the queue buffer first, and puts the waiting thread warps into the back of the queue buffer; The queue controller inserts a waiting thread warp into the active thread warp queue, wherein the insertion strategy is a fixed insertion strategy, and the fixed insertion strategy is to insert a waiting thread warp every other active thread warp; Schedule the thread warps for execution according to the adjusted queue order.
2. The GPGPU warp scheduling optimization method based on reordering according to claim 1, characterized in that: The uncertain operation warp is a warp whose operation time is uncertain, the long access delay warp is a warp whose off-chip memory access operation time is long, and the short warp is a warp whose operation time is certain or whose operation time is short and access delay is small.
3. The GPGPU warp scheduling optimization method based on reordering according to claim 1, characterized in that: The thread warp classification method is as follows: the instruction executed by the current thread warp is received through the decoding module, and a type flag is issued according to the instruction type. The type flag is received by the type register, and the type register stores the thread warp number after matching it with its type.
4. The GPGPU warp scheduling optimization method based on reordering according to claim 1, characterized in that: The short warp is used as the active warp, and the long warp is used as the waiting warp.
5. The GPGPU warp scheduling optimization method based on reordering according to claim 1, characterized in that: The queue buffer adopts a priority queue mechanism. For the thread warps that are also in the active thread warp queue or the waiting thread warp queue, the priority is specified according to the thread warp number size.
6. The GPGPU warp scheduling optimization method based on reordering according to claim 1, characterized in that: It also includes backtracking and rescheduling steps. When it is found that the waiting time remains unchanged or the spatial locality deteriorates after scheduling, the queue controller reclassifies, arranges and inserts the thread warps.
7. The GPGPU warp scheduling optimization method based on reordering according to claim 1, characterized in that: The insertion strategy also includes a software insertion strategy that takes into account the current state of the pipeline, cache pressure, and expected execution time.
8. The GPGPU warp scheduling optimization method based on reordering according to claim 7, characterized in that: The insertion strategy is as follows: the priority value U is calculated by the following formula, and the thread warp with a large priority value is inserted to the front of the queue. The formula is as follows: U=(1-T)*H / (S-1) Among them, T indicates whether the processing is in a stagnant state, a value of 1 indicates a stagnant state, a value of 0 indicates a non-stagnant state, H indicates the cache hit rate, and S indicates the estimated execution time; For the thread bundle with a priority value of 0, continue to calculate its stagnation consumption value X, and insert the thread bundle with a large stagnation consumption value to the front of the queue. The formula is as follows: X=H / (S-1)。 9. A GPGPU warp scheduling optimization device based on reordering, comprising a processor and a memory storing program instructions, characterized in that: The processor is configured to execute the reordering-based GPGPU warp scheduling optimization method according to any one of claims 1 to 8 when running the program instructions.
10. A storage medium storing program instructions, characterized in that: When the program instructions are executed, the reordering-based GPGPU warp scheduling optimization method according to any one of claims 1 to 8 is executed.
Citation Information
Cited By
GPU Warp scheduling method and device based on chessboard scheduling strategy
CN121722532A