A conflicting thread warp scheduling method and a GPGPU register file access method and system
By identifying the address conflicts of thread bundle operation registers and adjusting the scheduling sequence in real time, combined with the design of the conflict controller and operand collector, the problems of register access conflicts and thread bundle scheduling delay in GPGPU are solved, which significantly improves the execution efficiency and performance of GPGPU.
Patent Information
- Application Number
- CN202510026863.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-08
AI Technical Summary
While meeting the high bandwidth memory access requirements, register files in GPGPU are difficult to effectively utilize multi-port design, resulting in register access collisions and thread bundle scheduling delays.
By obtaining and identifying address conflicts in thread bundle operation registers, adjusting the scheduling order of thread bundles in real time to avoid conflicts, and combining the design of conflict controllers and operand collectors, efficient parallel access to register files is achieved.
It effectively avoids thread bundle scheduling delay caused by register access conflicts, improves the execution efficiency and performance of GPGPUs, and realizes the full utilization of GPGPU resources.
Smart Images

Figure CN119440774B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of GPGPU warp scheduling, and in particular relates to a conflict warp scheduling method and a GPGPU register file memory access method and system. Background Art
[0002] GPGPU is the abbreviation of General-Purpose Graphics Processing Unit, a general-purpose graphics processor.
[0003] GPGPU processes data at high speed by scheduling a large number of warps alternately. Therefore, the register files in GPGPU are usually designed with large capacity, so that the warps in the chassis can be switched with zero overhead. In addition to the capacity requirement, the register files of GPGPU also need high memory access bandwidth to meet the needs of parallel reading and writing of warp operands. Warp scheduling selects ready warps for scheduling and execution based on the available storage resources and computing resources in GPGPU. Therefore, the way warps are scheduled will directly affect the execution efficiency of the entire GPGPU task.
[0004] However, the area of the register file expands rapidly with the increase of read and write ports, making it difficult to adopt a multi-port design. In order to meet the high-bandwidth memory access requirements of registers and limit the area of register files, GPGPU register files usually use a multi-block storage structure to achieve parallel access. In the multi-block design using dual-port memory, the number of register read and write ports is the same, but the GPGPU program mainly includes: branches, jumps, and control instructions that only read multiple operands; control instructions that only write back a single operand; and operation instructions that require both multiple operand reads and single operand write back. Due to the high frequency of branch, jump, and operation instructions, the multi-block design of the register read port has significantly improved access performance, but the frequency of control instructions that only write back is low and operation instructions only write back a single operand. The multi-block design of the register write port has low storage bandwidth utilization, resulting in a waste of port resources. Therefore, the multi-block single-port storage structure commonly used by GPGPU at present is bound to have the conflict between the high-bandwidth memory access requirements of registers and the thread bundle on a single port. Summary of the invention
[0005] In a first aspect, an embodiment of the present application provides a conflict warp scheduling method, comprising the following steps:
[0006] S1. Obtain the first address of the register that the scheduled warp needs to operate, and the second address of the register that the to-be-scheduled warp needs to operate;
[0007] S2. Identify whether there is a conflict between the second addresses, and when there is a conflict, select a schedule from the two to-be-scheduled thread warps corresponding to the two conflicting second addresses, and when there is no conflict, proceed to step S3;
[0008] S3. Identify whether there is a conflict between the first address and the second address, and when there is no conflict, schedule the thread warp to be scheduled in a polling manner, and when there is a conflict, suspend the corresponding thread warp to be scheduled and execute the part of the thread warp to be scheduled that does not conflict with the first address. In this solution, by obtaining and identifying the address conflict of the thread warp operation register, the thread warp scheduling delay caused by the register access conflict is avoided, and the execution efficiency of the GPGPU is improved.
[0009] Furthermore, the method further comprises the following steps:
[0010] SS1. Set a conflict controller to record the read and write register information of the thread warp instructions, mark the status of each register address being used, and adjust the scheduling order of the thread warp to be scheduled in real time according to the usage status of the marked register address. In this preferred solution, the setting of the conflict controller can record the read and write register information of the thread warp instructions in real time, mark the usage status of the register address, and adjust the scheduling order of the thread warp accordingly, reducing the probability of conflict.
[0011] Further, step SS1 comprises the following steps:
[0012] SS11. The conflict controller marks the status of each register address being used through the conflict control table;
[0013] SS12. The conflict controller selects threads to be scheduled that do not conflict according to the conflict control table through the scheduler. In this preferred solution, the conflict control table and the scheduler work together to accurately select threads to be scheduled that do not conflict, further improving the performance of GPGPU.
[0014] In a second aspect, an embodiment of the present application further provides a GPGPU register file memory access method using the conflicting warp scheduling method of the first aspect, comprising the following steps:
[0015] S101. The thread warp scheduling unit selects a thread warp to be scheduled that does not conflict according to the conflicting thread warp scheduling method for scheduling;
[0016] S102. The operand access unit reads or writes operands from the single-port multi-address register file in parallel according to the instructions of the selected scheduling warp. In this solution, combined with the conflict warp scheduling method, efficient parallel access to the single-port multi-address register file is achieved, improving the data reading and writing speed.
[0017] Furthermore, the specific steps of step S101 are as follows:
[0018] S1011. The thread warp scheduling unit fetches and decodes instructions from the thread warp to be scheduled through the instruction prefetch queue, and stores the decoded instructions into the instruction queue;
[0019] S1012. The warp scheduling unit records the read and write register information of the warp instruction through the conflict controller, and marks the status of each register address being used, and then adjusts the scheduling order of the warp to be scheduled in real time according to the usage status of the marked register address, and selects the warp to be scheduled without address conflict;
[0020] S1013. The thread warp scheduling unit stores the selected thread warps to be scheduled in the scheduling queue through the conflict controller. In this preferred solution, the coordinated work of the instruction prefetch queue, the conflict controller and the scheduling queue realizes the accurate scheduling of thread warp instructions, avoids register access conflicts, and improves the computing efficiency of GPGPU.
[0021] Furthermore, the specific steps of step S102 are as follows:
[0022] S1021. The operand access unit records the register address of the operand of the instruction in the thread warp, the status of whether the operand is executed, and the operand data through the operand collector;
[0023] S1022. The operand access unit arbitrates the address entries of the register file that can be accessed according to the operand collector, reads and writes operands from the single-port multi-address register file in parallel, and sets the status of whether the operand is executed in the operand collector when the operand of a certain register address is successfully read or written back. In this preferred solution, the setting of the operand collector can accurately record the address, status and data of the operand, and the memory access executor can access the register file in parallel based on this, thereby improving the parallelism and efficiency of data access.
[0024] In a third aspect, an embodiment of the present application further provides a GPGPU register file memory access system using the conflicting warp scheduling method described in the second aspect, including:
[0025] The thread warp scheduling unit selects a thread warp to be scheduled without conflict by using a conflicting thread warp scheduling method for scheduling;
[0026] an operand access unit, which reads or writes back operands in parallel from a single-port multi-address register file according to the instructions of the selected scheduling warp;
[0027] After completing the operation, the execution unit sends the operand write-back request back to the warp scheduling unit and enters the warp to be scheduled. In this solution, the conflict warp scheduling method and the efficient register file access method are combined to achieve full utilization of GPGPU resources and improve the overall performance of the system.
[0028] Furthermore, the warp scheduling unit includes:
[0029] The instruction pre-fetch queue fetches instructions from the thread warp to be scheduled for execution, decodes them and stores them in the instruction queue;
[0030] The conflict controller records the read and write register information of the thread bundle instruction, marks the status of each register address being used, and schedules the threads to be scheduled that do not have conflicts according to the marked register addresses;
[0031] The scheduling queue stores selected non-conflicting thread warps to be scheduled, waiting for scheduling. In this preferred solution, the coordinated work of the instruction prefetch queue, the conflict controller and the scheduling queue realizes accurate scheduling and conflict avoidance of thread warp instructions, thereby improving the computing efficiency of GPGPU.
[0032] Furthermore, the operand access unit includes:
[0033] An operand collector records the register address of the operands of instructions in the thread warp, the status of whether the operands are executed, and the operand data;
[0034] A single-port multi-block register file, including a number of single-port memories, stores the operands corresponding to the operation objects of the thread warp and allows read or write operations to be performed in the same cycle;
[0035] The memory access executor reads and writes operands from the single-port multi-address register file in parallel according to the address entries of the register file that can be accessed by the operand collector, and sets the execution completion status of the operand in the operand collector when the operand of a certain register address is successfully read or written back. In this preferred solution, through the combination of the operand collector and the single-port multi-block register file, efficient parallel access and storage of operands are achieved, thereby improving the speed and efficiency of data processing.
[0036] Furthermore, the operand collector includes:
[0037] Operand read collection table, including N operand entries;
[0038] The operand write-back collection table includes an operand entry;
[0039] Each operand entry has:
[0040] The valid bit field marks whether the operation data corresponding to each entry needs to be executed;
[0041] The register ID field records the register number of the operand corresponding to each entry;
[0042] The data field stores the operand data corresponding to each entry;
[0043] The ready bit field marks whether each entry has completed the operation of the corresponding operation data. In this preferred solution, through the setting of the operand read collection table and the operand write back collection table, as well as the detailed field design of each operand entry, the state and data of the operand can be accurately recorded and managed, providing strong support for efficient data access and processing.
[0044] It can be seen from the above technical solutions that the present invention has the following advantages:
[0045] The conflicting thread warp scheduling method and GPGPU register file access method and system provided by this application accurately identify the address conflicts of thread warp operation registers, adjust the scheduling order of thread warps in real time, and avoid thread warp scheduling delays caused by register access conflicts; at the same time, combined with the design of efficient register file access methods and operand collectors, efficient parallel access and storage of operands are achieved, improving the speed and efficiency of data processing. Overall, this application can significantly improve the execution efficiency and performance of GPGPU, and provide strong support for applications in fields such as high-performance computing and graphics processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solution of the present invention, the accompanying drawings required for use in the description will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0047] Figure 1 It is a flowchart of the conflicting warp scheduling method of the present invention.
[0048] Figure 2 The present invention is a flowchart of a GPGPU register file memory access method of the application conflict warp scheduling method.
[0049] Figure 3 The present invention provides a GPGPU register file memory access system for an application conflict warp scheduling method. DETAILED DESCRIPTION
[0050] In the specific steps of the conflicting warp scheduling method described in detail below, various embodiments of the present disclosure will be described more fully. The present disclosure may have various embodiments, and adjustments and changes may be made therein. However, it should be understood that there is no intention to limit the various embodiments of the present disclosure to the specific embodiments disclosed herein, but rather the present disclosure should be understood to cover all adjustments, equivalents and / or alternatives that fall within the spirit and scope of the various embodiments of the present disclosure.
[0051] For example, a general-purpose graphics processing unit (GPGPU) performs high-speed data processing by alternately scheduling a large number of thread bundles. Therefore, the register file in the GPGPU is usually designed with large capacity to achieve zero-overhead switching of thread bundles. In addition to the capacity requirement, the GPGPU register file also requires high memory access bandwidth to meet the needs of parallel reading and writing of thread bundle operands. Thread bundle scheduling selects ready thread bundles for scheduling and execution based on the status of available storage resources and computing resources in the GPGPU. The thread bundle scheduling strategy will directly affect the execution efficiency of the entire GPGPU task.
[0052] However, the area of the register file expands rapidly with the increase of read and write ports, making it difficult to adopt a more radical multi-port design. In order to meet the high-bandwidth memory access requirements of registers and limit the area of register files, GPGPU register files usually use a multi-block storage structure to achieve parallel access. In a multi-block design using dual-port memory, the number of register read and write ports is the same, but the GPGPU program mainly includes: branches, jumps, and control instructions that only read multiple operands; control instructions that only write back a single operand; and operation instructions that require both multiple operand reads and single operand write back. Due to the high frequency of branch, jump, and operation instructions, the multi-block design of the register read port achieves significant access performance improvement, but the control instructions that only write back have a low frequency of occurrence and the operation instructions only write back a single operand. The multi-block design of the register write port has low storage bandwidth utilization, resulting in a waste of port resources.
[0053] In view of the above problems, this embodiment provides a conflicting warp scheduling method, which avoids warp scheduling delays caused by register access conflicts by acquiring and identifying address conflicts of warp operation registers, thereby improving the execution efficiency of GPGPU.
[0054] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0055] See also Figure 1 The figure is a flowchart of a method for scheduling conflicting warps in a specific embodiment, the method comprising the following steps:
[0056] S1. Obtain the first address of the register that the scheduled warp needs to operate, and the second address of the register that the to-be-scheduled warp needs to operate;
[0057] S2. Identify whether there is a conflict between the second addresses, and when there is a conflict, select a schedule from the two to-be-scheduled thread warps corresponding to the two conflicting second addresses, and when there is no conflict, proceed to step S3;
[0058] S3. Identify whether there is a conflict between the first address and the second address, and when there is no conflict, schedule the thread bundle to be scheduled in a round-robin manner, and when there is a conflict, suspend the corresponding thread bundle to be scheduled and execute the part of the thread to be scheduled that does not conflict with the first address;
[0059] It should be noted that if there are no unconflicted thread warps to be scheduled, the number of conflicts of each thread warp to be scheduled is calculated, and the thread warps to be scheduled are selected for scheduling in order of the number of conflicts from low to high, or the thread warps to be scheduled are selected for scheduling in a round-robin manner;
[0060] Specifically, a calculation cycle is set, and the conflict situation of the thread warps to be scheduled is re-judged after each calculation cycle.
[0061] This embodiment avoids warp scheduling delays caused by register access conflicts by acquiring and identifying address conflicts of warp operation registers, thereby improving the execution efficiency of GPGPU.
[0062] Further, as a refinement and extension of the specific implementation of the above embodiment, in order to fully illustrate the specific implementation process in this embodiment, another conflicting warp scheduling method is provided, which also includes the following steps:
[0063] SS1. Set a conflict controller to record the read and write register information of the thread warp instruction, mark the status of each register address being used, and adjust the scheduling order of the thread warp to be scheduled in real time according to the status of the marked register address being used; step SS1 includes the following steps:
[0064] SS11. The conflict controller marks the status of each register address being used through the conflict control table;
[0065] SS12. The conflict controller selects a thread to be scheduled that does not have a conflict for scheduling through the scheduler and according to the conflict control table;
[0066] It should be noted that the setting of the conflict controller can record the read and write register information of the thread warp instructions in real time, mark the use status of the register address, and adjust the scheduling order of the thread warp accordingly, reducing the probability of conflict;
[0067] Through the collaborative work of the conflict control table and the scheduler, it is possible to accurately select thread warps to be scheduled that do not have conflicts, further improving the performance of GPGPU.
[0068] Figure 2 As shown, the following is a GPGPU register file access method using the conflicting warp scheduling method described in the above embodiment, including the following steps:
[0069] S101. The thread warp scheduling unit selects a thread warp to be scheduled that does not conflict according to the conflicting thread warp scheduling method for scheduling;
[0070] S102. The operand access unit reads or writes back operands in parallel from the single-port multi-address register file according to the instruction of the selected scheduling warp;
[0071] It should be noted that, in combination with the conflicting warp scheduling method, efficient parallel access to the single-port multi-address register file is achieved, thereby improving the data reading and writing speed.
[0072] Further, as a refinement and extension of the specific implementation of the above embodiment, in order to fully illustrate the specific implementation process in this embodiment, another GPGPU register file memory access method is provided, including the following steps:
[0073] S101. The thread warp scheduling unit selects a thread warp to be scheduled that does not conflict according to the conflict thread warp scheduling method for scheduling; the specific steps of step S101 are as follows:
[0074] S1011. The thread warp scheduling unit fetches and decodes instructions from the thread warp to be scheduled through the instruction prefetch queue, and stores the decoded instructions into the instruction queue;
[0075] S1012. The warp scheduling unit records the read and write register information of the warp instruction through the conflict controller, and marks the status of each register address being used, and then adjusts the scheduling order of the warp to be scheduled in real time according to the usage status of the marked register address, and selects the warp to be scheduled without address conflict;
[0076] S1013. The thread warp scheduling unit stores the selected thread warps to be scheduled into the scheduling queue through the conflict controller;
[0077] It should be noted that, through the collaborative work of the instruction prefetch queue, conflict controller and scheduling queue, the thread warp instructions are accurately scheduled, register access conflicts are avoided, and the computing efficiency of GPGPU is improved;
[0078] S102. The operand access unit reads or writes back operands in parallel from the single-port multi-address register file according to the instruction of the selected scheduling warp; the specific steps of step S102 are as follows:
[0079] S1021. The operand access unit records the register address of the operand of the instruction in the thread warp, the status of whether the operand is executed, and the operand data through the operand collector;
[0080] S1022. The operand access unit arbitrates the address entries of the register file that can be accessed according to the operand collector, reads and writes operands from the single-port multi-address register file in parallel, and sets the execution completion status of the operand in the operand collector when the operand of a certain register address is successfully read or written back;
[0081] It should be noted that the setting of the operand collector can accurately record the address, status and data of the operand, and the memory access executor can access the register file in parallel based on this, thereby improving the parallelism and efficiency of data access.
[0082] Exemplarily, when the GPU task is initialized, since there is no register write-back address conflict, the read address conflict is also ignored. At this time, assuming that there is no conflicting thread warp, the scheduler in the thread warp scheduling unit will select the ready thread warp for instruction fetching according to the polling strategy. At this time, two instructions are taken out continuously. After the fetched instructions are decoded, the read and write register information of the instructions is stored in the back end of the instruction prefetch queue; the conflict control table synchronizes each control entry according to the register information contained in the front end instruction of the prefetch queue, marks the register address read and written by the thread warp, and is used for thread warp scheduling in the next cycle; at the same time, the instruction at the front end of the prefetch queue is stored in the scheduling queue, and the operand collector in the operand access unit synchronizes the operand read collection table according to the thread warp instruction (due to There is no write-back operation during initial execution, and there is no need to synchronize the operand write-back collection table). If the instruction needs to read a register, the valid field of an entry in the table is set, and the register ID of the entry is set to the number of the register to be read; the memory access unit will arbitrate the entries that can access the register block according to the numbers of the registers to be read in all operand read tables, and read the source operands from the register file in parallel; when the operand corresponding to a valid entry in an operand read collection table is successfully read, the ready field of the entry is set, and the operand is filled into the data field, indicating that the operand has been collected; when all valid entries in an operand read collection table are ready, it means that the source operands of the thread warp have been obtained;
[0083] When the GPGPU task is executed normally, in the thread warp scheduling unit, if there is a conflicting thread warp (which may include register read and write back address conflicts), the scheduler will use the polling strategy to query the thread warp without address conflicts from the ready thread warp according to the conflict control table (at this time, it is preferred to query from the ready and pre-fetched thread warp, and if it is found, the thread warp is selected, otherwise the thread warp that is ready but has not pre-fetched instructions is selected). If the thread warp without address conflicts is not found, the thread warp with fewer conflicts is selected first, or the conflicting thread warp is selected by polling; the instruction prefetch queue fetches instructions for the selected thread warp (the thread warp pre-fetched for the first time fetches two consecutive instructions, and the thread warp that has been pre-fetched fetches one instruction). After the fetched instructions are decoded, the instruction prefetch queue is at the back end; the conflict control table marks the register addresses read and written back by the thread warp according to the register information contained in the instruction at the front end of the prefetch queue; at the same time, the instruction at the front end of the prefetch queue is stored in the scheduling queue. In the operand access unit, the operand collector reads the collection table synchronously according to the thread bundle instruction. The entry setting of the read collection table is the same as the initialization stage. The operand is written back to the collection table synchronously according to the write-back request. If the request needs to write back to the register, the valid field of the entry is set, and the register ID of the entry is set to the number of the register to be written back. The memory access unit arbitrates the collection entries that can access the register file according to the two types of operand collection tables, and reads and writes operands from the register file in parallel. When the operand corresponding to a valid entry is successfully read or written back, the ready field of the entry is set, indicating that the operand has been accessed. When all valid entries in the collection table are ready, it means that the corresponding thread bundle has completed the reading and writing of the operand.
[0084] It should be understood that the order of execution of the steps in the above embodiment does not necessarily mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention.
[0085] like Figure 3 As shown, the following is an embodiment of the GPGPU register file memory access system provided by the embodiment of the present disclosure. The system and the GPGPU register file memory access method of the above-mentioned embodiments belong to the same inventive concept. For details not described in detail in the embodiment of the GPGPU register file memory access system, please refer to the embodiment of the above-mentioned GPGPU register file memory access method.
[0086] The system includes:
[0087] The thread warp scheduling unit selects a thread warp to be scheduled without conflict by using a conflicting thread warp scheduling method for scheduling;
[0088] an operand access unit, which reads or writes back operands in parallel from a single-port multi-address register file according to the instructions of the selected scheduling warp;
[0089] The execution unit, after completing the operation, sends the operand write-back request back to the thread warp scheduling unit and enters the thread warp to be scheduled.
[0090] This embodiment combines the conflicting warp scheduling method with the efficient register file memory access method, thereby achieving full utilization of GPGPU resources and improving the overall performance of the system.
[0091] Further, as a refinement and extension of the specific implementation of the above embodiment, in order to fully illustrate the specific implementation process in this embodiment, another GPGPU register file memory access system is provided, which also includes:
[0092] The thread warp scheduling unit selects a thread warp to be scheduled without conflict by using a conflicting thread warp scheduling method for scheduling; the thread warp scheduling unit includes:
[0093] The instruction pre-fetch queue fetches instructions from the thread warp to be scheduled for execution, decodes them and stores them in the instruction queue;
[0094] The conflict controller records the read and write register information of the thread bundle instruction, marks the status of each register address being used, and schedules the threads to be scheduled that do not have conflicts according to the marked register addresses;
[0095] The scheduling queue stores the selected non-conflicting thread bundles to be scheduled, waiting to be scheduled;
[0096] It should be noted that, through the collaborative work of the instruction prefetch queue, conflict controller and scheduling queue, accurate scheduling and conflict avoidance of thread warp instructions are achieved, thus improving the computing efficiency of GPGPU;
[0097] The operand access unit reads or writes back operands in parallel from the single-port multi-address register file according to the instructions of the selected scheduling thread warp; the operand access unit includes:
[0098] An operand collector records the register address of the operands of instructions in the thread warp, the status of whether the operands are executed, and the operand data;
[0099] A single-port multi-block register file, including a number of single-port memories, stores the operands corresponding to the operation objects of the thread warp and allows read or write operations to be performed in the same cycle;
[0100] Specifically, the single-port multi-block register file consists of 8 single-port memories. Although the single-port structure only allows read or write operations to be performed in the same cycle, it effectively increases the number of ports that can access memory in a single cycle compared to the 4 dual-port register files;
[0101] The memory access executor, according to the address entries of the register file that can be accessed by the operand collector, reads and writes operands from the single-port multi-address register file in parallel, and sets the execution completion status of the operand in the operand collector when the operand of a certain register address is successfully read or written back;
[0102] It should be noted that the combination of the operand collector and the single-port multi-block register file enables efficient parallel access and storage of operands, improving the speed and efficiency of data processing;
[0103] The execution unit, after completing the operation, sends the operand write-back request back to the thread warp scheduling unit and enters the thread warp to be scheduled.
[0104] In one embodiment of the present invention, based on the operand collector, a possible embodiment is given below to illustrate its specific implementation in a non-limiting manner.
[0105] The operand collectors include:
[0106] An operand read collection table, comprising N operand entries; illustratively, since an instruction includes at most 3 source operands, each operand read collection table may include 3 entries;
[0107] The operand write-back collection table includes an operand entry;
[0108] Each operand entry has:
[0109] The valid bit field is used to mark whether the operation data corresponding to each entry needs to be executed; for example, not all instructions have three source operands, and the valid bit field of the unused source operand can be set to be invalid;
[0110] The register ID field records the register number of the operand corresponding to each entry;
[0111] The data field stores the operand data corresponding to each entry;
[0112] The ready bit field marks whether each entry has completed the operation of the corresponding operation data;
[0113] It should be noted that the setting of the operand read collection table and the operand write back collection table, as well as the detailed field design of each operand entry, can accurately record and manage the status and data of the operands, providing strong support for efficient data access and processing.
[0114] This application achieves efficient use of GPGPU resources and conflict avoidance. By accurately identifying the address conflicts of thread warp operation registers and adjusting the scheduling order of thread warps in real time, thread warp scheduling delays caused by register access conflicts are avoided. At the same time, combined with the design of efficient register file access methods and operand collectors, efficient parallel access and storage of operands are achieved, improving the speed and efficiency of data processing. Overall, this application can significantly improve the execution efficiency and performance of GPGPU, and provide strong support for applications in fields such as high-performance computing and graphics processing.
[0115] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A conflicting warp scheduling method, characterized in that: The steps include: S1. Obtain the first address of the register that the scheduled warp needs to operate, and the second address of the register that the to-be-scheduled warp needs to operate; S2. Identify whether there is a conflict between the second addresses, and when there is a conflict, select a schedule from the two to-be-scheduled thread warps corresponding to the two conflicting second addresses, and when there is no conflict, proceed to step S3; S3. Identify whether there is a conflict between the first address and the second address, and when there is no conflict, schedule the thread bundles to be scheduled in a polling manner, and when there is a conflict, suspend the corresponding thread bundles to be scheduled and execute the part of the thread bundles to be scheduled that do not conflict with the first address.
2. The conflicting warp scheduling method according to claim 1, characterized in that: The following steps are also included: SS1. Set up a conflict controller to record the read and write register information of the thread warp instructions, mark the status of each register address being used, and adjust the scheduling order of the thread warps to be scheduled in real time according to the usage status of the marked register addresses.
3. The conflicting warp scheduling method according to claim 2, characterized in that: Step SS1 includes the following steps: SS11. The conflict controller marks the status of each register address being used through the conflict control table; SS12. The conflict controller selects a thread to be scheduled that does not have a conflict for scheduling through the scheduler and according to the conflict control table.
4. A GPGPU register file access method using the conflicting warp scheduling method according to any one of claims 1 to 3, characterized in that: The steps include: S101. The thread warp scheduling unit selects a thread warp to be scheduled that does not conflict according to the conflicting thread warp scheduling method for scheduling; S102. The operand access unit reads or writes back operands in parallel from the single-port multi-address register file according to the instructions of the selected scheduling warp.
5. The GPGPU register file access method according to claim 4, characterized in that: The specific steps of step S101 are as follows: S1011. The thread warp scheduling unit fetches and decodes instructions from the thread warp to be scheduled through the instruction prefetch queue, and stores the decoded instructions into the instruction queue; S1012. The warp scheduling unit records the read and write register information of the warp instruction through the conflict controller, and marks the status of each register address being used, and then adjusts the scheduling order of the warp to be scheduled in real time according to the usage status of the marked register address, and selects the warp to be scheduled without address conflict; S1013. The warp scheduling unit stores the selected warp to be scheduled in the scheduling queue through the conflict controller.
6. The GPGPU register file access method according to claim 5, characterized in that: The specific steps of step S102 are as follows: S1021. The operand access unit records the register address of the operand of the instruction in the thread warp, the status of whether the operand is executed, and the operand data through the operand collector; S1022. The operand access unit arbitrates the address entries of the register file that can be accessed according to the operand collector, reads and writes operands from the single-port multi-address register file in parallel, and sets the execution completion status of the operand in the operand collector when the operand of a certain register address is successfully read or written back.
7. A GPGPU register file memory access system using the conflicting warp scheduling method according to any one of claims 1 to 3, characterized in that: include: A thread warp scheduling unit selects a thread warp to be scheduled that does not have a conflict and schedules it by using a conflict thread warp scheduling method; an operand access unit, which reads or writes back operands in parallel from a single-port multi-address register file according to the instructions of the selected scheduling warp; The execution unit, after completing the operation, sends the operand write-back request back to the thread warp scheduling unit and enters the thread warp to be scheduled.
8. The GPGPU register file memory access system according to claim 7, characterized in that: The thread warp scheduling unit includes: The instruction pre-fetch queue fetches instructions from the thread warp to be scheduled for execution, decodes them and stores them in the instruction queue; The conflict controller records the read and write register information of the thread bundle instruction, marks the status of each register address being used, and schedules the threads to be scheduled that do not have conflicts according to the marked register addresses; The scheduling queue stores the selected non-conflicting thread bundles to be scheduled, waiting to be scheduled.
9. The GPGPU register file memory access system according to claim 8, characterized in that: The operand access unit includes: An operand collector records the register address of the operands of instructions in the thread warp, the status of whether the operands are executed, and the operand data; A single-port multi-block register file, including a number of single-port memories, stores the operands corresponding to the operation objects of the thread warp and allows read or write operations to be performed in the same cycle; The memory access executor arbitrates the address entries of the register file that can be accessed according to the operand collector, reads and writes operands from the single-port multi-address register file in parallel, and sets the execution completion status of the operand in the operand collector when the operand of a certain register address is successfully read or written back.
10. The GPGPU register file memory access system according to claim 7, characterized in that: The operand collectors include: Operand read collection table, including N operand entries; The operand write-back collection table includes an operand entry; Each operand entry has: The valid bit field marks whether the operation data corresponding to each entry needs to be executed; The register ID field records the register number of the operand corresponding to each entry; The data field stores the operand data corresponding to each entry; The ready bit field marks whether each entry has completed the operation of the corresponding operation data.
Citation Information
Patent Citations
File access method based on GPGPU multi-channel register and storage medium
CN118689538A
Key thread bundle scheduling method and device
CN119127447A