Operand collection method and device for GPGPU (General Purpose Graphics Processing Unit) and medium
By pre-allocating the operand collector in the operand collector of the GPGPU and matching in the content addressable memory unit, the problem of long and low utilization of access to register file data in the traditional operand collection method is solved, and more efficient execution efficiency and resource utilization are achieved.
Patent Information
- Application Number
- CN202510010680.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-06
AI Technical Summary
In traditional operand collection methods, accessing data in register files takes a long time, and you need to wait until the source operand becomes ready before continuing to execute, resulting in a low utilization rate of the register file section.
After the thread beam scheduler is started, the instructions are pre-allocated by the operand collector, pre-allocated information is generated, and matched in the content addressable memory unit to determine the data transmission link of the result data, reducing waiting time and unnecessary data transmission.
It improves overall execution efficiency, reduces the data's dwell time in the system, reduces the dynamic power consumption of register files, and improves resource utilization and system stability.
Smart Images

Figure CN119938141A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of GPGPU technology, and more particularly to a method, device and medium for collecting operands for GPGPU. Background Art
[0002] With the increasing number of computationally intensive tasks, general-purpose graphics processing units (GPGPUs) have become an important tool for accelerating various applications such as machine learning and data analysis. The efficient performance of GPGPUs not only benefits from their large number of computing modules, but also relies on their large register files (RFs), which are essential for achieving high parallel processing capabilities. In GPGPUs, the context of each thread is stored in the register file for fast access and execution. Threads are scheduled in warps to ensure that all threads in a warp can be scheduled to the computing unit at the same time. When all instructions are ready, the operand collector quickly loads the required source operands from the register file and sends them to the execution unit for processing.
[0003] However, due to the limitation of wiring resources, the number of register file blocks is limited, which means that it takes dozens of clock cycles to access the data in the register file. In addition, many instructions have data dependencies and must wait for their source operands to become ready before they can continue to execute, resulting in a low utilization rate of the register file blocks. At present, the streaming multiprocessor (SM) of GPGPU adopts a scheduling strategy based on thread warps, allowing each SM to execute multiple thread warps simultaneously. In order to maintain the context of all thread warps and support fast context switching, the SM is equipped with a large register file, but its number of modules and data bit width are limited by hardware resources, and the energy consumption of the register file accounts for a large part of the total energy consumption of GPGPU. Therefore, simply increasing the read and write bit width and the number of blocks to reduce the source operand loading delay may significantly increase energy consumption and affect the overall performance of GPGPU. Therefore, in the traditional operand collection method, it takes a long time to access the data in the register file, and it is necessary to wait for the source operand to become ready before continuing to execute, resulting in a low utilization rate of the register file blocks. Summary of the invention
[0004] One or more embodiments of the present specification provide an operand collection method, device and medium for GPGPU, which are used to solve the following technical problems: In the traditional operand collection method, it takes a long time to access the data in the register file, and it is necessary to wait for the source operand to become ready before continuing execution, resulting in low utilization of the register file block.
[0005] One or more embodiments of this specification adopt the following technical solutions:
[0006] One or more embodiments of the present specification provide an operand collection method for GPGPU, the method comprising: after a warp scheduler is started, pre-allocating an operand collector for an instruction, generating pre-allocation information, wherein the pre-allocation information comprises a warp identifier, a register identifier of an unready source operand, an operand collector identifier, and a register refresh signal; based on the pre-allocation information, generating entry tag information corresponding to the instruction, and after the instruction is executed, generating execution information, wherein the execution information comprises result data, an execution warp identifier, and an execution register identifier; matching is performed in a content addressable memory unit according to the execution warp identifier, the execution register identifier, and the entry tag information, so as to determine a data sending link for the result data through a matching result, wherein the data sending link comprises any one of an operand collector or a register file.
[0007] Furthermore, after the thread warp scheduler is started, an operand collector is pre-allocated to the instruction to generate pre-allocation information, which specifically includes: obtaining a program counter value of the instruction and sending it to a static random access memory unit in a compiler support module; using the program counter value as an entry index to determine a corresponding search entry in the static random access memory unit, wherein the search entry includes a plurality of register refresh bits, each of which corresponds to a source operand and is used to indicate whether the source operand is allowed to skip a register write process; monitoring the execution process of the instruction through the thread warp scheduler to determine location information of a not-ready source operand; and determining a register refresh signal corresponding to the not-ready source operand through a multiplexer based on the location information of the not-ready source operand and the plurality of register refresh bits.
[0008] Furthermore, before pre-allocating operand collectors for instructions, the method further includes: screening instructions among a plurality of candidate instructions to determine a designated instruction for pre-allocation, wherein the designated instruction is an instruction without long-latency memory dependency.
[0009] Furthermore, based on the pre-allocated information, entry tag information corresponding to the instruction is generated, specifically including: obtaining the thread warp identifier and the register identifier of the not-ready source operand in the pre-allocated information, concatenating the thread warp identifier and the register identifier of the not-ready source operand to determine the entry tag; determining the entry index with the operand collector identifier; and generating the entry tag information corresponding to the instruction through the entry tag, the entry index and the register refresh signal.
[0010] Furthermore, matching is performed in a content addressable memory unit according to the execution thread bundle identifier, the execution register identifier and the entry tag information, specifically including: obtaining an entry tag, an entry index and a register refresh signal in the entry tag information; performing tag matching according to the execution thread bundle identifier, the execution register identifier and the entry tag, and when a bitwise AND result is true, determining that a matching tag exists; when a bitwise AND result is false, determining that a matching tag does not exist.
[0011] Furthermore, the data sending link of the result data is determined through the matching result, specifically including: when the matching result is that there is a matching tag, determining that the data sending link of the result data is an operand collector; when the matching result is that there is no matching tag, performing an AND operation check with the matching result according to the register refresh signal in the entry tag information, and when the AND operation result is false, determining that the data sending link of the result data is a register file.
[0012] Furthermore, before using the program counter value as an entry index to determine the corresponding search entry in the static random access memory unit, the method also includes: obtaining the maximum number of source operands contained in the instruction, and determining the number of entries in the static random access memory unit based on the maximum number of source operands, wherein each entry corresponds to a source operand of the instruction; obtaining the corresponding inter-instruction dependency of the compiler in the instruction generation stage, and identifying the inter-instruction dependency to determine the register refresh bit corresponding to each source operand.
[0013] Further, performing instruction screening among multiple candidate instructions to determine the designated instruction to be pre-assigned specifically includes: obtaining a pre-built scoreboard, wherein the scoreboard is used for dependencies between instructions and the execution status and waiting time of each instruction; monitoring the instruction dependency of each candidate instruction through the pre-built scoreboard, determining the candidate instructions with dependencies, and determining the dependency waiting time; determining the designated instruction according to the dependency waiting time and a preset long delay time threshold. One or more embodiments of this specification provide an operand collection device for GPGPU, including:
[0014] at least one processor; and,
[0015] a memory communicatively connected to the at least one processor; wherein,
[0016] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the above method.
[0017] One or more embodiments of the present specification provide a non-volatile computer storage medium storing computer executable instructions, wherein the computer executable instructions are configured to execute the above method.
[0018] At least one of the above technical solutions adopted in the embodiments of this specification can achieve the following beneficial effects: through the above technical solution, by pre-allocating the operand collector after the thread warp scheduler is started, the waiting time caused by insufficient resources during the instruction execution process is reduced. When the instruction is executed, execution information including result data, execution thread warp identifier and execution register identifier is generated, which can be quickly used for subsequent data processing or instruction execution, reducing the retention time of data in the system and improving the overall execution efficiency; by matching through the content addressable memory unit (CAM), the data sending link of the result data can be accurately determined, avoiding unnecessary data transmission and register write operations, especially in the read-after-write dependency, when the data is only used once and overwritten by subsequent data, the dynamic power consumption of the register file can be significantly reduced; pre-allocating the operand collector also avoids the overhead of frequently applying for and releasing resources during the instruction execution process, improving the resource utilization and system stability; without changing the register structure, the number and complexity of the operand collector will not be increased, and the performance optimization is achieved by introducing additional data links for the operand collector and register bandwidth that are not fully utilized. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art description. Obviously, the drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative labor. In the drawings:
[0020] Figure 1 A flowchart of an operand collection method for GPGPU provided in an embodiment of this specification;
[0021] Figure 2 A schematic diagram of the architecture of an operand collection method provided in an embodiment of this specification;
[0022] Figure 3 A schematic diagram of the structure of a static random access memory unit provided in an embodiment of this specification;
[0023] Figure 4 A schematic diagram of the structure of a content addressable memory unit provided in an embodiment of this specification;
[0024] Figure 5A schematic diagram of the structure of an operand collection device for GPGPU provided in an embodiment of this specification. DETAILED DESCRIPTION
[0025] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments of this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this specification.
[0026] With the increase of computationally intensive tasks, general-purpose graphics processing units (GPGPUs) are widely used to accelerate various applications, such as machine learning and data analysis. In addition to its large number of computing modules, its high parallel capability also relies on its large register file (RF). The context of each thread is stored in the register file. In GPGPU, threads are scheduled in units of thread bundles. All threads in a thread bundle will be scheduled to the computing unit at the same time. When all instructions are ready, all source operands will be loaded from the register file by the operand collector and sent to the execution unit. Due to limited wiring resources, the number of register file blocks is also limited. It often takes dozens of clock cycles to access data in the register file. In addition, the utilization of register file blocks is usually very limited because many instructions have data dependencies and need to wait for their source operands to become ready.
[0027] The embodiments of this specification provide an operand collection method for GPGPU. It should be noted that the execution subject in the embodiments of this specification can be a server or any device with data processing capabilities. Figure 1 A flowchart of a method for collecting operands for GPGPU provided in an embodiment of this specification is shown in FIG. Figure 1 As shown, it mainly includes the following steps:
[0028] Step S101 : after the warp scheduler is started, operand collectors are pre-allocated to instructions to generate pre-allocation information.
[0029] The pre-allocation information includes a thread warp identifier, a register identifier of an unready source operand, an operand collector identifier, and a register refresh signal;
[0030] In one embodiment of the present specification, an early operand collector allocation strategy is proposed, under which all source operands of the instruction are loaded into the operand collector even if a source operand is not yet ready. At the same time, in order to reduce the impact on the overall performance, this strategy will only take effect on instructions that do not have long-latency memory dependencies. In traditional GPGPUs, the warp scheduler assigns an available operand collector to a ready instruction and enables register file access operations. In order to check the ready state, the warp scheduler consults a scoreboard, which is essentially a lookup table that tracks the status of a single register, including whether the register is ready or whether the production instruction has completed execution. In an embodiment of the present specification, the warp scheduler does not wait for all source operands of the instruction to become ready, but actively allocates operand collectors. Figure 2 A schematic diagram of the architecture of an operand collection method provided in an embodiment of this specification, such as Figure 2 As shown in ①, even if the source operand is not ready, the pre-allocation operation is still performed. By reading the source operand immediately after the operand collector allocates it, the bandwidth of the register file can be fully utilized, thereby reducing the extraction latency of the ready source operand.
[0031] After the thread warp scheduler is started, operand collectors are pre-allocated to the instructions to generate pre-allocation information, which specifically includes: obtaining the program counter value of the instruction and sending it to the static random access memory unit in the compiler support module; using the program counter value as an entry index to determine the corresponding search entry in the static random access memory unit, wherein the search entry includes a plurality of register refresh bits, each of which corresponds to a source operand and is used to indicate whether the source operand is allowed to skip the register write process; monitoring the execution process of the instruction through the thread warp scheduler to determine the position information of the unready source operand; and determining the register refresh signal corresponding to the unready source operand through a multiplexer according to the position information of the unready source operand and the plurality of register refresh bits.
[0032] In one embodiment of the present specification, the warp scheduler starts and prepares to schedule the ready instructions. The warp scheduler consults the scoreboard to track the status of the registers, including whether the registers are ready or whether the production instruction has completed execution. The warp scheduler does not wait for all the source operands of the instruction to become ready, but actively allocates the operand collector. Even if the source operand is not ready, the warp scheduler sends the corresponding warp ID, the register ID of the unready source operand, the operand collector ID and the register refresh (RR) signal to the data control module. When the warp scheduler actively allocates the operand collector, the program counter (PC) of the instruction is sent to the static random access memory (SRAM) as an entry index. The three RR bits contained in the indexed entry are sent to the cache in the multiplexer. The warp scheduler provides the location of the unready source operand to the multiplexer to control the output of the multiplexer. Finally, the RR state of the unready source operand is sent to the warp scheduler to participate in the subsequent logic.
[0033] When the warp scheduler decides to actively assign an operand collector to an instruction, it first obtains the program counter (PC) value of the instruction. This PC value represents the unique location of the instruction in the program and is the key to subsequent search and indexing. The warp scheduler then sends this PC value to the SRAM static random access memory unit in the compiler support module. After receiving the PC value, the SRAM unit uses it as an entry index to find the corresponding entry.
[0034] Before determining the corresponding search entry in the static random access memory unit using the program counter value as an entry index, the method also includes: obtaining a maximum number of source operands contained in an instruction, and determining the number of entries in the static random access memory unit using the maximum number of source operands, wherein each entry corresponds to a source operand of an instruction; and obtaining the dependency between instructions corresponding to the compiler in the instruction generation stage, and identifying the dependency between instructions to determine a register refresh bit corresponding to each source operand.
[0035] In one embodiment of this specification, Figure 3 This is a schematic diagram of the structure of a static random access memory unit provided in an embodiment of the present specification. The compiler support module is mainly composed of a small SRAM unit. The entry size is determined by the maximum number of source operands of an instruction. Since the number of source operands in a conventional instruction will not exceed three, each SRAM entry contains three bits, such as Figure 3As shown in the figure, each entry contains three bits, which are called RR bits (register refresh bits). Each RR indicates whether the corresponding source operation can skip the register write process. This flag is generated by the compiler. It should be noted that due to the existence of the data transmission link, data that is not related to maintaining the complete system state does not need to be written into the register file, such as Figure 2 As shown in ③. For example, when the value of the source operand register that the instruction is waiting for is only used once by the instruction and then overwritten and refreshed by the updated instruction data, there is no need to write the result value back to the register file. This read-after-write dependency is partially recognized by the compiler and a flag is sent, which is processed by the compiler support module to avoid unnecessary register write operations. The register file utilization is thus improved, while the dynamic power consumption of the register file is reduced.
[0036] After the SRAM unit reads the corresponding entry, it sends the three RR bits to the multiplexer for caching. The multiplexer is an important data selection device that can select and output specific data according to the control signal. The warp scheduler provides the multiplexer with the location information of the unready source operands. This location information is one or more bit masks that indicate which source operands are not ready. The multiplexer selects the corresponding RR state based on this location information and the cached RR bits and sends it to the warp scheduler. After the warp scheduler receives the RR state of the unready source operand, it determines whether to allow the register write process to be skipped based on this state and the location of the unready source operand. If the RR state indicates that a source operand can skip the register write process, the warp scheduler will avoid unnecessary register read and write operations on the source operand in subsequent logical processing. This can reduce the number of register file accesses, improve bandwidth utilization, and reduce power consumption.
[0037] In traditional GPGPU, the warp scheduler usually waits for all source operands of the instruction to become ready before allocating the operand collector. However, this approach may waste the bandwidth of the register file because some source operands may not be ready yet, but the register file is idle. In order to make full use of the bandwidth of the register file, it is proposed that the warp scheduler can actively allocate the operand collector even if the source operand is not ready. This allows the register file to continue processing other ready source operands while waiting for the source operand to be ready, thereby improving bandwidth utilization. By reading the source operand immediately after the operand collector is allocated, the extraction latency of the ready source operand can be further reduced. The operand collector is ready to receive data, and once the source operand is ready, it can be sent to the operand collector immediately without waiting for further scheduling by the warp scheduler. For unready source operands, they must be the target registers of a previous instruction. The values of these registers are either calculated by the functional unit or loaded from the memory. By designing additional data transmission data streams, these data values can be directly forwarded to the corresponding operand collectors without writing / reading registers, which not only saves read and write cycles, reduces source operand extraction latency, but also saves register file resources. Because for data that is only used once and then updated, there is no need to write it back to the register file.
[0038] Before pre-allocating operand collectors for instructions, the method also includes: screening instructions among multiple candidate instructions to determine the designated instructions for pre-allocation, wherein the designated instructions are instructions without long-delay memory dependencies. Screening instructions among multiple candidate instructions to determine the designated instructions for pre-allocation specifically includes: obtaining a pre-built scoreboard, wherein the scoreboard is used for dependencies between instructions and the execution status and waiting time of each instruction; monitoring the instruction dependencies of each candidate instruction through the pre-built scoreboard, determining the candidate instructions with dependencies, and determining the dependency waiting time; determining the designated instruction based on the dependency waiting time and a preset long delay time threshold.
[0039] The early operand collector allocation strategy has drawbacks. When all operand collectors are occupied by this strategy, instructions for which all source operands are ready cannot be processed. When the waiting source operands require data loaded from the memory hierarchy, longer read and write delays exacerbate this problem.
[0040] In one embodiment of the present specification, a scoreboard data structure is established to record the dependencies between instructions and the waiting time and state of each instruction, and the data in the scoreboard is initialized to ensure that the dependencies and waiting times of all instructions are in the initial state. The instruction stream is monitored, and the dependencies and operand requirements of each instruction are identified in the instruction stream. The dependency and operand requirement information of the instructions are recorded in the scoreboard. The scoreboard is a data structure used to record the dependencies between instructions in the processor and the execution status and waiting time of each instruction, and is used to identify which instructions have data dependencies and the nature of these dependencies (such as long delay or short delay). The scoreboard usually contains the following information: instruction ID, dependency (recording the data dependency between each instruction and other instructions), waiting time (recording the time each instruction waits for the operand or resource it depends on to become available), etc.
[0041] Long-latency data dependency means that the instruction waits for the data it depends on to become available for a long time. In order to identify long-latency data dependency, the processor will periodically check the instruction dependency in the scoreboard. First, traverse each instruction in the scoreboard to check the dependency and waiting time; for instructions with dependencies, calculate their waiting time and compare it with the preset long-latency time threshold. If the waiting time of the instruction exceeds the long-latency time threshold, it is determined to be a long-latency data dependency. It should be noted that the long-latency time threshold here can be set according to the performance goals, application scenarios and hardware characteristics of the processor. It can be a fixed threshold set based on experience or test results, or the long-latency time threshold can be dynamically adjusted according to real-time performance data and hardware status. For instructions that do not have long-latency memory dependencies, the advance operand collector allocation strategy is directly applied to allocate operand collectors to these instructions so that operands can be collected as early as possible. When all candidate instructions depend on long-latency memory access instructions, the conditional advance operand collector allocation strategy is adopted. According to the waiting time of the instructions, the instructions with the longest waiting time are processed first, and the operand collector is allocated to the instructions with the longest waiting time to reduce the blocking cycle and improve the utilization efficiency of the register file.
[0042] By introducing a scoreboard mechanism to monitor the latency of data dependencies, and based on this, decide whether to apply the early operand collector allocation strategy to instructions. The scoreboard is used to record various dependencies, especially the latency of memory access, and is used to identify which instructions have long-latency dependencies. Usually, instructions need to collect the operands they need before execution. Early allocation can reduce the time instructions wait for operands, thereby improving execution efficiency. Early allocation is only applied when there is no long-latency memory dependency. When all candidate instructions depend on long-latency memory access instructions, early allocation is conditionally allowed based on the waiting time of the instructions, reducing processor blocking caused by long-latency dependencies and improving the utilization efficiency of register files.
[0043] Step S102, based on the pre-allocated information, generating entry tag information corresponding to the instruction, and generating execution information after the instruction is executed.
[0044] The execution information includes result data, execution warp identifier and execution register identifier;
[0045] Based on the pre-allocated information, the entry tag information corresponding to the instruction is generated, specifically including: obtaining the thread warp identifier and the register identifier of the not-ready source operand in the pre-allocated information, concatenating the thread warp identifier and the register identifier of the not-ready source operand to determine the entry tag; determining the entry index with the operand collector identifier; and generating the entry tag information corresponding to the instruction through the entry tag, the entry index and the register refresh signal.
[0046] In one embodiment of this specification, Figure 4 A structural diagram of a content addressable memory unit provided in an embodiment of this specification is shown in FIG. Figure 4 As shown, a small content addressable memory (CAM) unit constitutes the data control module, and each entry stores a tag as the unique identifier of the source operand (composed of the thread warp ID and the register ID), the operand collector ID as the entry index, and the register refresh flag RR (Register Refresh). The number of entries in the memory is equal to the number of operand collector IDs. When the thread warp scheduler actively allocates an operand collector, it will simultaneously send the corresponding thread warp ID, the register ID of the unready source operand, the operand collector ID, and the RR signal to the control module. The memory is indexed by the operand collector ID to determine the entry index, and the thread warp ID and the register ID are concatenated to form the tag of the entry to determine the entry tag.
[0047] Instructions are executed in the execution module. When the thread warp scheduler assigns an instruction to an execution module, the execution module starts to execute the instruction. The execution module contains multiple computing units, which work in parallel to accelerate the execution of instructions. During the execution of instructions, the register file or operand collector is accessed to obtain the required source operands. The computing unit calculates according to the instruction's opcode and source operands to generate result data.
[0048] When the instruction is executed, the execution module sends the result data and the corresponding thread bundle and register ID back to the data control module, that is, generates execution information, which includes the result data, the execution thread bundle identifier and the execution register identifier. When the instruction is executed, the execution module sends the execution information back to the data control module. The data control module is a bridge connecting the execution module and the register file / operand collector, responsible for receiving the result data sent by the execution module and sending it to the correct location as needed.
[0049] Step S103 : matching is performed in the content addressable memory unit according to the execution warp identifier, the execution register identifier and the entry tag information, so as to determine the data transmission link of the result data according to the matching result.
[0050] The data transmission link includes any one of an operand collector and a register file.
[0051] According to the execution thread bundle identifier, the execution register identifier and the entry tag information, matching is performed in the content addressable memory unit, specifically including: obtaining the entry tag, entry index and register refresh signal in the entry tag information; performing tag matching according to the execution thread bundle identifier, the execution register identifier and the entry tag, and when the bitwise AND result is true, determining that a matching tag exists; when the bitwise AND result is false, determining that no matching tag exists.
[0052] In one embodiment of the present specification, in the data control module, each entry stores a tag, which is composed of a warp ID and a register ID, and is used as a unique identifier of the source operand for matching in subsequent data processing. When the warp scheduler actively allocates an operand collector, it will simultaneously send the corresponding warp ID, the register ID of the unready source operand, the operand collector ID, and the RR (register refresh) signal to the data control module, and this information is stored in the content addressable memory CAM unit and is associated with the tag.
[0053] When the instruction is executed, the data control module needs to search for a tag that matches the result data from the CAM entry, which is achieved by performing a bitwise AND operation on the stored tag and a new tag formed by concatenating the thread bundle ID and the register ID. That is, after receiving the data, the data control module will immediately search for the tag associated with the result data in its content addressable memory (CAM). These tags usually correspond one-to-one to the entries in the operand collector and are used to identify which operand collectors are waiting for specific data. During the search process, the data control module will perform a bitwise AND operation to compare the tags. Only when all bits match, it is considered that a matching tag is found. It should be noted that bitwise AND is a basic bit operation. It performs an AND operation on each bit of two numbers. Only when both corresponding bits are 1, the bit of the result is 1, otherwise it is 0. If the result of the bitwise AND operation is all 1 (that is, all bits are 1), then it can be considered that the tag match is successful and there is a matching tag, that is, the operand collector is waiting for the data. When the bitwise AND result is false, it is determined that there is no matching tag.
[0054] The data sending link of the result data is determined through the matching result, specifically including: when the matching result is that there is a matching tag, determining that the data sending link of the result data is an operand collector; when the matching result is that there is no matching tag, performing an AND operation check with the matching result according to the register refresh signal in the entry tag information, and when the AND operation result is false, determining that the data sending link of the result data is a register file.
[0055] In one embodiment of the present specification, if the CAM search result shows that there is a matching tag, that is, the bitwise AND result of the tag comparison is true, this indicates that an operand collector is waiting for the data. At this time, the data control module will use the data transmission link to directly send the result data to the corresponding operand collector, avoiding the unnecessary overhead of writing data to the register file. If there is no matching tag, or the RR bit (register refresh bit) indicates that the data must be written to the register file, the data control module will write the result data to the register file. The write position of the register file is determined by the thread warp ID and the register ID. In addition, when deciding whether the data is written to the register file, the data control module will also check the RR bit of the entry. The RR bit is used to indicate whether a source operand can skip the register write process. By default, the result data needs to perform a write operation. However, if the RR bit is compared with the tag and the result is true after the AND operation, it means that the register is no longer referenced by subsequent instructions, then the data will not be written to the register file, thereby improving the utilization of the register file and reducing dynamic power consumption.
[0056] Through the above technical solution, by pre-allocating the operand collector after the thread warp scheduler is started, the waiting time caused by insufficient resources during the instruction execution process is reduced. When the instruction is executed, execution information including result data, execution thread warp identifier and execution register identifier is generated, which can be quickly used for subsequent data processing or instruction execution, reducing the retention time of data in the system and improving the overall execution efficiency; by matching through the content addressable memory unit (CAM), the data sending link of the result data can be accurately determined, avoiding unnecessary data transmission and register write operations, especially in the read-after-write dependency, when the data is only used once and overwritten by subsequent data, the dynamic power consumption of the register file can be significantly reduced; pre-allocating the operand collector also avoids the overhead of frequently applying for and releasing resources during the instruction execution process, improving the resource utilization and system stability; without changing the register structure, the number and complexity of the operand collector will not be increased, and the performance optimization is achieved by introducing additional data links for the operand collector and register bandwidth that are not fully utilized.
[0057] The embodiment of this specification also provides an operand collection device for GPGPU, such as Figure 5 As shown, the device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above method.
[0058] The embodiments of the present specification also provide a non-volatile computer storage medium storing computer executable instructions, wherein the computer executable instructions are configured to execute the above method.
[0059] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device, equipment, and non-volatile computer storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0060] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0061] The devices and media provided in the embodiments of this specification correspond one-to-one to the methods. Therefore, the devices and media also have similar beneficial technical effects as the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be repeated here.
[0062] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0063] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0064] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0065] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0066] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0067] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0068] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0069] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0070] The above description is only one or more embodiments of this specification and is not intended to limit this specification. For those skilled in the art, one or more embodiments of this specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of one or more embodiments of this specification shall be included in the scope of the claims of this specification.
Claims
1. An operand collection method for GPGPU, characterized in that: The method comprises: After the thread warp scheduler is started, the operand collector is pre-allocated to the instruction to generate pre-allocation information, wherein the pre-allocation information includes the thread warp identifier, the register identifier of the unready source operand, the operand collector identifier and the register refresh signal; Based on the pre-allocated information, generating entry tag information corresponding to the instruction, and when the instruction is executed, generating execution information, wherein the execution information includes result data, an execution warp identifier, and an execution register identifier; According to the execution warp identifier, the execution register identifier and the entry tag information, a match is performed in a content addressable memory unit to determine a data sending link of the result data through a matching result, wherein the data sending link includes any one of an operand collector or a register file.
2. The method for collecting operands for GPGPU according to claim 1, characterized in that: After the thread warp scheduler is started, the operand collector is pre-allocated for the instruction and pre-allocation information is generated, including: Obtaining a program counter value of the instruction and sending it to a static random access memory unit in a compiler support module; Using the program counter value as an entry index, determining a corresponding search entry in the static random access memory unit, wherein the search entry includes a plurality of register refresh bits, each of the register refresh bits corresponds to a source operand, and is used to indicate whether the source operand is allowed to skip a register write process; The execution process of the instruction is monitored by the thread warp scheduler to determine the location information of the unready source operand; A register refresh signal corresponding to the not-ready source operand is determined through a multiplexer according to the position information of the not-ready source operand and the plurality of register refresh bits.
3. The method for collecting operands for GPGPU according to claim 2, characterized in that: Before pre-allocating an operand collector for an instruction, the method further includes: Instruction screening is performed among multiple candidate instructions to determine a designated instruction for pre-allocation, wherein the designated instruction is an instruction without long-latency memory dependency.
4. The method for collecting operands for GPGPU according to claim 1, characterized in that: Based on the pre-allocated information, generating entry tag information corresponding to the instruction specifically includes: Obtaining the thread warp identifier and the register identifier of the unready source operand in the pre-allocated information, concatenating the thread warp identifier and the register identifier of the unready source operand, and determining an entry tag; Determine an entry index using the operand collector identifier; Entry tag information corresponding to the instruction is generated through the entry tag, the entry index and the register refresh signal.
5. The method for collecting operands for GPGPU according to claim 1, characterized in that: According to the execution warp identifier, the execution register identifier and the entry tag information, matching is performed in a content addressable memory unit, specifically comprising: Acquire an entry tag, an entry index and a register refresh signal in the entry tag information; Perform tag matching according to the execution warp identifier, the execution register identifier and the entry tag, and when a bitwise AND result is true, determine that a matching tag exists; When the bitwise AND result is false, it is determined that there is no matching tag.
6. The method for collecting operands for GPGPU according to claim 5, characterized in that: Determining a data transmission link of the result data through the matching result specifically includes: When the matching result is that a matching tag exists, determining that a data transmission link of the result data is an operand collector; When the matching result is that there is no matching tag, an AND operation check is performed with the matching result according to the register refresh signal in the entry tag information, and when the AND operation result is false, it is determined that the data transmission link of the result data is a register file.
7. The method for collecting operands for GPGPU according to claim 2, characterized in that: Before determining a corresponding search entry in the static random access memory unit using the program counter value as an entry index, the method further includes: Obtaining a maximum number of source operands included in an instruction, and determining a number of entries in the static random access memory unit based on the maximum number of source operands, wherein each entry corresponds to a source operand of a corresponding instruction; Obtain the inter-instruction dependency corresponding to the compiler in the instruction generation stage, identify the inter-instruction dependency, and determine the register refresh bit corresponding to each source operand.
8. The method for collecting operands for GPGPU according to claim 3, characterized in that: Instructions are screened among multiple candidate instructions to determine the specified instructions for pre-allocation, including: Obtaining a pre-built scoreboard, wherein the scoreboard is used for dependencies between instructions and the execution status and latency of each instruction; Through the pre-built scoreboard, the instruction dependency of each candidate instruction is monitored, the candidate instructions with dependencies are identified, and the dependency waiting time is determined; The designated instruction is determined according to the dependent waiting time and a preset long delay time threshold.
9. An operand collection device for GPGPU, characterized in that: The device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the method according to any one of claims 1 to 8.
10. A non-volatile computer storage medium storing computer executable instructions, characterized in that: The computer executable instructions are configured to execute the method according to any one of claims 1 to 8.
Citation Information
Cited By
Memory configuration method and device, electronic equipment, storage medium and program product
CN121541935A