A system for reading and writing inter-thread operands
Through the unified management of inter-thread data by registers, the synchronization delay problem of GPGPU in data sharing scenarios is solved, efficient cross-thread data sharing is achieved, and the parallel performance of GPGPU is significantly improved.
Patent Information
- Application Number
- CN202510246184.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-03-04
AI Technical Summary
The read and write delay and cache usage problems caused by synchronous operations in data sharing scenarios affect parallel performance.
By using registers as the only medium for data sharing, thread remapping and cross-thread marking methods are used to uniformly manage register data and status between threads, realizing cross-thread data reading and writing, and avoiding synchronous operations.
It reduces synchronization overhead, reduces read and write delay, improves cache resource utilization, ensures data security and consistency, and significantly improves the parallel performance of GPGPU in multi-threaded parallel computing scenarios.
Smart Images

Figure CN119718426B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cross-thread data sharing of GPGPU, in particular to a system for cross-thread reading and writing of inter-thread operands. Background Art
[0002] The General-Purpose Graphics Processing Unit (GPGPU) uses a single instruction multiple thread (SIMT) architecture, which allows multiple threads to execute the same instruction in parallel. The SIMT design enables GPGPU to efficiently handle large-scale parallel computing tasks while maintaining programming flexibility and scalability. GPGPU organizes multiple threads (usually 32) into a warp. The threads in the warp execute the same instruction synchronously, but the data between threads is independent and cannot be directly shared between threads.
[0003] GPGPU provides a fast on-chip storage resource called shared memory for data sharing between warps. The access speed of shared memory is much higher than that of global memory, so in scenarios where data needs to be exchanged frequently, using shared memory can significantly improve performance. Threads in a warp can write data to shared memory, and threads in other warps can read the data from shared memory to achieve data sharing. However, when using shared memory to share data, synchronization operations must be performed between warps to ensure data consistency and correctness. However, during the synchronization process, all warps need to wait for the slowest warp to complete the operation, which will undoubtedly bring additional overhead. Especially when there are a large number of warps, the synchronization overhead will be more significant, thereby reducing the efficiency of program execution.
[0004] In recent years, artificial intelligence technology represented by neural networks has developed rapidly. Various neural network models use matrix operations as basic operators, and matrix operations usually rely on shared memory to improve performance. Although shared memory provides a convenient way to exchange data between thread bundles, in some GPGPU architectures, the initial data loading path of shared memory is inconvenient. When programmers use shared memory, they need to use register Rx as a medium to first read data from global memory. The data will be written to the medium register after passing through multiple storage levels such as multi-level cache; finally, the data is read from the register and written to the shared memory. In addition, since shared memory relies on registers as a medium, and the cache is mostly inclusive, some data may be retained in the L1 cache, resulting in a waste of valuable on-chip cache resources. Summary of the invention
[0005] In view of the needs and shortcomings of current technological development, the present invention provides a system for cross-thread reading and writing of inter-thread operands, which solves the problems of read and write delays and cache occupancy caused by synchronous operations in GPGPU in data sharing scenarios and improves parallel performance.
[0006] The system for reading and writing inter-thread operands between threads of the present invention adopts the following technical solutions to solve the above technical problems:
[0007] A system for reading and writing inter-thread operands, the implementation architecture of which includes:
[0008] The decoding unit is composed of a custom instruction decoder and a thread remapping module, wherein: the decoder decodes the custom instruction and extracts the control information; the thread remapping module calculates the real thread index according to the extracted control information;
[0009] The emission unit includes an instruction queue, a scoreboard, and an operand collector, wherein: the instruction queue caches decoded instructions; the scoreboard creates a score table for each thread bundle, controls the instruction execution order, and avoids data correlation risks by querying the register write-back status flag to achieve data synchronization control; the operand collector manages the registers of all threads in a unified manner, creates a collection table for each emission channel, initializes the collection table according to the instruction control information, calculates the offset address, reads the register from the unified storage space, and completes cross-thread data reading; after completing the data reading, the emission unit sends the instruction and data to the execution unit, and the result of the execution unit is then passed to the write-back unit for write-back operation;
[0010] The write-back unit extracts the destination operand from the instruction execution result, calculates the destination register offset address according to the destination thread index and the destination register label, writes the data into the storage space, and notifies the scoreboard to modify the write-back status of the corresponding register, completing cross-thread data writing and forming a complete cross-thread data read and write path.
[0011] Optionally, the RISC-V instruction set is extended to define inter-warp operand memory access instructions, inter-thread operand memory access instructions of different warps, and inter-thread operand memory access instructions in the same warp;
[0012] The decoder decodes the custom instructions after the RISC-V instruction set is extended and extracts control information, including thread index and thread index offset; the thread remapping module adds the thread index and thread index offset of the instruction to calculate the actual thread index.
[0013] Optionally, when decoding a custom instruction, the decoding unit involved needs to verify the legality of the instruction: if it is found that the instruction contains illegal control information or does not meet the format requirements of the custom instruction set, the decoding unit refuses to decode the instruction and feeds back the error information to the system;
[0014] When calculating the actual thread index, the thread remapping module will consider the boundary conditions of the thread warp: if the calculated thread index exceeds the valid range of the thread warp, the thread remapping module will correct the thread index according to the preset rules to ensure the validity and accuracy of the thread index.
[0015] Optionally, the instruction queues involved cache decoded instructions corresponding to the same warp.
[0016] Further optionally, the scoreboard involved specifically controls the execution order of the instructions corresponding to each thread warp through the following operations:
[0017] First, the score table allocates a 1-bit status bit to each register of the thread to record the write-back status of the register;
[0018] Then, when an instruction needs to read a register across threads, the scoreboard queries the status flag corresponding to the register to be accessed based on the decoded thread index and register number to determine whether the data has been written back;
[0019] Finally, the scoreboard traverses the instructions in the instruction queue and gives priority to instructions that have completed data writeback for operand collection, avoiding correlation hazards and achieving data synchronization control.
[0020] Preferably, the scoreboard involved will evaluate the priority of the instructions when controlling the execution order of the instructions: for instructions with high real-time requirements or critical to the entire computing process, the scoreboard will preferentially mark them as executable on the basis of satisfying the data write-back status judgment, even if there are instructions in the queue that have completed data write-back at this time;
[0021] At the same time, the scoreboard will record the waiting time of the instruction in the queue. If the waiting time of an instruction exceeds the preset threshold and the required data has been written back, the scoreboard will adopt the preset processing mechanism to adjust the execution order of the instructions or issue a warning message to avoid long-term blocking of instructions and affect system performance.
[0022] Optionally, the operand collector involved performs the following operations to complete the inter-thread operand reading:
[0023] First, each collection table contains n collection entries, where n is determined by the maximum number of registers that may be used by the instruction. Each collection entry includes the following fields: a valid bit field, indicating whether the operand corresponding to the entry needs to be collected; a source register number, indicating the register where the operand is located; an offset address field, indicating the address of the register in the storage space; a data field, used to temporarily store the read operand; and a ready bit field, indicating whether the operand corresponding to the entry has been collected.
[0024] Then, the operand collector initializes the collection table according to the control information of the instruction, where the offset address is the absolute address of the register to be read in the storage space, which needs to be calculated based on the decoded thread index and register number. The calculation system is shown in formula (1):
[0025] addr offset ={sRsx[bit:0],wIndex}Formula (1),
[0026] In the formula, addr offset Indicates the offset address; sRsx indicates the information related to the source register, sRsx[bit:0] indicates taking the part of the source register sRsx from the bit to the 0th bit, which may contain part of the register address or thread-related index information; wIndex indicates the index of the thread bundle, which is used to identify different thread bundles and locate the position of the thread bundle in the entire computing resource;
[0027] Finally, the operand collector reads the register from the unified storage space according to the offset address to complete the cross-thread data reading.
[0028] Preferably, when the operand collector involved initializes the collection table, the setting of the valid bit field follows the following rules:
[0029] a) When the instruction explicitly requires the collection of a certain operand, the valid bit field of the corresponding collection entry is set to a valid state;
[0030] b) When the instruction clearly does not need to collect a certain operand, the corresponding valid bit field is set to an invalid state;
[0031] At the same time, the operand collector initializes the ready bit field in the collection table to the not-ready state to ensure that the operand will not be used incorrectly when it has not completed collection;
[0032] After initializing the collection table according to the instruction control information, the operand collector checks the integrity of the collection table. If it is found that key information is missing or wrong in the collection table, an error signal will be issued and the current operation will be stopped, waiting for further processing.
[0033] Preferably, the operand collector in the emission unit involved adopts a cache mechanism when reading register data from the unified storage space;
[0034] When the operand collector reads the register data for the first time according to the offset address, it stores the data in the local cache; if the subsequent instruction needs to read the register data of the same address again, the operand collector gives priority to obtaining the data from the cache, and only when the data does not exist in the cache will it be read again from the unified storage space.
[0035] Compared with the prior art, the system for reading and writing inter-thread operands of the present invention has the following beneficial effects:
[0036] 1. The present invention uses registers as the only medium for data sharing, and manages register data and states between threads in a unified manner through the "thread remapping" and "cross-thread marking" methods, thereby achieving safe data sharing between threads. This avoids the synchronization operations between thread warps required when using shared memory, eliminates the extra overhead caused by synchronization waiting, and improves the execution efficiency of the program, especially in multi-threaded parallel computing scenarios;
[0037] 2. The operand collector of the present invention directly reads registers from a unified storage space to complete cross-thread data reading, thereby reducing the data transmission process between multiple storage levels and reducing read and write delays. In addition, the operand collector adopts a cache mechanism. When a subsequent instruction needs to read register data at the same address again, it is obtained from the cache first, further reducing the access time to the storage space and improving data reading efficiency.
[0038] 3. The present invention creates a score table for each thread bundle through a scoreboard, controls the instruction execution order, queries the register write-back status flag, avoids data correlation risks, and realizes data synchronization control; during the instruction execution process, the instruction that has completed data write-back is preferentially selected for operand collection, which ensures the consistency and correctness of the data and guarantees the security of data sharing between threads;
[0039] 4. The present invention significantly improves the parallel performance of GPGPU in multi-threaded parallel computing scenarios by optimizing multiple aspects such as reducing synchronization overhead, lowering read and write latency, improving cache resource utilization, and ensuring data security consistency. It is particularly suitable for artificial intelligence applications such as neural networks based on matrix operations, and can better meet the needs of large-scale parallel computing tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Attached Figure 1 is a system architecture diagram of Embodiment 1 of the present invention;
[0041] Attached Figure 2 is a detailed architecture diagram of Embodiment 2 of the present invention;
[0042] Attached Figure 3 Schematic diagram of read-after-write hazard of data in an embodiment of the present invention;
[0043] Attached Figure 4 It is a schematic diagram of the write-after-write hazard of data in an embodiment of the present invention. DETAILED DESCRIPTION
[0044] In order to make the technical solution, the technical problem solved and the technical effect of the present invention more clearly understood, the technical solution of the present invention is clearly and completely described below in conjunction with specific embodiments.
[0045] Embodiment 1:
[0046] Combined with Figure 1 , this embodiment proposes a system for cross-thread reading and writing of inter-thread operands, and its implementation architecture includes a decoding unit, a transmitting unit, and a write-back unit. These three units cooperate with each other to form a complete cross-thread data reading and writing path. Among them, the decoding unit is responsible for decoding custom instructions and remapping thread indexes; the transmitting unit is responsible for instruction caching, data synchronization control, and cross-thread data reading; the write-back unit is responsible for writing the instruction execution results back to the storage space and updating the status of related registers. Data is transmitted in sequence between the units, and the results processed by the decoding unit are passed to the transmitting unit. After the transmitting unit completes the data reading, it sends the instructions and data to the execution unit, and the results of the execution unit are passed to the write-back unit for write-back operation.
[0047] In this embodiment, the decoding unit is composed of a custom instruction decoder and a thread remapping module, wherein: the decoder decodes the custom instruction and extracts control information; the thread remapping module calculates the real thread index according to the extracted control information.
[0048] Before the decoder decodes the custom instruction, it is necessary to expand the RISC-V instruction set and define the memory access instructions for operands between warps, the memory access instructions for operands between threads of different warps, and the memory access instructions for operands between threads in the same warp. The opcodes of these custom instructions are allocated in the encoding space not used by the original RISC-V instruction set to ensure compatibility with the original instruction set. The format of the custom instruction retains the RISC-V basic instruction format framework and adds fields for identifying control information such as thread index and thread index offset.
[0049] The decoder decodes the custom instructions after the RISC-V instruction set is extended and extracts control information, including thread index and thread index offset; the thread remapping module adds the thread index and thread index offset of the instruction to calculate the actual thread index.
[0050] It should be added that: when decoding custom instructions, the legality of the instructions needs to be verified: if it is found that the instruction contains illegal control information (such as the thread index exceeds the reasonable range, the opcode does not comply with the custom instruction set definition, etc.) or does not meet the format requirements of the custom instruction set (such as field length error, required field missing, etc.), the decoding unit refuses to decode the instruction and feeds back error information (such as error type, error location and other detailed information) to the system.
[0051] The thread remapping module takes the boundary of the thread bundle into consideration when calculating the actual thread index. The preset rules are as follows: If the calculated thread index exceeds the valid range of the thread bundle, when it exceeds the upper limit, the thread index is modulo the maximum valid index value of the thread bundle; when it exceeds the lower limit, the thread index is corrected by adding the maximum valid index value of the thread bundle to ensure the validity and accuracy of the thread index.
[0052] In this embodiment, the emission unit includes an instruction queue, a scoreboard, and an operand collector, wherein: the instruction queue caches decoded instructions; the scoreboard creates a score table for each thread bundle, controls the instruction execution order, and avoids data correlation hazards by querying the register write-back status flag to achieve data synchronization control; the operand collector manages the registers of all threads in a unified manner, creates a collection table for each emission channel, initializes the collection table according to the instruction control information, calculates the offset address, reads the register from the unified storage space, and completes cross-thread data reading.
[0053] (1) The instruction queue caches the decoded instructions corresponding to the same thread warp to prepare for subsequent instruction execution order control and operand collection.
[0054] (2) The scoreboard controls the execution order of the instructions corresponding to each thread bundle through the following operations:
[0055] First, the score table allocates a 1-bit status bit to each register of the thread to record the write-back status of the register;
[0056] Then, when an instruction needs to read a register across threads, the scoreboard queries the status flag corresponding to the register to be accessed based on the decoded thread index and register number to determine whether the data has been written back;
[0057] Finally, the scoreboard traverses the instructions in the instruction queue and gives priority to instructions that have completed data writeback for operand collection, avoiding correlation hazards and achieving data synchronization control.
[0058] refer to Figure 3 and Figure 4 , a schematic diagram of data dependency hazards in GPGPU.
[0059] Figure 3This is a typical read-after-write hazard diagram. Figure 3 Instruction i and instruction j belong to the same thread, and instruction i is executed before instruction j. The register to be written back by instruction i and the register to be read by instruction j are the same register. If instruction i is relatively complex and requires more execution time, instruction j will read the old value of the register before instruction i writes the calculation result back to the register (instruction j should read the calculation result of instruction i, that is, the new value), resulting in an error in the execution result.
[0060] Figure 4 A typical write-after-write adventure diagram. Figure 4 Instruction i and instruction j belong to the same thread, and instruction i is executed before instruction j. The register to be written back by instruction i and instruction j is the same register. If instruction i is relatively complex and requires more execution time, instruction j will update the value of the register in advance before instruction i writes the calculation result back to the register (instruction i should be written to the register first, instruction j should be written to the register later, and the final register will hold the execution result of instruction j), resulting in an error in the execution result.
[0061] These two types of risks are more likely to occur when collecting cross-thread operands. The present invention designs a scoreboard strategy for cross-threads to avoid the occurrence of the above data risks from the underlying logic and achieve cross-thread data synchronization.
[0062] It should be added that the scoreboard will evaluate the priority of instructions when controlling the execution order of instructions. For instructions with high real-time requirements (such as key instructions in real-time graphics rendering) or instructions that are critical to the entire calculation process (such as core calculation instructions in matrix multiplication), they will be judged according to preset evaluation criteria (such as the functional type of the instruction, the position in the calculation process, etc.). On the basis of satisfying the data write-back status judgment, it will be marked as executable first, even if there are other instructions in the queue that have completed data write-back at this time.
[0063] At the same time, the scoreboard will record the waiting time of the instruction in the queue. If the waiting time of an instruction exceeds the preset threshold (the threshold is dynamically adjusted according to system performance and actual application scenarios) and the required data has been written back, the scoreboard will adopt the preset processing mechanism. The processing mechanism includes: when the waiting time is too long and the system resources allow, the instruction will be advanced to the head of the queue for priority execution; if the system resources are tight, a warning message (such as a specific error code or log record) will be issued to prompt the system administrator to intervene.
[0064] (3) The operand collector performs the following operations to read operands between threads:
[0065] First, each collection table contains n collection entries, where n is determined by the maximum number of registers that may be used by the instruction. Each collection entry includes the following fields: a valid bit field, indicating whether the operand corresponding to the entry needs to be collected; a source register number, indicating the register where the operand is located; an offset address field, indicating the address of the register in the storage space; a data field, used to temporarily store the read operand; and a ready bit field, indicating whether the operand corresponding to the entry has been collected.
[0066] Then, the operand collector initializes the collection table according to the control information of the instruction, where the offset address is the absolute address of the register to be read in the storage space, which needs to be calculated based on the decoded thread index and register number. The calculation system is shown in formula (1):
[0067] addr offset ={sRsx[bit:0],wIndex}Formula (1),
[0068] In the formula, addr offset Indicates the offset address; sRsx indicates the information related to the source register, sRsx[bit:0] indicates taking the part of the source register sRsx from the bit to the 0th bit, which may contain part of the register address or thread-related index information; wIndex indicates the index of the thread bundle, which is used to identify different thread bundles and locate the position of the thread bundle in the entire computing resource;
[0069] Finally, the operand collector reads the register from the unified storage space according to the offset address to complete the cross-thread data reading.
[0070] Based on the above operations, after the transmit unit completes data reading, it sends the instruction and data to the execution unit, and the result of the execution unit is then passed to the write-back unit for write-back operation.
[0071] It should be added that when the operand collector initializes the collection table, the setting of the valid bit field follows the following rules:
[0072] a) When the instruction explicitly requires the collection of a certain operand, the valid bit field of the corresponding collection entry is set to the valid state;
[0073] b) When the instruction clearly does not need to collect a certain operand, the corresponding valid bit field is set to an invalid state;
[0074] At the same time, the operand collector initializes the ready bit field in the collection table to the not-ready state to ensure that the operand will not be used incorrectly when it has not completed collection;
[0075] After initializing the collection table according to the instruction control information, the operand collector checks the integrity of the collection table. If it is found that key information is missing (such as the source register label is empty, the offset address calculation is wrong, etc.) or wrong in the collection table, it will issue an error signal (such as a specific error code and prompt message) and stop the current operation, waiting for further processing.
[0076] It should be added that the operand collector in the emission unit adopts a cache mechanism when reading register data from the unified storage space.
[0077] The cache size is dynamically adjusted according to the system's memory resources and actual application scenarios, and is generally set to accommodate a certain amount of commonly used register data. The cache update strategy uses the LRU (least recently used) algorithm. When the cache occupancy reaches the preset threshold, the least recently used data is removed from the cache. When the operand collector reads the register data for the first time according to the offset address, the data is stored in the local cache; if the subsequent instruction needs to read the register data at the same address again, the operand collector will first obtain the data from the cache, and only when the data does not exist in the cache will it be read again from the unified storage space. Theoretical analysis and actual tests show that this cache mechanism can shorten the data reading time by an average of [X]%, significantly improving the overall performance of the system.
[0078] In this embodiment, the write-back unit extracts the destination operand from the instruction execution result, calculates the destination register offset address according to the destination thread index and the destination register label, writes the data into the storage space, and notifies the scoreboard to modify the write-back status of the corresponding register, thereby completing the cross-thread data writing and forming a complete cross-thread data reading and writing path.
[0079] Through the above-mentioned inter-thread operand collection strategy and application system, significant performance improvement has been achieved in multi-threaded parallel computing scenarios.
[0080] Embodiment 2:
[0081] Combination Figure 1 and Figure 2 , This embodiment proposes a system for reading and writing inter-thread operands across threads, and takes the inter-thread register read and write instructions in the matrix calculation scenario as an example to specifically describe the execution method of the inter-thread operand collection method.
[0082] 1. In the decoding unit:
[0083] (1.1) The custom decoder recognizes that the instruction is a cross-thread register read and write instruction, and extracts the thread index, thread index offset, and register label information from the instruction.
[0084] (1.2) The thread remapping module calculates the actual accessed thread index based on the thread index and thread index offset, and all control information is sent to the emission unit.
[0085] 2. In the transmitting unit:
[0086] (2.1) The instruction is cached in the instruction queue corresponding to the thread index.
[0087] (2.2) The scoreboard continuously takes instructions from the instruction queue and queries the score table where the register is located. If the data of the register has not been written back, the instruction is blocked. If the data of the register has been written back, the instruction is selected first for operand collection.
[0088] (2.3) The operand collector initializes the operand collection table according to the control information of the instruction, enables the valid bit field and calculates the address offset. When all operands in the instruction are collected, the data stored in the data field is packaged and sent to the execution unit for calculation.
[0089] 3. In the write-back unit:
[0090] (3.1) The calculation results of the execution unit (specifically the GPGPU core computing unit) and the control information related to the instruction are sent to the write-back unit.
[0091] (3.2) The write-back unit advances the destination operand from the execution result and calculates the destination offset address based on the destination thread index and the destination register number in the control information.
[0092] (3.3) The destination operand and offset address are sent back to the emission unit, and the operand collector directly writes the data into the storage space. At the same time, the scoreboard modifies the write-back status of the register, unlocks the blocked instruction, and completes data sharing between threads.
[0093] In this embodiment, the write-back unit extracts the destination operand from the instruction execution result, calculates the destination register offset address according to the destination thread index and the destination register label, writes the data into the storage space, and notifies the scoreboard to modify the write-back status of the corresponding register, thereby completing the cross-thread data writing and forming a complete cross-thread data reading and writing path.
[0094] It should be added that the storage space here refers to the global memory space of GPGPU or the designated register storage space, which depends on the system design and data storage requirements. ① For GPGPU global memory space: GPGPU has a large capacity global memory, which is used to store various types of data during program execution, including matrix data and intermediate results in matrix calculations. After completing the matrix calculation, the write-back unit calculates the destination offset address, and the operand collector writes the destination operand to the corresponding address location in the global memory. This is because the global memory can be accessed by all threads, and the data written here can be conveniently used by other subsequent threads or computing tasks. For example, in large-scale matrix calculations, the matrix sub-block results calculated by different thread warps will be written back to the global memory so that they can be finally combined into a complete matrix calculation result. ② Designated register storage space: In addition to the global memory, the system may also have a register space dedicated to storing specific data. These registers may be designed to store data that is frequently accessed or has a significant impact on computing performance. The operand collector writes the destination operand to the location specified by the offset address in the register storage space. For example, some registers used to store key parameters or intermediate results in the calculation process will have the new results written back to the corresponding register addresses after the calculation is completed so that subsequent instructions can quickly read and use them, reducing data access delays and improving computing efficiency.
[0095] It is important to know that the storage space written by the operand collector is mainly used to store calculation results for use in subsequent calculation processes. It is not the storage location inside the emission unit, but a data storage area related to the entire GPGPU computing architecture and accessible to multiple threads.
[0096] In summary, the system for cross-thread reading and writing of inter-thread operands using the present invention can solve the problem of data sharing between threads, reduce synchronization overhead, optimize data reading and writing processes, reduce latency, improve cache resource utilization, ensure data security and consistency, and enhance GPGPU parallel performance.
[0097] The above specific examples are used to explain the principles and implementation methods of the present invention in detail. These examples are only used to help understand the core technical content of the present invention. Based on the above specific embodiments of the present invention, any improvements and modifications made by technicians in this technical field without departing from the principles of the present invention should fall within the scope of patent protection of the present invention.
Claims
1. A system for reading and writing inter-thread operands, characterized in that: Its implementation architecture includes: The decoding unit is composed of a custom instruction decoder and a thread remapping module, wherein: the decoder decodes the custom instruction and extracts the control information; the thread remapping module calculates the real thread index according to the extracted control information; The emission unit includes an instruction queue, a scoreboard, and an operand collector, wherein: the instruction queue caches decoded instructions; the scoreboard creates a score table for each thread bundle, controls the instruction execution order, and avoids data correlation risks by querying the register write-back status flag to achieve data synchronization control; the operand collector manages the registers of all threads in a unified manner, creates a collection table for each emission channel, initializes the collection table according to the instruction control information, calculates the offset address, reads the register from the unified storage space, and completes cross-thread data reading; after completing the data reading, the emission unit sends the instruction and data to the execution unit, and the result of the execution unit is then passed to the write-back unit for write-back operation; The write-back unit extracts the destination operand from the instruction execution result, calculates the destination register offset address according to the destination thread index and the destination register label, writes the data into the storage space, and notifies the scoreboard to modify the write-back status of the corresponding register, completing the cross-thread data writing and forming a complete cross-thread data reading and writing path.
2. A system for reading and writing inter-thread operands according to claim 1, characterized in that: Extend the RISC-V instruction set to define inter-warp operand memory access instructions, inter-thread operand memory access instructions of different warps, and inter-thread operand memory access instructions within the same warp. The decoder decodes the custom instruction after the RISC-V instruction set is extended, and extracts control information, including thread index and thread index offset; The thread remapping module adds the thread index and thread index offset of the instruction to calculate the actual thread index.
3. A system for reading and writing inter-thread operands according to claim 2, characterized in that: When the decoding unit decodes the custom instruction, it is necessary to verify the legality of the instruction: if it is found that the instruction contains illegal control information or does not meet the format requirements of the custom instruction set, the decoding unit refuses to decode the instruction and feeds back the error information to the system; When calculating the actual thread index, the thread remapping module will consider the boundary conditions of the thread warp: if the calculated thread index exceeds the valid range of the thread warp, the thread remapping module will correct the thread index according to the preset rules to ensure the validity and accuracy of the thread index.
4. A system for reading and writing inter-thread operands according to claim 1, characterized in that: The instruction queue caches decoded instructions corresponding to the same warp.
5. A system for reading and writing inter-thread operands according to claim 4, characterized in that: The scoreboard specifically controls the execution order of the instructions corresponding to each thread warp through the following operations: First, the score table allocates a 1-bit status bit to each register of the thread to record the write-back status of the register; Then, when an instruction needs to read a register across threads, the scoreboard queries the status flag corresponding to the register to be accessed based on the decoded thread index and register number to determine whether the data has been written back; Finally, the scoreboard traverses the instructions in the instruction queue and gives priority to instructions that have completed data writeback for operand collection, avoiding correlation hazards and achieving data synchronization control.
6. A system for reading and writing inter-thread operands according to claim 5, characterized in that: When controlling the execution order of instructions, the scoreboard will evaluate the priority of the instructions: for instructions with high real-time requirements or critical to the entire computing process, the scoreboard will prioritize marking them as executable on the basis of satisfying the data write-back status judgment, even if there are instructions in the queue that have completed data write-back at this time; At the same time, the scoreboard will record the waiting time of the instruction in the queue. If the waiting time of an instruction exceeds the preset threshold and the required data has been written back, the scoreboard will adopt the preset processing mechanism to adjust the execution order of the instructions or issue a warning message to avoid long-term blocking of instructions and affect system performance.
7. A system for reading and writing inter-thread operands according to claim 5, characterized in that: The operand collector specifically completes the inter-thread operand reading by the following operations: First, each collection table contains n collection entries, where n is determined by the maximum number of registers that may be used by the instruction. Each collection entry includes the following fields: a valid bit field, indicating whether the operand corresponding to the entry needs to be collected; The source register number indicates the register where the operand is located; The offset address field indicates the address of the register in the storage space; the data field is used to temporarily store the operands read; the ready bit field indicates whether the operands corresponding to the entry have been collected; Then, the operand collector initializes the collection table according to the control information of the instruction, where the offset address is the absolute address of the register to be read in the storage space, which needs to be calculated based on the decoded thread index and register number. The calculation system is shown in formula (1): addr offset ={sRsx[bit:0],wIndex}Formula (1), In the formula, addr offset Indicates the offset address; sRsx indicates the information related to the source register, sRsx[bit:0] indicates taking the part of the source register sRsx from the bit to the 0th bit, which may contain part of the register address or thread-related index information; wIndex indicates the index of the thread bundle, which is used to identify different thread bundles and locate the position of the thread bundle in the entire computing resource; Finally, the operand collector reads the register from the unified storage space according to the offset address to complete the cross-thread data reading.
8. A system for reading and writing inter-thread operands according to claim 7, characterized in that: When the operand collector initializes the collection table, the setting of the valid bit field follows the following rules: a) When the instruction explicitly requires the collection of a certain operand, the valid bit field of the corresponding collection entry is set to a valid state; b) When the instruction clearly does not need to collect a certain operand, the corresponding valid bit field is set to an invalid state; At the same time, the operand collector initializes the ready bit field in the collection table to the not-ready state to ensure that the operand will not be used incorrectly when it has not completed collection; After initializing the collection table according to the instruction control information, the operand collector checks the integrity of the collection table. If it is found that key information is missing or wrong in the collection table, an error signal will be issued and the current operation will be stopped, waiting for further processing.
9. A system for reading and writing inter-thread operands according to claim 7, characterized in that: The operand collector in the emission unit adopts a cache mechanism when reading register data from the unified storage space; When the operand collector reads the register data for the first time according to the offset address, it stores the data in the local cache; If a subsequent instruction needs to read register data at the same address again, the operand collector will first obtain the data from the cache. Only when the data does not exist in the cache will it be read again from the unified storage space.
Citation Information
Patent Citations
Robust, efficient multiprocessor-coprocessor interface
CN110858387A
Method for GPU read-write unit to access register file through operand collector
CN112817639A