Multi-thread processor, synchronization method, electronic equipment and storage medium
By introducing the synchronization control unit and the first barrier instruction in the multithreaded processor, the synchronization problem between thread sets in the multithreaded processor is solved, data consistency and correct execution order are achieved, and higher-level thread set synchronization is supported.
Patent Information
- Application Number
- CN202510330480.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-08
AI Technical Summary
In a multi-core multi-threaded processor, how to ensure the data accuracy and consistency of multiple threads or thread groups when accessing and modifying shared resources, existing barrier instructions cannot effectively achieve synchronization between thread sets at a higher level than execution groups.
A synchronization control unit independent of the computing unit is designed. By receiving the first type of information and execution completion information, the barrier synchronization between multiple thread sets is controlled. The first barrier instruction and the second barrier instruction are used for the barrier synchronization of the thread set and the execution group respectively to ensure the correct order execution between the multiple thread sets.
It realizes barrier synchronization between multiple thread collections to ensure data consistency and support more flexible application scenario design.
Smart Images

Figure CN120276877A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to a multi-threaded processor, a synchronization method, an electronic device, and a storage medium. Background Art
[0002] In the field of modern computing, with the rapid development of technology, multi-core and multi-threaded processors have become one of the key technologies to improve computing efficiency and processing power. Multi-core and multi-threaded processors such as Central Processing Unit (CPU), Graphics Processing Unit (GPU), and General-purpose computing on graphics processing units (GPGPU) can process a large amount of data in parallel. For example, GPGPU utilizes its highly parallel architecture to process a large amount of data in parallel through thousands of computing units (stream processors), which can greatly improve the processing speed in fields such as scientific computing, artificial intelligence, and big data analysis.
[0003] A multi-core processor can execute multiple threads simultaneously. By dividing threads into thread groups, multiple threads can be organized at a higher level, enabling multiple threads to share resources and cooperate with each other, thereby optimizing resource utilization and management. When a multi-core and multi-threaded processor executes multiple tasks in parallel, multiple threads or thread groups may simultaneously access and modify shared resources. Therefore, how to ensure the correctness and consistency of data has become an urgent problem to be solved. Summary of the Invention
[0004] At least one embodiment of the present disclosure provides a multi-threaded processor, the multi-threaded processor comprising: a plurality of computing units and a synchronization control unit coupled to the plurality of computing units, wherein each of the plurality of computing units includes a scheduling module and an execution module, each scheduling module being configured to, in response to the currently executed instruction being a first instruction of a first type that needs to participate in a thread set barrier, send first type information of the first instruction to the synchronization control unit, and send the first instruction to the execution module in the same computing unit, and in response to the currently executed instruction being a first barrier instruction of a second type, send second type information of the first barrier instruction to the synchronization control unit; each execution module being configured to receive the first instruction from the scheduling module, and send execution completion information of the first instruction to the synchronization control unit after executing the first instruction; the synchronization control unit being configured to receive the first type information and the second type information from the scheduling modules of the plurality of computing units, receive the execution completion information from the execution modules of the plurality of computing units, and control barrier synchronization between a plurality of thread sets corresponding to the plurality of computing units according to the first type information, the second type information, and the execution completion information.
[0005] For example, in the multi-threaded processor provided by at least one embodiment of the present disclosure, the plurality of thread sets are a plurality of thread groups, each of the plurality of thread groups includes a plurality of execution groups, and each execution group executes at least one first barrier instruction, and the first barrier instruction is used to synchronize the plurality of thread groups divided into the same thread group block.
[0006] For example, in the multi-threaded processor provided by at least one embodiment of the present disclosure, the synchronization control unit includes a first register group, the first register group includes a first barrier storage area, a second barrier storage area, a third barrier storage area, and a fourth barrier storage area, the first barrier storage area is used to store a total number of first execution groups, and the total number of first execution groups represents the number of execution groups in the plurality of thread groups; the second barrier storage area is used to store a first current number of execution groups, and the first current number of execution groups represents the number of execution groups that have reached the first barrier instruction in the plurality of thread groups; the third barrier storage area is used to store a total number of first instructions, and the total number of first instructions represents the number of first instructions in the plurality of thread groups; the fourth barrier storage area is used to store a number of completed instructions, and the number of completed instructions is used to represent the number of first instructions that have been executed and completed in the plurality of thread groups.
[0007] For example, in the multi-threaded processor provided by at least one embodiment of the present disclosure, the synchronization control unit is further configured to: in response to receiving the first type of information, increase the total number of the first instructions by a first value to update the total number of the first instructions; in response to receiving the second type of information, increase the first current number of execution groups by a second value to update the first current number of execution groups; in response to receiving the execution completion information, increase the number of completed instructions by a third value to update the number of completed instructions.
[0008] For example, in the multi-threaded processor provided by at least one embodiment of the present disclosure, the synchronization control unit is further configured to: in response to the first current number of execution groups being equal to the total number of the first execution groups, determine whether the number of completed instructions is equal to the total number of the first instructions; in response to the number of completed instructions being equal to the total number of the first instructions, send synchronization completion information to the scheduling module, and initialize the first current number of execution groups, the total number of the first instructions, and the number of completed instructions for the next barrier synchronization.
[0009] For example, in the multi-threaded processor provided by at least one embodiment of the present disclosure, the scheduling module is further configured to: in response to the currently executed instruction being the first barrier instruction and not receiving the synchronization completion information from the synchronization control unit, set the barrier state of the execution group corresponding to the first barrier instruction to a first state, and stop executing the instructions in the execution group that are after the first barrier instruction; in response to receiving the synchronization completion information from the synchronization control unit, set the barrier state of the execution group corresponding to the first barrier instruction to a second state, and continue to execute the instructions in the execution group that are after the first barrier instruction.
[0010] For example, in the multi-threaded processor provided by at least one embodiment of the present disclosure, the synchronization control unit includes a plurality of first register groups assigned to a plurality of thread group blocks, and each thread group block includes at least two thread groups.
[0011] For example, in the multi-threaded processor provided by at least one embodiment of the present disclosure, the scheduling module includes an instruction fetch sub-module, a decoding sub-module, and an issuing sub-module. The instruction fetch sub-module is configured to obtain the currently executed instruction, and the decoding sub-module is configured to parse the currently executed instruction to obtain the instruction type information of the instruction, where the instruction type information includes the first type of information and the second type of information.
[0012] The transmitting sub-module is configured to: in response to the instruction type information being the first type of information, determine that the instruction is the first instruction, send the first type of information of the first instruction to the corresponding first register group in the synchronization control unit according to the thread group block index information corresponding to the first instruction, and send the first instruction and the execution group information corresponding to the first instruction to the execution module in the same computing unit; in response to the instruction type information being the second type of information, determine that the instruction is the first barrier instruction, and send the second type of information of the first barrier instruction to the corresponding first register group in the synchronization control unit according to the thread group block index information corresponding to the first barrier instruction.
[0013] For example, in the multi-threaded processor provided in at least one embodiment of the present disclosure, the instruction includes a first coding bit, and the first coding bit is used to indicate whether the instruction is the first instruction that needs to participate in the thread set barrier.
[0014] For example, in the multi-threaded processor provided in at least one embodiment of the present disclosure, the first instruction includes a memory read instruction, a memory write instruction, or a computing instruction.
[0015] For example, in the multi-threaded processor provided in at least one embodiment of the present disclosure, each of the multiple thread groups includes multiple execution groups, and each execution group executes at least one second barrier instruction, and the second barrier instruction is used to synchronize the multiple execution groups in each thread group.
[0016] For example, in the multi-threaded processor provided in at least one embodiment of the present disclosure, the scheduling module further includes a second register group, and the second register group includes a fifth barrier storage area and a sixth barrier storage area. The fifth barrier storage area is used to store the total number of second execution groups, and the total number of second execution groups represents the number of execution groups in the thread group corresponding to the computing unit where the scheduling module is located; the sixth barrier storage area is used to store the current number of second execution groups, and the current number of second execution groups represents the number of execution groups that have executed to the second barrier instruction in the thread group corresponding to the computing unit where the scheduling module is located.
[0017] At least one embodiment of the present disclosure further provides a synchronization method for a multi-threaded processor. The multi-threaded processor includes a plurality of computing units and a synchronization control unit coupled to the plurality of computing units. The synchronization method includes: through a scheduling module in each computing unit, in response to the currently executed instruction being a first instruction of a first type that needs to participate in a thread set barrier, sending first type information of the first instruction to the synchronization control unit, and sending the first instruction to an execution module in the same computing unit, and in response to the currently executed instruction being a first barrier instruction of a second type, sending second type information of the first barrier instruction to the synchronization control unit; through an execution module in each computing unit, receiving the first instruction from the scheduling module, and sending execution completion information of the first instruction to the synchronization control unit after executing the first instruction; through the synchronization control unit, receiving the first type information and the second type information from the scheduling modules of the plurality of computing units, receiving the execution completion information from the execution modules of the plurality of computing units, and controlling barrier synchronization between a plurality of thread sets corresponding to the plurality of computing units according to the first type information, the second type information, and the execution completion information.
[0018] For example, in the synchronization method provided by at least one embodiment of the present disclosure, the plurality of thread sets are a plurality of thread groups, each thread group in the plurality of thread groups includes a plurality of execution groups, and each execution group executes at least one first barrier instruction, and the first barrier instruction is used to synchronize the plurality of thread groups divided into the same thread group block.
[0019] For example, in the synchronization method provided by at least one embodiment of the present disclosure, the synchronization control unit includes a first register group. The first register group includes a first barrier storage area, a second barrier storage area, a third barrier storage area, and a fourth barrier storage area. The first barrier storage area is used to store a total number of first execution groups, and the total number of first execution groups represents the number of execution groups in the plurality of thread groups; the second barrier storage area is used to store a first current number of execution groups, and the first current number of execution groups represents the number of execution groups in the plurality of thread groups that have executed to the first barrier instruction; the third barrier storage area is used to store a total number of first instructions, and the total number of first instructions represents the number of first instructions in the plurality of thread groups; the fourth barrier storage area is used to store a number of completed instructions, and the number of completed instructions is used to represent the number of first instructions that have been executed and completed in the plurality of thread groups.
[0020] The synchronization method further includes: in response to the synchronization control unit receiving the first type of information, increasing the total number of the first instructions by a first value to update the total number of the first instructions; in response to the synchronization control unit receiving the second type of information, increasing the first current number of execution groups by a second value to update the first current number of execution groups; in response to the synchronization control unit receiving the execution completion information, increasing the number of completed instructions by a third value to update the number of completed instructions.
[0021] For example, the synchronization method provided by at least one embodiment of the present disclosure further includes: in response to the first current number of execution groups stored in the synchronization control unit being equal to the total number of the first execution groups, determining whether the number of completed instructions is equal to the total number of the first instructions; in response to the number of completed instructions being equal to the total number of the first instructions, sending the synchronization completion information in the synchronization control unit to the scheduling module, and initializing the first current number of execution groups, the total number of the first instructions, and the number of completed instructions for the next barrier synchronization.
[0022] For example, the synchronization method provided by at least one embodiment of the present disclosure further includes: in response to the currently executed instruction being the first barrier instruction and the scheduling module not receiving the synchronization completion information from the synchronization control unit, setting the barrier state of the execution group corresponding to the first barrier instruction to a first state, and stopping the execution of the instructions in the execution group that are after the first barrier instruction; in response to the scheduling module receiving the synchronization completion information from the synchronization control unit, setting the barrier state of the execution group corresponding to the first barrier instruction to a second state, and continuing the execution of the instructions in the execution group that are after the first barrier instruction.
[0023] For example, the synchronization method provided by at least one embodiment of the present disclosure further includes: dividing the plurality of thread groups into a plurality of thread group blocks through the synchronization control unit, and respectively allocating a plurality of first register groups to the plurality of thread group blocks, wherein each thread group block includes at least two thread groups.
[0024] At least one embodiment of the present disclosure further provides an electronic device, including: at least one memory, non-transiently storing computer-executable instructions; at least one processor, configured to run the computer-executable instructions, wherein the computer-executable instructions, when run by the processor, implement the synchronization method described in any of the above embodiments.
[0025] At least one embodiment of the present disclosure further provides a non-transient computer-readable storage medium, wherein the non-transient computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions, when executed by at least one processor, implement the synchronization method described in any of the above embodiments. Description of the Drawings
[0026] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description only relate to some embodiments of the present disclosure and do not limit the present disclosure.
[0027] Figure 1A It is a schematic structural diagram of a general graphics processing unit;
[0028] Figure 1B It is a schematic diagram of the execution pipeline of a scheduling module;
[0029] Figure 1C It is a schematic diagram of some registers in a scheduling module provided by at least one embodiment of the present disclosure;
[0030] Figure 2 It is a schematic block diagram of a multi-threaded processor provided by at least one embodiment of the present disclosure;
[0031] Figure 3A It is a schematic diagram of a thread group block provided by at least one embodiment of the present disclosure;
[0032] Figure 3B It is a schematic diagram of a set of thread group blocks provided by at least one embodiment of the present disclosure;
[0033] Figure 4 It is a schematic diagram of some registers in a synchronization control unit provided by at least one embodiment of the present disclosure;
[0034] Figure 5 It is a schematic diagram of an exemplary instruction sequence provided by at least one embodiment of the present disclosure;
[0035] Figure 6 It is a flowchart of a synchronization method provided by at least one embodiment of the present disclosure;
[0036] Figure 7 It is a schematic diagram of a non-transitory computer-readable storage medium provided by at least one embodiment of the present disclosure;
[0037] Figure 8 It is a schematic structural diagram of an electronic device provided by at least one embodiment of the present disclosure. Detailed Embodiments
[0038] In order to make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.
[0039] Unless otherwise defined, the technical terms or scientific terms used in the present disclosure shall have the ordinary meanings understood by those of ordinary skill in the art to which the present disclosure pertains. The terms "first", "second", and similar terms used in the present disclosure do not denote any order, quantity, or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or items appearing before this word cover the elements or items listed after this word and their equivalents, without excluding other elements or items. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left", "right", etc. are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly. It should be understood that the steps recorded in the method embodiments of the present disclosure may be executed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.
[0040] The present disclosure will be described below through several specific embodiments. To keep the following description of the embodiments of the present disclosure clear and concise, detailed descriptions of known functions and known components may be omitted. When any component of the embodiments of the present disclosure appears in more than one drawing, the component is denoted by the same or similar reference numerals in each drawing.
[0041] The design of multi-core and multi-thread processors is for parallel processing, allowing the processor to execute multiple tasks or instruction streams simultaneously, thereby improving computing efficiency and system performance. For example, a central processing unit (CPU) usually includes multiple computing units, which are also called cores. Each core is a relatively independent processing unit, having its own arithmetic logic unit (ALU), control unit, register set, and other components, and can execute instructions independently. Similarly, a graphics processing unit (GPU) and a general-purpose graphics processing unit (GPGPU) also include multiple computing units, and each computing unit can execute instructions independently and is usually used for parallel computing.
[0042] Figure 1AIt is a schematic structural diagram of a general - purpose graphics processor. As Figure 1A shown, the general - purpose graphics processor includes a command processor 110, a thread - group distribution unit 120, multiple computing units 130, and a cache 140. Each computing unit 130 includes a scheduling module 131, multiple arithmetic - logic units (ALUs) 132, and general - purpose registers 133.
[0043] In parallel computing, the tasks to be executed generally include multiple threads (workitems). Before these threads are executed in the general - purpose graphics processor (or called parallel computing processor) as Figure 1A shown, they will be divided into multiple thread groups (workgroups) in the command processor 110 first, and then distributed to each computing unit 130 via the thread - group distribution unit 120. All threads in a thread group must be assigned to the same computing unit 130 for execution. Multiple thread groups can be executed in the same computing unit 130.
[0044] At the same time, the thread group will be split into the smallest execution thread groups (hereinafter referred to as execution groups for short), and each execution group contains a fixed number (or less than this fixed number) of threads. For example, 32 threads. In the computing unit 130, according to the number of ALUs 132 and other modules in the computing unit 130, multiple execution groups in a thread group can be executed simultaneously or time - shared. Multiple threads in each execution group will execute the same instruction.
[0045] The reading, decoding, and issuing of instructions are all completed in the scheduling module 131. Different types of instructions are completed in different execution modules. For example, computing instructions will be issued to the ALU 132 for computing operations, and memory instructions will be issued to the cache 140 for memory - related operations. Memory instructions include memory read instructions, memory write instructions, and other instructions related to memory operations. Memory read instructions are used to implement the read operation of reading data from memory, and memory write instructions are used to implement the write operation of writing data into memory.
[0046] When an execution group executes a compute instruction, the source data required (from general register 133) may come from a previous memory read instruction. In this case, a wait instruction is needed to ensure that the data read back by the previous memory read instruction is ready. In the same thread group, there is also such a synchronization relationship among multiple execution groups. For example, a thread group includes two execution groups, namely execution group 0 and execution group 1. The computations of these two execution groups both require reading data from memory area A. To save read time and bandwidth, a common optimization method is for execution group 0 to read half of the data in memory area A and execution group 1 to read the other half of the data in memory area A. However, for any execution group, all the data in memory area A is required when the execution group executes a compute instruction. Therefore, before the execution group executes the compute instruction, it is necessary to wait until the memory read instructions of execution group 0 and execution group 1 have read back the data. At this time, a barrier instruction is needed to prevent the execution group from continuing to execute instructions until the memory read instructions of execution group 0 and execution group 1 are completed, in order to obtain all the data in memory area A.
[0047] Figure 1B It is a schematic diagram of the execution pipeline of a scheduling module. As Figure 1B shown, multiple execution groups can simultaneously fetch instructions (instruction fetch) and decode instructions (instruction decode) in the scheduling module. In the decode stage, if a barrier instruction or a memory wait instruction is found, it will be determined whether to advance the Program Counter (PC) to execute the next instruction according to whether the conditions are met. For example, Figure 1B the line with the letter a shown. In the issue stage, according to the types and quantities of the execution modules, a suitable execution group will be selected from multiple execution groups according to a certain arbitration logic to issue the instruction to the execution module for execution, and at the same time, the PC used for instruction fetch will be advanced forward. For example, Figure 1B the line with the letter b shown.
[0048] Figure 1C It is a schematic diagram of some registers in a scheduling module provided by at least one embodiment of the present disclosure. As Figure 1C shown, in the scheduling module, a certain number of registers are reserved for the barrier mechanism, so as to force all execution groups in the same thread group to continue to execute subsequent instructions only when they reach the barrier instruction, thereby realizing the barrier synchronization of multiple execution groups within the same thread group. For example, as Figure 1CAs shown, 3 register groups can be reserved in the scheduling module, and these 3 register groups respectively correspond to barrier 0, barrier 1, and barrier 2. Then the computing unit where the scheduling module is located can support 3 different thread groups to process barrier instructions inside their respective units. For example, thread group 0 can be divided into 4 execution groups. When dividing the execution groups, an available barrier register group (such as barrier 0) will be selected and assigned to thread group 0, and the corresponding barrier register group will be initialized according to the number of execution groups in this thread group 0. For example, the total number of execution groups in the barrier register group (barrier 0) corresponding to thread group 0 is initialized to 4, the current number of execution groups is initialized to 0, and the barrier ids corresponding to all execution groups within thread group 0 are set to barrier 0.
[0049] In the decoding stage, in response to the instruction currently executed by the execution group being a barrier instruction, the current number of execution groups stored in the barrier register group corresponding to the barrier id of the execution group is incremented by 1. Additionally, the barrier instruction of this execution group will block the execution group from continuing to execute instructions (i.e., the PC will not advance). For example, in response to the instruction currently executed by any one of execution groups 0 to 3 being a barrier instruction, the current number of execution groups in barrier 0 is incremented by 1. When the current number of execution groups in barrier 0 is equal to the total number of execution groups, it indicates that all execution groups 0 to 3 of thread group 0 have executed up to this barrier instruction. At this time, the scheduling module advances the PCs of all execution groups 0 to 3 of this thread group 0 and clears the current number of execution groups stored in the barrier register group (barrier 0) corresponding to thread group 0 for the next barrier.
[0050] The inventors of the present disclosure have noticed that this type of barrier instruction can only function as a barrier within the same thread group and cannot serve as a barrier between thread groups. That is, this type of barrier instruction can only achieve a barrier between multiple execution groups within the same thread group and cannot achieve a barrier between thread sets at a higher level than the execution group (such as thread groups or thread group blocks, etc.). For example, in some cases, thread group 0 writes data to memory, thread group 1 reads the data from memory and uses the data for calculation. The behavior of thread group 0 writing data to memory needs to be completed first (for example, the memory write instruction is executed by the execution module), and then thread group 1 reads the data. However, this barrier instruction cannot ensure that thread group 0 finishes writing the data first and then thread group 1 reads the data.
[0051] To address the deficiencies of the above solution, at least one embodiment of the present disclosure provides a new barrier mechanism. For a first instruction that needs to participate in a thread set barrier at a higher level than the execution group, a first barrier instruction for synchronizing multiple thread sets is designed. A synchronization control unit independent of the scheduling module of the computing unit is added. The synchronization control unit receives the first type information and execution completion information of the first instruction and the second type information of the first barrier instruction to control the barrier synchronization between multiple thread sets, thereby ensuring that the first instructions between multiple thread sets can be executed in the correct order and ensuring the consistency of data between multiple thread sets, facilitating users to design application scenarios more flexibly.
[0052] For example, in some embodiments of the present disclosure, a first encoding bit can be added to the instruction. The first encoding bit is used to indicate whether the instruction needs to participate in the thread set barrier. For example, when encoding the instruction, a 1-bit first encoding bit can be added to the instruction. When the first encoding bit is 1, it indicates that the instruction needs to participate in the thread set barrier. For ease of description, the embodiments of the present disclosure refer to the instruction that needs to participate in the thread set barrier as the "first instruction"; when the first encoding bit is 0, it indicates that the instruction does not need to participate in the thread set barrier. Here, the first instruction that needs to participate in the thread set barrier can be understood as an instruction that has an association relationship with the execution of instructions in other thread sets. This association relationship can be reflected, for example, in the execution order of two instructions in different thread sets having a sequential relationship. For example, in an example, the execution group in thread set B1 executes instruction 1, and the execution group in thread set B2 executes instruction 2. The association relationship between instruction 1 and instruction 2 requires that after executing instruction 1 (the first instruction) in thread set B1, instruction 2 in thread set B2 can be executed.
[0053] For example, in some embodiments of the present disclosure, the first instruction can be a memory instruction, a computing instruction, or other types of instructions that need to participate in the thread set barrier. The embodiments of the present disclosure do not limit the type of the first instruction.
[0054] For example, in some embodiments of the present disclosure, memory instructions include memory read instructions, memory write instructions, and other instructions related to accessing memory. Here, the memory can be, for example, shared memory, global memory, or cache. Shared memory is located on each stream processor, close to the execution unit. The size of shared memory is relatively small, and the access speed is fast, which is conducive to data sharing and communication between threads in the same thread block. Global memory is the largest memory area on the graphics processor and can be accessed by all threads. Compared with shared memory, the capacity of global memory is much larger, but the access latency is higher. In addition, to improve efficiency, a cache can also be used to reduce the impact of global memory access latency.
[0055] Figure 2 A schematic block diagram of a multi-threaded processor provided in at least one embodiment of the present disclosure. Figure 2 As shown, the multi-threaded processor provided by at least one embodiment of the present disclosure includes a synchronization control unit 210 and multiple computing units 220 , and each computing unit 220 of the multiple computing units 220 includes a scheduling module 221 and an execution module 222 .
[0056] Each scheduling module 221 is configured to, in response to the currently executed instruction being a first instruction of the first type that needs to participate in the thread set barrier, send first type information of the first instruction to the synchronization control unit 210, and send the first instruction to the execution module 222 in the same computing unit 220, and in response to the currently executed instruction being a first barrier instruction of the second type, send second type information of the first barrier instruction to the synchronization control unit 210.
[0057] Each execution module 222 is configured to receive a first instruction from the scheduling module 221 , and send execution completion information of the first instruction to the synchronization control unit 210 after executing the first instruction.
[0058] The synchronization control unit 210 is configured to receive first type information and second type information from the scheduling module 221 of multiple computing units 220, receive execution completion information from the execution module 222 of multiple computing units 220, and control the barrier synchronization between multiple thread sets corresponding to the multiple computing units 220 based on the first type information, the second type information and the execution completion information.
[0059] For example, in some embodiments of the present disclosure, the currently executed instruction may be an instruction of a different type, for example, the currently executed instruction may be a first instruction, a first barrier instruction, or another instruction. The first instruction is an instruction that needs to participate in the thread set barrier; the first barrier instruction is a barrier instruction used to prevent the execution group of the thread set to which it belongs from continuing to execute instructions. In other words, the first barrier instruction is an instruction used to implement a barrier between thread sets; other instructions are instructions that do not need to participate in the thread set barrier. Here, other instructions may be, for example, memory wait instructions or barrier instructions used to implement barrier synchronization between multiple execution groups of the same thread group. For ease of description, the embodiments of the present disclosure refer to barrier instructions used to implement barrier synchronization between multiple execution groups of the same thread group as "second barrier instructions."
[0060] For example, in some embodiments of the present disclosure, an instruction has instruction type information, which includes first type information and second type information. The first type information is used to indicate that the currently executed instruction is a first instruction of the first type. The first type information is sent by a scheduling module 221 in a computing unit 220 to a synchronization control unit 210, so that the synchronization control unit 210 can record the number of first instructions that appear in a plurality of thread sets corresponding to the computing unit 220. The second type information is used to indicate that the currently executed instruction is a first barrier instruction of the second type. Similarly, the second type information is sent by the scheduling module 221 in the computing unit 220 to the synchronization control unit 210, so that the synchronization control unit 210 can record the number of first barrier instructions that appear in a plurality of thread sets corresponding to the computing unit 220, which is the same as the number of execution groups that execute to the first barrier instruction.
[0061] Embodiments of the present disclosure do not limit the specific forms (data types) of the first type information and the second type information. For example, the first type information and the second type information can be represented by numbers such as "0" and "1", or can be represented by strings such as "instruction" and "execution group", respectively.
[0062] The synchronization control unit 210 and multiple computing units 220 are coupled through an interface to exchange information. Embodiments of the present disclosure add some information for thread set barriers on the interfaces between the synchronization control unit 210 and multiple computing units 220, and on the interfaces between the scheduling module 221 and the execution module 222 of the computing unit 220. Such information includes, but is not limited to, the above-mentioned first type information, second type information, and instruction completion information, etc.
[0063] For example, index information can also be added on an interface 11 between the synchronization control unit 210 and the scheduling module 221, an interface 12 between the synchronization control unit 210 and the execution module 222, and an interface 13 between the scheduling module 221 and the execution module 222. The index information is used to find the thread set barriers corresponding to a plurality of thread sets.
[0064] For example, instruction type information and synchronization completion information for indicating the end of a thread group barrier can be added on interface 11 between the synchronization control unit 210 and the scheduling module 221 in each computing unit 220. The instruction type information includes first type information of a first instruction and second type information of a first barrier instruction. For example, during a thread group barrier process, the scheduling module 221 can send the first type information of the first instruction and the second type information of the first barrier instruction to the synchronization control unit 210 through interface 11, so that the synchronization control unit 210 records the number of first instructions to be executed and the number of execution groups that have reached the first barrier instruction during this thread group barrier process; after the synchronization of this thread group barrier ends, the synchronization control unit 210 can send the synchronization completion information to the scheduling module 221 through interface 11, so that the scheduling module 221 advances the program counter (PC) to continue executing the instructions after the first barrier instruction.
[0065] For example, execution completion information of the first instruction can be added on interface 12 between the synchronization control unit 210 and the execution module 222 in each computing unit 220. For example, during a thread group barrier process, after the execution module 222 finishes executing the first instruction, it can send the execution completion information of the first instruction to the synchronization control unit 210 through interface 12, so that the synchronization control unit 210 records the number of first instructions that have been executed.
[0066] For example, instruction information of the first instruction can be added on interface 13 between the scheduling module 221 and the execution module 222 in each computing unit 220. During a thread group barrier process, the scheduling module 221 can send the first instruction and the instruction information of the first instruction to the execution module 222 through interface 13, so that the first instruction can be executed in the execution module 222. For example, the execution module 222 and the scheduling module 221 are located in the same computing unit 220.
[0067] Embodiments of the present disclosure can, by designing the first instruction and the first barrier instruction, and adding information for thread group barriers on the interfaces between the synchronization control unit 210 and multiple computing units 220 and on the interfaces between the scheduling module 221 and the execution module 222 of the computing unit 220, record the number of first instructions to be executed according to the first type information by the synchronization control unit 210, record the number of first instructions that have been executed according to the execution completion information, and record the number of execution groups that have reached the first barrier instruction according to the second type information, so that after it is determined that all execution groups of multiple related thread groups have reached the first barrier instruction and all first instructions to be executed have been executed by the execution module 222, the execution groups are allowed to continue executing the instructions after the first barrier instruction, thereby achieving barrier synchronization between multiple thread groups.
[0068] For example, in an embodiment of the present disclosure, a "thread set" refers to a set of multiple threads whose grouping level is higher than that of an execution group, and the level of a thread set barrier is higher than that of an execution group barrier. Since the first barrier instruction and the second barrier instruction are barrier instructions respectively used for the thread set barrier mechanism and the execution group barrier mechanism, the barrier level corresponding to the first barrier instruction is higher than the barrier level corresponding to the second barrier instruction.
[0069] For example, in some examples, the thread set can be a "thread group". At this time, the thread set barrier refers to a barrier between multiple thread groups. Multiple thread groups can be divided into a thread group block. After multiple thread groups are divided into a thread group block, the multiple thread groups and their affiliated thread group block index information are sent to the computing unit together, and the thread group block index information is recorded in the status register of each thread group for subsequent use.
[0070] Figure 3A A schematic diagram of a thread group block provided for at least one embodiment of the present disclosure. As Figure 3A shown, the multiple thread sets participating in the thread set barrier are multiple thread groups in a thread group block, and each thread group includes multiple execution groups. For example, thread group block 0 includes 4 thread groups (thread groups 0 to 3), and thread group 0 includes 2 execution groups.
[0071] For example, in some other examples, the thread set can be a "thread group block", and the thread group block includes multiple thread groups. At this time, the thread set barrier refers to a barrier between multiple thread group blocks. Multiple thread group blocks can be divided into a thread group block set. Similarly, after multiple thread group blocks are divided into a thread group block set, the multiple thread group blocks and their affiliated thread group block set index information are sent to the computing unit together, and the thread group block set index information is recorded in the status register of each thread group for subsequent use.
[0072] Figure 3B A schematic diagram of a thread group block set provided for at least one embodiment of the present disclosure. As Figure 3B shown, the multiple thread sets participating in the thread set barrier are multiple thread group blocks in a thread group block set, each thread group block includes multiple thread groups, and each thread group includes multiple execution groups. For example, thread group block set 0 includes 4 thread group blocks (thread group blocks 0 to 3), thread group block 0 includes 2 thread groups (thread group 0 and thread group 1), and thread group 0 includes 2 execution groups.
[0073] For the convenience of description, in the following, the thread set is taken as an example of a thread group to illustrate the embodiments of the present disclosure. In the following, the thread set barrier is referred to as a thread group barrier, that is, the first barrier instruction mentioned below can be used to implement barrier synchronization between thread groups. Those skilled in the art should understand that the embodiments of the present disclosure can also be applied to implement barrier synchronization between thread group blocks.
[0074] For example, in some embodiments of the present disclosure, multiple thread sets are multiple thread groups, and each thread group in the multiple thread groups includes multiple execution groups. It should be noted that the number of execution groups included in a thread group can be set according to the actual hardware circuit situation, and the embodiments of the present disclosure do not limit this.
[0075] For example, in some embodiments of the present disclosure, each execution group may include multiple threads, for example, 32 threads, 64 threads, etc. The embodiments of the present disclosure do not limit the number of threads in an execution group.
[0076] For example, in some embodiments of the present disclosure, some registers are reserved in the synchronization control unit 210 for thread group barriers. For example, some registers can be directly allocated to multiple thread groups with an association relationship for the barrier synchronization of these thread groups, or multiple thread groups can be first divided into a thread group block, and then thread group barrier resources can be allocated to this thread group block. As Figure 3A shown, multiple thread groups 0 to 3 correspond to thread group block 0, and a first register group can be allocated to thread group block 0 for the barriers of the thread groups within this thread group block.
[0077] Figure 4 is a schematic diagram of some registers in a synchronization control unit 210 provided by at least one embodiment of the present disclosure. As Figure 4As shown in the figure, the synchronization control unit 210 includes a plurality of first register groups to be used as thread group barrier resources for a plurality of thread group blocks. For example, thread group barriers 0 to 3 can be allocated to thread group blocks 0 to 3 respectively, so that 4 thread group blocks can be supported to process the first barrier instruction. For example, each first register group may include a first barrier storage area 201, a second barrier storage area 202, a third barrier storage area 203, and a fourth barrier storage area 204. The first barrier storage area 201 is used to store the total number of first execution groups, and the total number of first execution groups represents the number of all execution groups in a plurality of thread groups divided into the same thread group block. The second barrier storage area 202 is used to store the current number of first execution groups, and the current number of first execution groups represents the number of execution groups that have reached the first barrier instruction among the plurality of thread groups. The third barrier storage area 203 is used to store the total number of first instructions, and the total number of first instructions represents the number of first instructions that need to be executed among the plurality of thread groups; the fourth barrier storage area 204 is used to store the number of completed instructions, and the number of completed instructions is used to represent the number of first instructions that have been executed and completed among the plurality of thread groups.
[0078] For example, when the synchronization control unit 210 allocates available first register groups to thread group blocks, it can initialize the total number of first execution groups, the current number of first execution groups, the total number of first instructions, and the number of completed instructions in the first barrier storage area 201, the second barrier storage area 202, the third barrier storage area 203, and the fourth barrier storage area 204 of the first register group.
[0079] For example, in one example, as Figure 4 shown, thread group block 0 consists of 4 thread groups, and each thread group includes two execution groups. That is, thread group block 0 includes 8 execution groups. During initial configuration, the synchronization control unit 210 sets the thread group block index in thread group barrier 0 to 0 (that is, set to the thread group block id) to indicate that this first register group is allocated to thread group block 0 to be used as the barrier resource for a plurality of thread groups in thread group block 0; sets the total number of first execution groups in the first barrier storage area 201 to 8, sets the current number of first execution groups in the second barrier storage area 202 to 0, sets the total number of first instructions in the third barrier storage area 203 to 0, and sets the number of completed instructions in the fourth barrier storage area 204 to 0. For example, during initial configuration, the scheduling module 221 can also set the barrier status of all execution groups in this thread group block 0 to the initial state (the second state in the embodiments of the present disclosure), for example, 0. In this state, the execution group can execute instructions normally.
[0080] For example, in some embodiments of the present disclosure, at least one first barrier instruction is included in the instruction sequence of each execution group in a plurality of thread groups within the same thread group block.
[0081] For example, in one example, if the instruction sequence of each execution group in multiple thread groups within the same thread group block includes a first barrier instruction, then the multiple thread groups in this thread group block can experience a thread group barrier. During this thread group barrier process, any thread group in this thread group block will be in a waiting state after executing the first barrier instruction until all thread groups in this thread group block have executed the first barrier instruction, and then the PCs in each execution group can continue to advance.
[0082] For example, in another example, if the instruction sequence of each execution group in multiple thread groups within the same thread group block includes multiple first barrier instructions, then the multiple thread groups in this thread group block can experience multiple thread group barriers. After each thread group barrier ends, all execution groups in this thread group block can continue to execute the instructions after the first barrier instruction. At the same time, the synchronization control unit 210 can re-initialize and configure the first register group for the next thread group barrier to start a new round of thread group barrier process until all instructions in the instruction queue of this thread group block are executed.
[0083] For example, after all instructions in a thread group block have been executed, the thread group barrier resources allocated to this thread group block can be released, and the released thread group barrier resources can be re-allocated to other thread group blocks that need to perform thread group barriers.
[0084] For example, after the synchronization control unit 210 completes the initial configuration, it can receive the first type information of the first instruction and the second type information of the first barrier instruction from the scheduling modules 221 of each computing unit 220, receive the execution completion information of the first instruction from the execution modules 222 of each computing unit 220, and record this information in their corresponding first register groups. Specifically, when the synchronization control unit 210 responds to receiving the first type information, it can increase the total number of first instructions by a first value to update the total number of first instructions; when it responds to receiving the second type information, it can increase the first current number of execution groups by a second value to update the first current number of execution groups; when it responds to receiving the execution completion information, it can increase the number of completed instructions by a third value to update the number of completed instructions. For example, the first value, the second value, and the third value can be the same and are all 1.
[0085] For example, the synchronization control unit 210 may control the barrier synchronization of the thread groups in the thread group blocks corresponding to the respective first register groups according to the information recorded in the respective first register groups. For example, after a certain thread group block completes the current thread group barrier synchronization, the synchronization control unit 210 may send the synchronization completion information of this thread group block to the scheduling module 221 in the computing unit 220 where it is located, so that all execution groups of this thread group block can cross the first barrier instruction and continue to execute new instructions. Specifically, the synchronization control unit 210 may, in response to the first current number of execution groups being equal to the first total number of execution groups, determine whether the number of completed instructions is equal to the first total number of instructions; if the number of completed instructions is equal to the first total number of instructions, it sends the synchronization completion information to the scheduling module 221. If the number of completed instructions is not equal to the first total number of instructions, it means that there are still first instructions not executed yet, then each execution group in this thread group block needs to continue waiting until all the first instructions participating in this thread group barrier are executed, that is, until the execution module 222 sends the execution completion information of all the first instructions to the synchronization control unit 210 to make the number of completed instructions equal to the first total number of instructions.
[0086] For example, in the case where the instruction sequence of the execution groups of a certain thread group block includes multiple first barrier instructions, when the first current number of execution groups in the first register group corresponding to this thread group block is equal to the first total number of execution groups and the number of completed instructions is equal to the first total number of instructions, the synchronization control unit 210 may also initialize the first current number of execution groups, the first total number of instructions, and the number of completed instructions in this first register group for the next barrier synchronization of multiple thread groups in this thread group block.
[0087] For example, the scheduling module 221 may set the barrier state of the execution group according to the first barrier instruction and the control of the synchronization control unit 210, so as to control the execution group to pause or continue executing instructions. Specifically, the scheduling module 221 may, in response to the currently executed instruction being the first barrier instruction and not receiving the synchronization completion information from the synchronization control unit 210, set the barrier state of the execution group corresponding to the first barrier instruction to the first state and stop executing the instructions after the first barrier instruction in this execution group; it may, in response to receiving the synchronization completion information from the synchronization control unit 210, set the barrier state of the execution group corresponding to the first barrier instruction to the second state and continue executing the instructions after the first barrier instruction in this execution group.
[0088] For example, the initial state of the barrier state of the execution group is the second state (e.g., 0). In the second state, the execution group can execute instructions normally (including the first instruction, the first barrier instruction, etc.). When the execution group executes the first barrier instruction, the scheduling module 221 can query whether all the execution groups in the thread group block where the execution group is located have executed the first barrier instruction, or can query whether the synchronization completion information of the thread group block where the execution group is located is received from the synchronization control unit 210. If there are still execution groups in the thread group block that have not executed the first barrier instruction, or the scheduling module 221 does not receive the synchronization completion information from the synchronization control unit 210, the scheduling module 221 can set the barrier state of the execution group to the first state (e.g., 1). Since the barrier state of the execution group is the first state, the first barrier instruction of the execution group will always block the execution of the instructions after the first barrier instruction. Until the scheduling module 221 receives the synchronization completion information of the thread group block from the synchronization control unit 210, the scheduling module 221 sets the barrier state of the execution group to the second state (e.g., 0), so that the execution group can continue to execute the instructions after the first barrier instruction.
[0089] For example, in some embodiments of the present disclosure, the scheduling module 221 includes an instruction fetching sub-module, a decoding sub-module, and an issuing sub-module. The instruction fetching sub-module is configured to obtain the currently executed instruction. The decoding sub-module is configured to parse the currently executed instruction to obtain the instruction type information of the instruction (including the first type information and the second type information). The issuing sub-module is configured to: in response to the instruction type information being the first type information, determine that the instruction is the first instruction, send the first type information of the first instruction to the corresponding first register group in the synchronization control unit 210 according to the thread group block index information corresponding to the first instruction, and send the first instruction and the execution group information corresponding to the first instruction to the execution module 222 in the same computing unit 220, and, in response to the instruction type information being the second type information, determine that the instruction is the first barrier instruction, and send the second type information of the first barrier instruction to the corresponding first register group in the synchronization control unit 210 according to the thread group block index information corresponding to the first barrier instruction.
[0090] For example, after performing group decoding, if it is found that the first instruction (e.g., a memory instruction) that needs to participate in the thread group barrier, the scheduling module 221 sends the thread group block index information and the first type information with the instruction type information being "instruction" to the synchronization control unit 210. At the same time, the scheduling module 221 sends the instruction information and the execution group information (e.g., execution group id) to the execution module 222. After receiving the thread group index information (e.g., thread group block id), the synchronization control unit 210 finds the corresponding first register group according to the thread group block id. Since the received instruction type information is the first type information "instruction", the total number of the first instructions in the first register group is incremented by 1. After receiving the instruction information and the execution group information, the execution module 222 executes the instruction to implement the required operation.
[0091] After the operation of each memory instruction participating in the thread group barrier is completed, the execution module 222 sends the thread group block index information and the execution completion information of the memory instruction to the synchronization control unit 210. After receiving the information, the synchronization control unit 210 finds the corresponding thread group barrier register group according to the thread group block id and increments the number of completed instructions by 1.
[0092] After performing group decoding, if it is found that it is a thread group barrier instruction, the scheduling module 221 sends the thread group block index information and the second type information with the instruction type information being "execution group" to the synchronization control unit 210. At the same time, the barrier state of the execution group is set to 1. After receiving the information, the synchronization control unit 210 finds the corresponding thread group barrier register group according to the thread group block id. Since the instruction type information is "execution group", the first current execution group number is incremented by 1. Since the barrier state of the execution group is 1, the thread group barrier instruction blocks the continued execution of the execution group (i.e., the PC does not advance).
[0093] When the synchronization control unit 210 finds that the first current execution group number in the thread group barrier resource is equal to the total number of execution groups, it checks whether the total number of the first instructions is equal to the number of completed instructions. If the total number of the first instructions is not equal to the number of completed instructions, then the thread group block continues to wait for the memory instructions participating in the thread group barrier to complete the memory operation. If the total number of the first instructions is equal to the number of completed instructions, it means that the execution module 222 has completed all the memory instructions participating in the thread group barrier before the barrier instruction of the thread group block. Then, the synchronization control unit 210 resets the first current execution group number in the thread group barrier resource to 0, the total number of the first instructions to 0, and the number of completed instructions to 0 for the next barrier. At the same time, the synchronization control unit 210 sends the thread group block index information and the synchronization completion information to the scheduling module 221. After receiving the information, the scheduling module 221 resets the barrier state of the execution group associated with the thread group block to 0, so that the execution group can continue to execute instructions.
[0094] Figure 5 A schematic diagram of an exemplary instruction sequence provided by at least one embodiment of the present disclosure. For example, as Figure 5 shown, a certain thread group block consists of thread group 0 and thread group 1. Each thread group has two execution groups. For example, thread group 0 includes execution group 0 and execution group 1, and thread group 1 includes execution group 2 and execution group 3.
[0095] As Figure 5 shown, the instruction sequence of each execution group includes multiple instructions, and the multiple instructions are executed in a certain order. There can be various types of instructions in the instruction sequence of the execution group. For example, the first instructions 0-3 are instructions that need to participate in the thread group barrier, and the first barrier instruction is a barrier instruction for implementing barrier synchronization between thread groups.
[0096] For example, there can be one or more first instructions in the instruction sequence of the execution group, or there can be no first instructions. For example, as Figure 5 shown, execution group 0 has two first instructions (first instruction 0 and first instruction 1), execution group 1 has one first instruction 2, execution group 2 has one first instruction 3, and execution group 3 has no first instructions.
[0097] For example, there is one first barrier instruction in the instruction queue of each execution group participating in the thread group barrier, and the positions of the respective first barrier instructions in the instruction queues of the execution groups can be different. For example, as Figure 5 shown, the first barrier instruction of execution group 3 is executed prior to the first barrier instructions of other execution groups.
[0098] In addition to the first instructions and the first barrier instructions, the instruction queue of the execution group can also include other instructions that do not need to participate in the thread group barrier. As Figure 5 shown, the other instruction blocks in each execution group are instructions that do not need to participate in the thread barrier. The memory instruction 0 in execution group 2 is also an instruction that does not need to participate in the thread group barrier. For example, the memory instruction 0 in execution group 2 can be an instruction that only needs to participate in the barrier synchronization between the two execution groups (execution group 2 and execution group 3) in thread group 1, but does not need to participate in the barrier synchronization between the two thread groups (thread group 0 and thread group 1) in this thread group block.
[0099] Next, in conjunction with Figure 5 the implementation of the thread group barrier mechanism of the embodiments of the present disclosure will be introduced in detail.
[0100] For example, thread group 0 and thread group 1 are distributed to two computing units 220 by the synchronization control unit 210. Now it is desired to synchronize the behaviors of these 2 thread groups, that is, execution groups 0, 1, and 2 first write data to memory by executing the first instructions, and then execution group 3 reads the data from memory.
[0101] Assume that the thread group block ID is 1. The synchronization control unit 210 assigns the thread group barrier 0 to thread group block 1 and performs an initial configuration on the thread group barrier 0, so that the initial state of the thread group barrier 0 is: thread group block index = 1, the total number of the first execution groups = 4, the current number of the first execution groups = 0, the total number of the first instructions = 0, and the number of completed instructions = 0.
[0102] During the decoding stage, execution group 0 discovers the first instruction 0 and the first instruction 1, then sends information to the memory-related module (equivalent to the execution module 222), and sends the first type of information for accumulating the number of instructions to the synchronization control unit 210 twice. Each time the synchronization control unit 210 receives the first type of information, it increments the total number of the first instructions of the thread group barrier 0 by 1. Each time the memory-related module completes a first instruction, it sends an execution completion message to the synchronization control unit 210, so that the synchronization control unit 210 increments the number of completed instructions of the thread group barrier 0 by 1.
[0103] During the decoding stage, execution group 0 discovers the first barrier instruction, sets the barrier state to 1 to stop the PC from advancing, and sends the second type of information for accumulating the number of execution groups to the synchronization control unit 210 once, so that the synchronization control unit 210 increments the current number of the first execution groups of the thread group barrier 0 by 1.
[0104] Similarly, execution group 1 increments the total number of the first instructions of the thread group barrier 0 by 1 and the current number of the first execution groups by 1. Execution group 2 increments the total number of the first instructions of the thread group barrier 0 by 1 and the current number of the first execution groups by 1. Execution group 3 increments the current number of the first execution groups of the thread group barrier 0 by 1.
[0105] After the synchronization control unit 210 discovers that the total number of the first execution groups is equal to the current number of the first execution groups and both are equal to 4, it checks whether the total number of the first instructions is equal to the number of completed instructions. When the total number of the first instructions is equal to the number of completed instructions and both are equal to 4, the synchronization control unit 210 resets the current number of the first execution groups, the total number of the first instructions, and the number of completed instructions to 0, and sends a synchronization completion message to the scheduling module 221, sets the barrier states of all 4 execution groups to 0 (reset to the initial state), so that execution groups 0 to 2 continue to execute other instruction blocks after the first barrier instruction, and execution group 3 continues to execute the memory instruction 1 after the first barrier instruction. In this way, the synchronization of thread group 0 and thread group 1 is achieved.
[0106] The multi-threaded processor provided by at least one embodiment of the present disclosure can achieve barrier synchronization between multiple thread sets, such as barrier synchronization between multiple thread groups or barrier synchronization between multiple thread group blocks, thereby ensuring that the first instructions between multiple thread sets can be executed in the correct order, ensuring the consistency of data between multiple thread sets, and facilitating users to design application scenarios more flexibly.
[0107] It should be noted that, in the embodiments of the present disclosure, the first barrier instruction and the second barrier instruction can coexist in the instruction queue to be executed by the same execution group, and independently act on the respective targeted barrier mechanisms.
[0108] Similarly to the synchronization control unit 210, some registers can be reserved in the scheduling module 221 for execution group barriers.
[0109] For example, in some embodiments of the present disclosure, each thread group includes multiple execution groups, and each execution group can execute at least one second barrier instruction, which is used to synchronize multiple execution groups in each thread group. The scheduling module 221 includes multiple second register groups to be used as execution group barrier resources for multiple thread groups. For example, each second register group (e.g., Figure 1C barrier 0 in) can include a fifth barrier storage area and a sixth barrier storage area. The fifth barrier storage area is used to store the total number of second execution groups, and the total number of second execution groups represents the number of execution groups in a thread group (e.g., thread group 0) of the computing unit 220 where the scheduling module 221 is located; the sixth barrier storage area is used to store the current number of second execution groups, and the current number of second execution groups represents the number of execution groups that have reached the second barrier instruction in this thread group of the computing unit 220 where the scheduling module 221 is located.
[0110] For example, in the embodiments of the present disclosure, a "first register group" can be added to the higher-level synchronization control unit 210, while retaining the "second register group" in the scheduling module 221 of the computing unit 220, so that barrier synchronization between thread groups can be achieved through the first register group and the first barrier instruction, and barrier synchronization between execution groups can also be achieved through the second register group and the second barrier instruction.
[0111] At least one embodiment of the present disclosure also provides a synchronization method, which can be used in the multi-threaded processor in the above embodiments. For example, the multi-threaded processor includes multiple computing units and a synchronization control unit coupled to the multiple computing units. For example, the synchronization control unit can be Figure 1A the command processor 110 or the thread distribution unit 120 in.
[0112] Figure 6 is a flowchart of a synchronization method provided by at least one embodiment of the present disclosure. As Figure 6 shown, the synchronization method includes the following steps S610 to S630.
[0113] Step S610: Through the scheduling module in each computing unit, in response to the currently executed instruction being a first-type first instruction that needs to participate in a thread set barrier, send the first-type information of the first instruction to the synchronization control unit, send the first instruction to the execution module in the same computing unit, and in response to the currently executed instruction being a second-type first barrier instruction, send the second-type information of the first barrier instruction to the synchronization control unit.
[0114] Step S610: Through the execution module in each computing unit, receive the first instruction from the scheduling module, and after executing the first instruction, send the execution completion information of the first instruction to the synchronization control unit.
[0115] Step S610: Through the synchronization control unit, receive the first-type information and the second-type information from the scheduling modules of multiple computing units, receive the execution completion information from the execution modules of multiple computing units, and control the barrier synchronization between multiple thread sets corresponding to multiple computing units according to the first-type information, the second-type information, and the execution completion information.
[0116] For example, the multiple thread sets in this synchronization method can be multiple thread groups or multiple thread group chunks.
[0117] For example, in the case where this synchronization method is used to implement the barrier synchronization between multiple thread groups, each thread group includes multiple execution groups, and each execution group executes at least one first barrier instruction. At this time, the first barrier instruction is used to synchronize multiple thread groups divided into the same thread group chunk.
[0118] For example, in the case where this synchronization method is used to implement the barrier synchronization between multiple thread group chunks, each thread group chunk includes multiple thread groups, each thread group includes multiple execution groups, and each execution group executes at least one first barrier instruction. At this time, the first barrier instruction is used to synchronize multiple thread group chunks divided into the same thread group chunk set.
[0119] For example, the synchronization control unit includes a first register group, and the first register group includes a first barrier storage area, a second barrier storage area, a third barrier storage area, and a fourth barrier storage area.
[0120] For example, when the synchronization method is used to implement barrier synchronization between multiple thread groups, the synchronization method further includes: dividing multiple thread groups into multiple thread group blocks through a synchronization control unit, and respectively allocating multiple first register groups to the multiple thread group blocks, where each thread group block includes at least two thread groups. At this time, the first barrier storage area in the first register group can be used to store the total number of first execution groups, where the total number of first execution groups represents the number of execution groups in multiple thread groups within the same thread group block; the second barrier storage area can be used to store the current number of first execution groups, where the current number of first execution groups represents the number of execution groups that have reached the first barrier instruction among multiple thread groups; the third barrier storage area can be used to store the total number of first instructions, where the total number of first instructions represents the number of first instructions in multiple thread groups; the fourth barrier storage area can be used to store the number of completed instructions, and the number of completed instructions is used to represent the number of first instructions that have been executed and completed among multiple thread groups.
[0121] For example, when the synchronization method is used to implement barrier synchronization between multiple thread group blocks, the synchronization method further includes: dividing multiple thread group blocks into multiple thread group block sets through a synchronization control unit, and respectively allocating multiple first register groups to the multiple thread group block sets, where each thread group block set includes at least two thread group blocks. At this time, the first barrier storage area in the first register group can be used to store the total number of first execution groups. At this time, the total number of first execution groups represents the number of execution groups in multiple thread group blocks within the same thread group block set; the second barrier storage area can be used to store the current number of first execution groups. At this time, the current number of first execution groups can represent the number of execution groups that have reached the first barrier instruction among multiple thread group blocks; the third barrier storage area can be used to store the total number of first instructions, where the total number of first instructions represents the number of first instructions in multiple thread group blocks; the fourth barrier storage area can be used to store the number of completed instructions, and the number of completed instructions is used to represent the number of first instructions that have been executed and completed among multiple thread group blocks.
[0122] For example, in at least one embodiment, the above synchronization method further includes: in response to the synchronization control unit receiving the first type of information, increasing the total number of first instructions by a first value to update the total number of first instructions; in response to the synchronization control unit receiving the second type of information, increasing the current number of first execution groups by a second value to update the current number of first execution groups; in response to the synchronization control unit receiving the execution completion information, increasing the number of completed instructions by a third value to update the number of completed instructions.
[0123] For example, in at least one embodiment, the above synchronization method further includes: in response to the first current execution group number stored in the synchronization control unit being equal to the total number of first execution groups, determining whether the number of completed instructions is equal to the total number of first instructions; in response to the number of completed instructions being equal to the total number of first instructions, sending the synchronization completion information in the synchronization control unit to the scheduling module, and initializing the first current execution group number, the total number of first instructions, and the number of completed instructions for the next barrier synchronization.
[0124] For example, in at least one embodiment, the above synchronization method further includes: in response to the currently executed instruction being the first barrier instruction and the scheduling module not receiving the synchronization completion information from the synchronization control unit, setting the barrier state of the execution group corresponding to the first barrier instruction to the first state, and stopping the execution of the instructions in the execution group that are after the first barrier instruction; in response to the scheduling module receiving the synchronization completion information from the synchronization control unit, setting the barrier state of the execution group corresponding to the first barrier instruction to the second state, and continuing to execute the instructions in the execution group that are after the first barrier instruction.
[0125] Those skilled in the art can understand that in the above method of the specific implementation manner, the writing order of each step does not mean a strict execution order and does not constitute any limitation to the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.
[0126] The above synchronization method provided by at least one embodiment of the present disclosure can control the barrier synchronization between multiple thread sets through the first barrier instruction and the synchronization control unit, so as to ensure that the first instructions between multiple thread sets can be executed in the correct order, ensure the consistency of data between multiple thread sets, and facilitate users to design application scenarios more flexibly.
[0127] At least one embodiment of the present disclosure further provides an electronic device, including a memory and a processor. The memory stores computer-executable instructions non-transiently, and the processor is configured to run the computer-executable instructions. Among them, when the computer-executable instructions are run by the processor, the synchronization method as described in any of the above embodiments is implemented. The technical effects of this electronic device are the same as those of the above synchronization method, and will not be elaborated here.
[0128] At least one embodiment of the present disclosure further provides a non-transient computer-readable storage medium.
[0129] Figure 7 It is a schematic diagram of a non-transient computer-readable storage medium provided by at least one embodiment of the present disclosure. For example, as Figure 7As shown, one or more computer-executable instructions 701 can be non-temporarily stored on a storage medium 700. For example, when the computer-executable instructions 701 are executed by a processor, one or more steps of the synchronization method described above can be performed. The technical effect of this non-temporary storage medium is the same as that of the above synchronization method and will not be elaborated here.
[0130] For example, the above non-transitory readable storage medium is implemented as a memory, such as a volatile memory and / or a non-volatile memory. In the above embodiments, the memory can be a volatile memory, for example, it can include a random access memory (RAM) and / or a cache, etc. The non-volatile memory can include, for example, a read-only memory (ROM), a hard disk, an erasable programmable read-only memory (EPROM), a portable compact disc read-only memory (CD-ROM), a USB memory, a flash memory, etc. Various application programs (codes, instructions) and data, as well as various data used and / or generated by the application programs, can also be stored in the memory.
[0131] The above computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to execute the method of the present disclosure described above.
[0132] Computer program code for performing the operations of the present disclosure can be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages - such as Java, Smalltalk, C++, and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0133] Some embodiments of the present disclosure also provide an electronic device, which includes the multi-threaded processor of any of the above embodiments or can execute the synchronization method of any of the above embodiments.
[0134] Figure 8Schematic block diagram of an electronic device provided by at least one embodiment of the present disclosure. The electronic device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 8 The illustrated electronic device 800 is merely an example and should not impose any limitation on the functions and scope of use of the embodiments of the present disclosure.
[0135] For example, as Figure 8 shown, in some examples, the electronic device 800 includes a processing device (such as a central processing unit, a graphics processing unit, etc.) 801, and the processing device 801 may include the multi-threaded processor of any of the above embodiments, or may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 802 or the program loaded from the storage device 808 into the random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the computer system are also stored. The processing device 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. The input / output (I / O) interface 805 is also connected to the bus 804.
[0136] For example, the following components may be connected to the I / O interface 805: an input device 806 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 807 including, such as a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 808 including, for example, a magnetic tape, a hard disk, etc.; a communication device 809 including, for example, a network interface card such as a LAN card, a modem, etc. The communication device 809 may allow the electronic device 800 to communicate with other devices wirelessly or wiredly to exchange data and perform communication processing via a network such as the Internet. The driver 810 is also connected to the I / O interface 805 as needed. A removable storage medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the driver 810 as needed so that the computer program read from it can be installed into the storage device 808 as needed. Although Figure 8 the illustrated electronic device 800 includes various devices, it should be understood that it is not required to implement or include all the illustrated devices, and more or fewer devices may be alternatively implemented or included.
[0137] For example, the electronic device 800 may further include a peripheral interface (not shown in the figure), etc. The peripheral interface may be various types of interfaces, such as a USB interface, a Lightning interface, etc. The communication device 809 may communicate with the network and other devices through wireless communication. The network may be, for example, the Internet, an intranet, and / or a wireless network such as a cellular phone network, a wireless local area network (LAN), and / or a metropolitan area network (MAN). The wireless communication may use any one of a variety of communication standards, protocols, and technologies, including but not limited to Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (W-CDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Wi-Fi (e.g., based on IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, and / or IEEE 802.11n standards), Voice over Internet Protocol (VoIP), WiMAX, protocols for email, instant messaging, and / or Short Message Service (SMS), or any other suitable communication protocol.
[0138] For example, the electronic device 800 may be any device such as a mobile phone, a tablet computer, a laptop computer, an e-book, a game console, a television, a digital photo frame, a navigator, etc., or may be any combination of a data processing device and hardware. The embodiments of the present disclosure are not limited thereto.
[0139] Although the present disclosure has been described in detail above with general descriptions and specific embodiments, some modifications or improvements can be made based on the embodiments of the present disclosure, which are obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present disclosure all fall within the scope of protection required by the present disclosure.
[0140] For the present disclosure, in addition to the above exemplary content, the following points need to be noted:
[0141] (1) The drawings of the embodiments of the present disclosure only relate to the structures involved in the embodiments of the present disclosure, and other structures can refer to the general design.
[0142] (2) For the sake of clarity, in the drawings used to describe the embodiments of the present disclosure, the thickness of layers or regions is enlarged or reduced, that is, these drawings are not drawn according to the actual scale.
[0143] (3) Without conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other to obtain new embodiments.
[0144] As described above, it is only the specific implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. The protection scope of the present disclosure shall be subject to the protection scope of the claims.
Claims
1. A multi-threaded processor, comprising: A plurality of computing units and a synchronization control unit coupled to the plurality of computing units, wherein each of the plurality of computing units includes a scheduling module and an execution module, each scheduling module is configured to, in response to the currently executed instruction being a first instruction of a first type that needs to participate in a thread set barrier, send the first type information of the first instruction to the synchronization control unit, send the first instruction to the execution module in the same computing unit, and in response to the currently executed instruction being a first barrier instruction of a second type, send the second type information of the first barrier instruction to the synchronization control unit; each execution module is configured to receive the first instruction from the scheduling module and send the execution completion information of the first instruction to the synchronization control unit after executing the first instruction; the synchronization control unit is configured to receive the first type information and the second type information from the scheduling modules of the plurality of computing units, receive the execution completion information from the execution modules of the plurality of computing units, and control the barrier synchronization between the plurality of thread sets corresponding to the plurality of computing units according to the first type information, the second type information, and the execution completion information.
2. The multi-threaded processor according to claim 1, wherein, The plurality of thread sets are a plurality of thread groups, each of the plurality of thread groups includes a plurality of execution groups, and each execution group executes at least one first barrier instruction, and the first barrier instruction is used to synchronize the plurality of thread groups divided into the same thread group block.
3. The multi-threaded processor according to claim 2, wherein, The synchronization control unit includes a first register group, and the first register group includes a first barrier storage area, a second barrier storage area, a third barrier storage area, and a fourth barrier storage area, the first barrier storage area is used to store a first total number of execution groups, and the first total number of execution groups represents the number of execution groups in the plurality of thread groups; the second barrier storage area is used to store a first current number of execution groups, and the first current number of execution groups represents the number of execution groups that have executed to the first barrier instruction in the plurality of thread groups; the third barrier storage area is used to store a first total number of instructions, and the first total number of instructions represents the number of first instructions in the plurality of thread groups; the fourth barrier storage area is used to store the number of completed instructions, and the number of completed instructions is used to represent the number of first instructions that have been executed and completed in the plurality of thread groups.
4. The multi-threaded processor according to claim 3, wherein, The synchronization control unit is further configured to: in response to receiving the first type information, increase the first total number of instructions by a first value to update the first total number of instructions; in response to receiving the second type information, increase the first current number of execution groups by a second value to update the first current number of execution groups; in response to receiving the execution completion information, increase the number of completed instructions by a third value to update the number of completed instructions.
5. The multi-threaded processor according to claim 3, wherein, The synchronization control unit is further configured to: in response to the first current number of execution groups being equal to the first total number of execution groups, determine whether the number of completed instructions is equal to the first total number of instructions; In response to the number of completed instructions being equal to the total number of the first instructions, send a synchronization completion message to the scheduling module, and initialize the first current execution group count, the total number of the first instructions, and the number of completed instructions for the next barrier synchronization.
6. The multi-threaded processor according to claim 5, wherein, The scheduling module is further configured to: In response to the currently executed instruction being the first barrier instruction and not receiving the synchronization completion message from the synchronization control unit, set the barrier state of the execution group corresponding to the first barrier instruction to the first state, and stop executing the instructions in the execution group that are after the first barrier instruction; In response to receiving the synchronization completion message from the synchronization control unit, set the barrier state of the execution group corresponding to the first barrier instruction to the second state, and continue executing the instructions in the execution group that are after the first barrier instruction.
7. The multi-threaded processor according to any one of claims 2-6, wherein, The synchronization control unit includes a plurality of first register groups assigned to a plurality of thread group blocks, and each thread group block includes at least two thread groups.
8. The multi-threaded processor according to claim 7, wherein, The scheduling module includes an instruction fetch sub-module, a decoding sub-module, and an issue sub-module. The instruction fetch sub-module is configured to obtain the currently executed instruction. The decoding sub-module is configured to parse the currently executed instruction to obtain the instruction type information of the instruction, where the instruction type information includes the first type information and the second type information. The issue sub-module is configured to: In response to the instruction type information being the first type information, determine that the instruction is the first instruction, send the first type information of the first instruction to the corresponding first register group in the synchronization control unit according to the thread group block index information corresponding to the first instruction, and send the first instruction and the execution group information corresponding to the first instruction to the execution module in the same computing unit. In response to the instruction type information being the second type information, determine that the instruction is the first barrier instruction, and send the second type information of the first barrier instruction to the corresponding first register group in the synchronization control unit according to the thread group block index information corresponding to the first barrier instruction.
9. The multi-threaded processor according to any one of claims 1-6, wherein, The instruction includes a first encoding bit, and the first encoding bit is used to indicate whether the instruction is the first instruction that needs to participate in the thread collective barrier.
10. The multi-threaded processor according to any one of claims 1-6, wherein, The first instruction includes a memory read instruction, a memory write instruction, or a compute instruction.
11. The multi-threaded processor according to any one of claims 2-6, wherein, Each of the plurality of thread groups includes a plurality of execution groups, and each execution group executes at least one second barrier instruction, and the second barrier instruction is used to synchronize the plurality of execution groups in each thread group.
12. The multi-threaded processor according to claim 11, wherein, The scheduling module further includes a second register group, and the second register group includes a fifth barrier storage area and a sixth barrier storage area. The fifth barrier storage area is used to store the total number of second execution groups, and the total number of second execution groups represents the number of execution groups in the thread group corresponding to the computing unit where the scheduling module is located. The sixth barrier storage area is used to store the current number of second execution groups, and the current number of second execution groups represents the number of execution groups in the thread group corresponding to the computing unit where the scheduling module is located that have executed up to the second barrier instruction.
13. A synchronization method for a multi-threaded processor, the multi-threaded processor including a plurality of computing units and a synchronization control unit coupled to the plurality of computing units, the synchronization method comprising: Through a scheduling module in each computing unit, in response to a first instruction of a first type that needs to participate in a thread set barrier among currently executed instructions, sending first type information of the first instruction to the synchronization control unit, and sending the first instruction to an execution module in the same computing unit, and in response to the currently executed instruction being a first barrier instruction of a second type, sending second type information of the first barrier instruction to the synchronization control unit; Through an execution module in each computing unit, receiving the first instruction from the scheduling module, and after executing the first instruction, sending execution completion information of the first instruction to the synchronization control unit; Through the synchronization control unit, receiving the first type information and the second type information from the scheduling modules of the plurality of computing units, receiving the execution completion information from the execution modules of the plurality of computing units, and controlling barrier synchronization between a plurality of thread sets corresponding to the plurality of computing units according to the first type information, the second type information, and the execution completion information.
14. The synchronization method according to claim 13, wherein, The plurality of thread sets are a plurality of thread groups, each thread group in the plurality of thread groups includes a plurality of execution groups, and each execution group executes at least one first barrier instruction, and the first barrier instruction is used to synchronize the plurality of thread groups divided into the same thread group block.
15. The synchronization method according to claim 14, wherein, The synchronization control unit includes a first register group, and the first register group includes a first barrier storage area, a second barrier storage area, a third barrier storage area, and a fourth barrier storage area. The first barrier storage area is used to store a total number of first execution groups, and the total number of first execution groups represents the number of execution groups in the plurality of thread groups; The second barrier storage area is used to store a first current number of executed groups, and the first current number of executed groups represents the number of execution groups in the plurality of thread groups that have executed up to the first barrier instruction; The third barrier storage area is used to store a total number of first instructions, and the total number of first instructions represents the number of first instructions in the plurality of thread groups; The fourth barrier storage area is used to store a number of completed instructions, and the number of completed instructions is used to represent the number of first instructions that have been executed and completed in the plurality of thread groups. The synchronization method further comprises: In response to the synchronization control unit receiving the first type information, increasing the total number of first instructions by a first value to update the total number of first instructions; In response to the synchronization control unit receiving the second type information, increasing the first current number of executed groups by a second value to update the first current number of executed groups; In response to the synchronization control unit receiving the execution completion information, increasing the number of completed instructions by a third value to update the number of completed instructions.
16. The synchronization method according to claim 15, further comprising: In response to the first current execution group number stored in the synchronization control unit being equal to the total number of the first execution groups, determine whether the number of completed instructions is equal to the total number of the first instructions; In response to the number of completed instructions being equal to the total number of the first instructions, send the synchronization completion information in the synchronization control unit to the scheduling module, and initialize the first current execution group number, the total number of the first instructions, and the number of completed instructions for the next barrier synchronization.
17. The synchronization method according to claim 16, further comprising: In response to the currently executed instruction being the first barrier instruction and the scheduling module not receiving the synchronization completion information from the synchronization control unit, set the barrier state of the execution group corresponding to the first barrier instruction to the first state, and stop executing the instructions in the execution group that are after the first barrier instruction; In response to the scheduling module receiving the synchronization completion information from the synchronization control unit, set the barrier state of the execution group corresponding to the first barrier instruction to the second state, and continue to execute the instructions in the execution group that are after the first barrier instruction.
18. The synchronization method according to any one of claims 14-17 further includes: Divide the plurality of thread groups into a plurality of thread group blocks through the synchronization control unit, and respectively allocate a plurality of first register groups to the plurality of thread group blocks, wherein each thread group block includes at least two thread groups.
19. An electronic device, comprising: At least one memory, non-transiently storing computer-executable instructions; At least one processor, configured to run the computer-executable instructions, wherein, when the computer-executable instructions are run by the processor, the synchronization method according to any one of claims 13-18 is implemented.
20. A non-transitory computer-readable storage medium, wherein, The non-transient computer-readable storage medium stores computer-executable instructions, When the computer-executable instructions are executed by at least one processor, the synchronization method according to any one of claims 13-18 is implemented.
Citation Information
Cited By
Data processing method and device in assembly line, medium, equipment and product
CN121050778A
Request scheduling method, device and system, storage medium and electronic equipment
CN122152540A