Thread synchronization method, device, processing unit and system

By leveraging the collaborative work of shared memory and the thread bundle scheduler, and utilizing target synchronization barriers to achieve thread synchronization, this solves the problems of increased circuit area and power consumption caused by the increased hardware structure in existing technologies. It achieves efficient thread synchronization and asynchronous transaction operation synchronization, thereby improving the system's execution efficiency and functional correctness.

CN121705046APending Publication Date: 2026-03-20T-HEAD (SHANGHAI) SEMICON CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511824192.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-04
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Existing processing architectures require the introduction of various hardware structures to achieve thread synchronization, which increases circuit area and power consumption, raises system implementation costs, and results in inefficient synchronization mechanisms.

Method used

By working together with shared memory, thread bundle schedulers, and asynchronous acceleration units, synchronization between threads and between threads and asynchronous transaction operations is achieved using target synchronization barriers, avoiding the introduction of additional hardware structures. Synchronization control is achieved using barrier update requests and thread wake-up information.

Benefits of technology

Without increasing the hardware structure, efficient thread synchronization was achieved, reducing the occupation of memory bandwidth and interconnect resources, and improving the system's execution efficiency and functional correctness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121705046A_ABST
    Figure CN121705046A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a thread synchronization method, device, processing unit and system. Wherein the processing unit comprises a shared memory, a thread beam scheduler and an asynchronous acceleration unit. The shared memory is configured to update an internally stored target synchronization barrier based on the barrier update request, and in response to detecting that the target synchronization barrier satisfies a release condition, send thread wake-up information to the thread beam scheduler. The thread beam scheduler is configured to execute a thread sleep operation based on the arrival state of each target thread and trigger a first update request for the target synchronization barrier, and execute a thread wake-up operation based on the thread wake-up information in response to receiving the thread wake-up information. The asynchronous acceleration unit is configured to trigger a second update request for the target synchronization barrier based on a completion state of each target asynchronous transaction operation. According to the embodiment of the invention, the synchronization between the threads and between the threads and the asynchronous transaction operation can be realized on the premise of not introducing too many hardware structures.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and specifically to a thread synchronization method, apparatus, processing unit, and system. Background Technology

[0002] In existing processing architectures, synchronization between threads and between threads and asynchronous transaction operations is crucial for ensuring the correctness and efficiency of system functions. An efficient synchronization mechanism needs to balance two objectives: firstly, once the synchronization condition is met, waiting threads need to be quickly notified to resume execution to minimize idle latency; secondly, while synchronization is not yet complete, frequent polling of the synchronization state needs to be avoided to reduce the unnecessary occupation of memory bandwidth and interconnect resources. However, to achieve these functions, existing processing architectures typically require the internal introduction and integration of various hardware structures such as shared memory, dedicated buffers, and caches. This multi-module architecture not only increases circuit area and power consumption but also significantly increases system implementation costs. Summary of the Invention

[0003] In view of this, embodiments of the present invention provide a thread synchronization method, apparatus, processing unit, and system to achieve synchronization between threads and between threads and asynchronous transaction operations without introducing excessive hardware structures.

[0004] In a first aspect, embodiments of the present invention aim to provide a processing unit, the processing unit comprising: Shared memory is configured to update a target synchronization barrier in internal storage based on a barrier update request, and to send a thread wake-up message to a thread bundle scheduler in response to detecting that the target synchronization barrier meets a release condition. The target synchronization barrier is used to synchronize the execution of at least one target thread and at least one target asynchronous transaction operation. The barrier update request includes a first update request and a second update request. A thread bundle scheduler is configured to perform a thread sleep operation and trigger a first update request for the target synchronization barrier based on the arrival status of each target thread, and to perform a thread wake-up operation based on the thread wake-up information received. An asynchronous acceleration unit is configured to trigger a second update request against the target synchronization barrier based on the completion status of each of the target asynchronous transaction operations.

[0005] Secondly, embodiments of the present invention aim to provide a thread synchronization method, the method being applicable to shared memory, the method comprising: Update the target synchronization barrier in internal storage based on the barrier update request; In response to the detection that the target synchronization barrier meets the release condition, a thread wake-up message is sent to the thread bundle scheduler so that the thread bundle scheduler performs a thread wake-up operation based on the thread wake-up message. The target synchronization barrier is used to synchronize the execution of at least one target thread and at least one target asynchronous transaction operation. The barrier update request includes a first update request and a second update request.

[0006] Thirdly, embodiments of the present invention aim to provide a thread synchronization method, the method being applicable to a thread bundle scheduler, the method comprising: Based on the arrival status of each target thread, a thread sleep operation is performed and a first update request for the target synchronization barrier is triggered, so that the shared memory updates the target synchronization barrier stored internally based on the first update request, wherein the target synchronization barrier is used to synchronize the execution of at least one target thread and at least one target asynchronous transaction operation. In response to receiving thread wake-up information, a thread wake-up operation is performed based on the thread wake-up information.

[0007] Fourthly, embodiments of the present invention aim to provide a shared memory configured to perform the method described in the second aspect.

[0008] Fifthly, embodiments of the present invention aim to provide a thread bundle scheduler configured to perform the method described in the third aspect.

[0009] In a sixth aspect, embodiments of the present invention aim to provide a processing system comprising at least one processing unit as described in the first aspect.

[0010] The processing unit of this invention includes shared memory, a thread beam scheduler, and an asynchronous acceleration unit. The shared memory is configured to update the target synchronization barrier in its internal storage based on a barrier update request, and to send thread wake-up information to the thread beam scheduler in response to detecting that the target synchronization barrier meets a release condition. The thread beam scheduler is configured to perform a thread sleep operation and trigger a first update request for the target synchronization barrier based on the arrival status of each target thread, and to perform a thread wake-up operation based on the thread wake-up information received. The asynchronous acceleration unit is configured to trigger a second update request for the target synchronization barrier based on the completion status of each target asynchronous transaction operation. This invention can achieve synchronization between threads and between threads and asynchronous transaction operations without introducing excessive hardware structures. Attached Figure Description

[0011] The above and other objects, features and advantages of the present invention will become clearer from the following description of embodiments of the invention with reference to the accompanying drawings, in which: Figure 1This is a schematic diagram of the processing unit according to an embodiment of the present invention; Figure 2 This is a schematic diagram of a synchronization barrier according to an embodiment of the present invention; Figure 3 This is a schematic diagram of another processing unit according to an embodiment of the present invention; Figure 4 This is a flowchart of a thread synchronization method according to an embodiment of the present invention; Figure 5 This is a flowchart of a thread synchronization method according to an embodiment of the present invention; Figure 6 This is a flowchart of the target thread wake-up method according to an embodiment of the present invention; Figure 7 This is a schematic diagram of a thread synchronization device according to an embodiment of the present invention; Figure 8 This is a schematic diagram of a thread synchronization device according to an embodiment of the present invention. Detailed Implementation

[0012] The present application is described below based on embodiments, but it is not limited to these embodiments. In the detailed description of the present application below, certain specific details are described in detail. Those skilled in the art can fully understand the present application without these details. To avoid obscuring the substance of the present application, well-known methods, processes, flows, elements, and circuits are not described in detail.

[0013] Furthermore, those skilled in the art should understand that the accompanying drawings provided herein are for illustrative purposes only and are not necessarily drawn to scale.

[0014] Unless the context explicitly requires it, words such as "including" or "contains" throughout the application should be interpreted as including rather than exclusive or exhaustive; that is, meaning "including but not limited to".

[0015] In the description of this application, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Furthermore, in the description of this application, unless otherwise stated, "a plurality of" means two or more.

[0016] The solutions described in this specification and embodiments, if involving the processing of personal information, will be processed only on the premise of having a legal basis (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be processed within the scope stipulated or agreed upon. A user's refusal to process personal information beyond what is necessary for basic functions will not affect the user's use of basic functions.

[0017] It should be noted that, in the embodiments of the present invention, a processing unit can refer to a hardware execution unit with basic data processing functions. The implementation of the processing unit can vary depending on the architecture of the processing system. For example, in a graphics processing unit (GPU), the processing unit can be a streaming multiprocessor (SM). In a dedicated accelerator, the processing unit can be a basic arithmetic module. In a central processing unit (CPU), the processing unit can be a processing core. The processing system can refer to a composite system composed of multiple processing units coupled together via an on-chip interconnect network. For example, the processing system can be a texture processing cluster (TPC), a computation processing cluster (CPC), a graphics processing unit (GPU), a tensor processing unit (TPU), a central processing unit (CPU), and a field-programmable gate array (FPGA), etc., and this application does not impose any limitations on this.

[0018] Furthermore, it should be noted that upon receiving a task to be processed, the processing system typically first divides the task into multiple subtasks (which can be understood as thread blocks), and then assigns each subtask to a corresponding processing unit for execution. This allows the processing system to achieve parallel processing of the task, thereby improving processing efficiency. In this embodiment of the invention, thread synchronization can refer to synchronization of threads from the same thread block or synchronization of threads from different thread blocks. When synchronization between different thread blocks is involved, these thread blocks can be thread blocks assigned to the same processing unit or thread blocks assigned to different processing units; this application does not impose any limitations on this.

[0019] Figure 1 This is a schematic diagram of the processing unit according to an embodiment of the present invention. Figure 1 As shown, the processing unit 1 may include a shared memory 11, a thread bundle scheduler 12, and an asynchronous acceleration unit 13.

[0020] The shared memory 21 can be a hardware storage resource within the processing unit, accessible to threads running on the processing unit to support data sharing or collaborative data exchange among threads. The thread bundle scheduler 12 can be a thread scheduling unit within the processing unit, capable of further dividing the thread blocks allocated to the processing unit into multiple warps and dynamically scheduling ready warps to corresponding computing units within the processing unit for execution, thus achieving efficient concurrent execution of threads within a warp. The asynchronous acceleration unit 13 can be a dedicated hardware module within the processing unit, capable of performing asynchronous transaction operations to provide the necessary data resources for execution to the threads running on the processing unit. Optionally, as an option, the asynchronous acceleration unit 13 can be configured as a direct memory access engine (DMAEngine), and the asynchronous transaction operation can be configured as an asynchronous data transfer operation performed by the DMAEngine, such as a multicast loading operation (i.e., loading data from global memory into the shared memory of each processing unit).

[0021] It is important to note that, in addition to implementing the aforementioned functions, shared memory 11, thread scheduler 12, and asynchronous acceleration unit 13 can constitute a thread synchronization system to achieve synchronization between threads and between threads and asynchronous transaction operations. In this thread synchronization system, shared memory 11 can support threads creating synchronization barriers internally based on programming instructions. This synchronization barrier can be viewed as a data structure stored within shared memory 11. After a thread creates a synchronization barrier, shared memory 11 can maintain its internally stored synchronization barrier based on the thread execution status of thread scheduler 12 and the completion status of asynchronous transaction operations of asynchronous acceleration unit 13, and achieve synchronization between threads and between threads and asynchronous transaction operations based on the synchronization barrier. For ease of understanding, the following description will use a target synchronization barrier and the target thread and target asynchronous transaction operation to achieve synchronization as examples.

[0022] Specifically, in Figure 1In the processing unit shown, the thread bundle scheduler 12 can perform a thread sleep operation and trigger a first update request for the target synchronization barrier based on the arrival status of each target thread. The asynchronous acceleration unit 13 can trigger a second update request for the target synchronization barrier based on the completion status of each target asynchronous transaction operation. The shared memory 11 can update the target synchronization barrier stored internally based on the barrier update request (i.e., including the first update request and the second update request), and send thread wake-up information to the thread bundle scheduler 12 in response to detecting that the target synchronization barrier meets the release condition. Furthermore, the thread bundle scheduler 12 can also perform a thread wake-up operation based on the thread wake-up information when it receives the thread wake-up information. Thus, the embodiments of the present invention can achieve synchronization between threads and between threads and asynchronous transaction operations without introducing more hardware structures.

[0023] It is worth noting that, although Figure 1 Only a single shared memory, thread bundle scheduler, and asynchronous acceleration unit are shown in the document, but this does not mean that the number of each is limited. In actual applications, the processing unit may also include multiple shared memory, multiple thread bundle schedulers, and multiple asynchronous acceleration units. This application does not limit the number of shared memory, thread bundle schedulers, and asynchronous acceleration units included in the processing unit.

[0024] It should be noted that for any synchronization barrier, the number of threads synchronized and the number of asynchronous transaction operations synchronized by the synchronization barrier can be single or multiple, which can be determined according to the synchronization execution requirements of each thread. This application does not impose any restrictions on this.

[0025] To clarify, synchronization between threads and between threads and asynchronous transaction operations can specifically refer to the following: for multiple threads and asynchronous transaction operations that have execution dependencies on each other, before all threads have been executed to the predetermined synchronization position and all the asynchronous transaction operations they depend on have been completed, any thread that has already been executed to the predetermined synchronization position is suspended from execution; then, after all threads have been executed to the specified synchronization position and all the asynchronous transaction operations they depend on have been completed, synchronization allows all waiting threads to continue execution.

[0026] In this embodiment of the invention, the arrival status of a thread can be used to characterize whether the thread has reached the corresponding predetermined synchronization position. The completion status of an asynchronous transaction operation can be used to characterize whether the asynchronous transaction operation has been completed. Specifically, executing a thread sleep operation means converting the thread into a sleep state. A thread in a sleep state will suspend execution. Executing a thread wake-up operation means waking up a thread in a sleep state. A thread awakened from a sleep state is scheduled to the corresponding computing unit by the thread bundle scheduler and can execute successfully. It should be noted that whether a thread bundle is ready does not depend on the activity of its internal threads. Thread bundles with inactive threads can also be considered ready thread bundles. When a thread bundle with inactive threads (i.e., including threads in a sleep state) is considered a ready thread bundle, the thread bundle scheduler can use an activity mask to mask the inactive threads in the thread bundle, thereby preventing inactive threads from executing the current instruction.

[0027] Optionally, for each target thread, the thread bundle scheduler may, upon detecting that the target thread has been executed to the corresponding predetermined synchronization position, put the target thread into a sleep state and trigger a first update request for the target synchronization barrier. For each target asynchronous transaction operation, the asynchronous acceleration unit may, upon detecting that the target asynchronous transaction operation has been completed, trigger a second update request for the target synchronization barrier.

[0028] Optionally, when a thread sleep operation needs to be performed, the thread bundle scheduler can issue loop instructions to each thread waiting to sleep, enabling each thread to read the release status of the corresponding synchronization barrier and, based on the release status, either transition to a sleep state or continue execution. Specifically, when a thread sleep operation needs to be performed, the thread bundle scheduler can issue loop instructions to each thread waiting to sleep. By executing the loop instructions, the thread waiting to sleep can read the release status of the corresponding synchronization barrier and, based on the release status, either transition to a sleep state or continue execution. If the read release status indicates that the corresponding synchronization barrier is in a released state, the thread waiting to sleep can remain in a woken-up state to continue executing instructions issued by the thread scheduler in subsequent scheduling cycles. If the read release status indicates that the corresponding synchronization barrier is not in a released state, the thread waiting to sleep can transition to a sleep state.

[0029] Optionally, for any synchronization barrier, the synchronization barrier can be configured to include at least a thread counter and a transaction counter. The thread counter can be updated in response to a first update request. The transaction counter can be updated in response to a second update request. The count values ​​of the thread counter and the transaction counter can be used to characterize the thread arrival status and the completion status of asynchronous transaction operations, respectively.

[0030] Figure 2 This is a schematic diagram of a synchronization barrier according to an embodiment of the present invention. Figure 2 As shown, the synchronization barrier may include a thread counter 21 and a transaction counter 22. The count values ​​stored in the thread counter 21 and the transaction counter 22 can be updated by the shared memory based on a first update request and a second update request, respectively. The count values ​​of the thread counter 21 and the transaction counter 22 can be used to characterize the thread arrival status and the completion status of asynchronous transaction operations, respectively. Optionally, the thread counter 21 and the transaction counter 22 in the synchronization barrier 2 may have a fixed bit width. The specific value of this bit width can be set by relevant personnel according to actual needs, and this application does not impose any restrictions on it.

[0031] In this embodiment of the invention, updating the target synchronization barrier in internal storage based on the barrier update request specifically refers to updating the thread counter or transaction counter value in the target synchronization barrier based on the barrier update request. Whether the synchronization barrier meets the release condition can be determined based on the thread counter value and the transaction counter value. Specifically, the shared memory can confirm that the synchronization barrier meets the release condition when it detects that the thread counter value represents the number of threads that have arrived reaching the expected thread arrival count, and the transaction counter value represents the number of asynchronous transaction operations that have been completed reaching the expected transaction completion count.

[0032] Schematic illustration: As one way to set the thread counter and transaction counter values, the thread counter and transaction counter values ​​can be initially set to the expected thread arrival count and expected transaction completion count, respectively. During subsequent updates, whenever a barrier update request is received, the shared memory can decrement the corresponding counter value by 1 based on the barrier update request. Therefore, the shared memory can confirm that the synchronization barrier meets the release condition when it detects that both the thread counter and transaction counter values ​​are 0. Alternatively, as another way to set the thread counter and transaction counter values, the thread counter and transaction counter values ​​can be initially set to a negative expected thread arrival count and a negative expected transaction completion count, respectively. During subsequent updates, whenever a barrier update request is received, the shared memory can increment the corresponding counter value by 1 based on the barrier update request. Therefore, the shared memory can confirm that the synchronization barrier meets the release condition when it detects that both the thread counter and transaction counter values ​​are 0. Here, the expected transaction completion count can be the number of asynchronous transaction operations expected to be completed, and the expected thread arrival count can be the number of threads expected to arrive.

[0033] Optionally, to facilitate the thread bundle scheduler in confirming the target threads to be woken up, in this embodiment of the invention, the thread wake-up information may include the thread block identifier of the target thread block corresponding to the target synchronization barrier. The thread block identifier may be a unique identifier for the thread block. Upon receiving the thread wake-up information, the shared memory can wake up each target thread based on the thread block identifier in the thread wake-up information. Specifically, when performing a thread wake-up operation based on the thread wake-up information, the thread bundle scheduler can determine the corresponding target thread block based on the thread block identifier in the thread wake-up information, and then wake up all sleeping threads (i.e., threads in a sleeping state) within the target thread block, thereby realizing the wake-up operation for each target thread.

[0034] Optionally, when there are sleeping threads corresponding to different synchronization barriers within the target thread block, waking up all sleeping threads within the target thread block by the thread bundle scheduler may lead to false wake-ups. For example, in a certain thread block, thread X is waiting for the synchronization barrier at address A, and thread Y is waiting for the synchronization barrier at address B. When the synchronization barrier at address A meets its release condition, both thread X and thread Y will be woken up by the thread bundle scheduler due to the thread wake-up information. However, at this time, the synchronization barrier at address B may not have met its release condition yet, and thread Y should not actually be woken up.

[0035] To avoid false wake-ups, as one implementation, after waking up each thread, the thread scheduler can also issue loop instructions for each awakened thread. This allows each awakened thread to read the release status of the corresponding synchronization barrier and, based on the release status, either switch to a sleep state or continue execution. Specifically, after waking up each thread (including the target thread and any potentially falsely awakened threads), the thread scheduler can issue loop instructions for each awakened thread. By executing the loop instructions, the awakened thread can read the release status of the corresponding synchronization barrier and, based on the release status, switch to a sleep state or continue execution. If the read release status indicates that the corresponding synchronization barrier is in a released state, the awakened thread can remain in the awakened state and continue executing instructions issued by the thread scheduler in subsequent scheduling cycles. If the read release status indicates that the corresponding synchronization barrier is not in a released state, the awakened thread can switch back to a sleep state. Thus, this embodiment of the invention ensures that falsely awakened sleeping threads can switch back to a sleep state.

[0036] Alternatively, as another implementation, each synchronization barrier can have corresponding classification information. In addition to the thread block identifier, the thread wake-up information can also include target classification information for the target synchronization barrier. When performing a thread wake-up operation based on the thread wake-up information, the thread bundle scheduler can selectively wake up each target thread based on the thread block identifier and target classification information in the thread wake-up information. Specifically, the thread bundle scheduler can determine the corresponding target thread block and multiple sleeping threads within that target thread block based on the thread block identifier in the thread wake-up information, and then selectively determine and wake up each target thread among the multiple sleeping threads based on the target classification information. Therefore, this embodiment of the invention can avoid mistakenly waking up sleeping threads within the same thread block.

[0037] Optionally, as a method for determining classification information, the storage space within shared memory used to store synchronization barriers can be pre-divided into multiple fixed barrier storage regions. Each barrier storage region can be set with corresponding classification information, and each barrier storage region can be used to store a single synchronization barrier. The classification information corresponding to each synchronization barrier can be determined based on the storage address of the synchronization barrier in shared memory.

[0038] Schematic illustration: As a method for determining the classification information corresponding to a storage address, the storage space within shared memory used to store synchronization barriers can be pre-divided into multiple fixed barrier storage regions. Each barrier storage region can have 8 bytes. For each barrier storage region, the 3rd and 4th binary characters in the binary sequence representation of the starting address of the barrier storage region can be determined as the classification information corresponding to the barrier storage region. For example, for a barrier storage region starting at address 0x00, the classification information corresponding to the barrier storage region can be determined as 00 (that is, the binary sequence of 0x00 represents the 3rd and 4th binary characters in "00000"). Therefore, the classification information corresponding to the synchronization barrier stored in this barrier storage region can be determined as 00. For a barrier storage region starting at address 0x08, the classification information corresponding to the barrier storage region can be determined as 01 (that is, the binary sequence of 0x08 represents the 3rd and 4th binary characters in "01000"). Therefore, the classification information corresponding to the synchronization barrier stored in this barrier storage region can be determined as 01. For the barrier storage region starting at address 0x10, the classification information corresponding to the barrier storage region can be determined as 10 (that is, the binary sequence of 0x10 represents the 3rd and 4th binary characters in "10000"). Therefore, the classification information corresponding to the synchronization barrier stored in this barrier storage region can be determined as 10. For the barrier storage region starting at address 0x18, the classification information corresponding to the barrier storage region can be determined as 11 (that is, the binary sequence of 0x18 represents the 3rd and 4th binary characters in "11000"). Therefore, the classification information corresponding to the synchronization barrier stored in this barrier storage region can be determined as 11. For the barrier storage region starting at address 0x20, the classification information corresponding to the barrier storage region can be determined as 00 (that is, the binary sequence of 0x20 represents the 3rd and 4th binary characters in "100000"). Therefore, the classification information corresponding to the synchronization barrier stored in this barrier storage region can be determined as 00. Further details are omitted here.

[0039] Alternatively, as another method for determining classification information, when each synchronization barrier is created, the relevant object (e.g., shared memory or thread) can determine unique classification information for the synchronization barrier and write this classification information into the synchronization barrier (this classification information can have a fixed bit width within the synchronization barrier, such as 2 bits). Therefore, embodiments of the present invention can further avoid false wake-ups.

[0040] Further optionally, in this embodiment of the invention, to ensure that each thread can successfully read the release status of the corresponding synchronization barrier, the synchronization barrier may also be configured to include a phase counter. The count value of the phase counter can characterize the release status of the synchronization barrier. Figure 2 As shown, the synchronization barrier 2 may include a phase counter 23. It should be understood that the phase counter 23 in the synchronization barrier 2 may also have a fixed bit width, the specific value of which can be set by relevant personnel according to actual needs, and this application does not impose any restrictions on this. Schematically, as one way to set the phase counter's count value, the phase counter can be set to have two different count values, which are used to characterize different release states of the synchronization barrier. Correspondingly, after detecting that the target synchronization barrier meets the release condition, the shared memory can also update the count value of the phase counter, thereby ensuring that the count value of the phase counter can correctly characterize the release state of the synchronization barrier.

[0041] Further optionally, after confirming that the synchronization barrier meets the release condition, to avoid repeatedly triggering thread wake-up messages or causing the phase counter value to be updated again, the shared memory can also initialize the thread counter value using the expected thread arrival count and / or the transaction counter value using the expected transaction completion count after detecting that the target synchronization barrier meets the release condition. Here, the expected transaction completion count can be the number of asynchronous transaction operations expected to be completed, and the expected thread arrival count can be the number of threads expected to arrive. The expected thread arrival count and / or the expected transaction completion count can be set in the synchronization barrier. Figure 2 As shown, the synchronization barrier 2 may also include an expected transaction completion count 24. It should be understood that the expected transaction completion count 24 in the synchronization barrier 2 may also have a fixed bit width, and the specific value of this bit width can be set by relevant personnel according to actual needs. This application does not impose any restrictions on this.

[0042] Optionally, to achieve synchronization across processing units, when threads executing on different processing units need to synchronize, shared memory within different processing units can create identical associated synchronization barriers for each thread within their respective units. Furthermore, shared memory within different processing units can synchronously maintain their respective associated synchronization barriers and achieve thread synchronization based on these barriers. Data exchange between shared memory within different processing units can be implemented using multi-level cross switches connecting different processing units, or it can be implemented using global memory as an intermediary; this application does not impose any limitations on this approach.

[0043] Optionally, as an implementation method, the operations of maintaining synchronization barriers and using synchronization barriers to achieve thread synchronization can be implemented by the shared memory through its internal atomic operation logic module. This atomic operation logic module can be a dedicated hardware logic unit integrated within the shared memory to support atomic access and modification of shared data.

[0044] Optionally, in addition to the shared memory 11, the thread bundle scheduler 12, and the asynchronous acceleration unit 13, the processing unit may also include multiple computing units, instruction fetching units, instruction caches, data caches, load / store units, and register files to ensure the normal operation of the processing unit.

[0045] Figure 3 This is a schematic diagram of another processing unit according to an embodiment of the present invention. (See diagram below.) Figure 3 As shown, a thread unit may include shared memory 31, a thread bundle scheduler 32, an asynchronous acceleration unit 33, multiple computing units 34, an instruction cache 35, an instruction fetching unit 36, a data cache 37, a load / store unit 38, and a register file 39.

[0046] The descriptions of shared memory 31, thread bundle scheduler 32, and asynchronous acceleration unit 33 can be found above and will not be repeated here.

[0047] The multiple computation units 34 can be computation execution units within the processing unit, used to perform corresponding computational operations. Each thread allocated by the thread bundle scheduler 32 can be executed on each computation unit 34. Optionally, the multiple computation units 34 may include vector arithmetic logic units (ALUs), scalar arithmetic logic units (ALUs), special function units (SFUs), and tensor cores, etc., and this application does not specifically limit them.

[0048] Instruction cache 35 can be a cache used to store programming instructions for each thread.

[0049] The instruction fetching unit 36 ​​can be positioned between the thread bundle scheduler 32 and the instruction cache 35. The instruction fetching unit 36 ​​can be used to transfer instructions between the thread bundle scheduler 32 and the instruction cache 35. Specifically, the instruction fetching unit 36 ​​can fetch programming instructions from the instruction cache 35 and pre-provide the fetched programming instructions to the thread bundle scheduler 32. Once a ready thread bundle is determined, the thread bundle scheduler 32 can issue the pre-fetched instruction pointed to by the current program counter of the ready thread bundle to the computation unit 34, enabling the computation unit 34 to concurrently execute the instruction on each active thread in the ready thread bundle.

[0050] Data cache 37 can be connected to global memory (not shown in the figure) and connected to shared memory 31 through asynchronous acceleration unit 33. Data cache 37 can be used to migrate data in global memory to shared memory 31 under the instruction of asynchronous acceleration unit 33.

[0051] The load / store unit 38 can be located between the thread bundle scheduler 32 and the shared memory 31. The load / store unit can be a dedicated hardware unit within the processing unit responsible for memory read / write operations. In this embodiment, since the thread bundle scheduler 32 typically does not have the ability to access the shared memory 31, it cannot send a synchronization barrier update request to the shared memory 31. The load / store unit 38 can then be used to send a synchronization barrier update request (specifically a first update request) to the shared memory 31 based on the indication information from the thread bundle scheduler 32.

[0052] Register file 39 can be a hardware storage resource within the processing unit. Register file 39 can be used to store temporary data generated by each thread during execution, such as local variables, intermediate calculation results, function parameters, and thread identifiers.

[0053] What I want to clarify is that, Figure 3 The processing unit shown is for illustrative purposes only. In actual applications, depending on the specific needs, the processing unit may include other hardware units or functional modules, and each hardware unit or functional module may also adopt different... Figure 3 Connect using the connection method shown.

[0054] Figure 4 This is a flowchart of a thread synchronization method according to an embodiment of the present invention. It is intended to be noted that... Figure 4 The execution entity of the thread synchronization method shown can specifically be the shared memory in the above embodiments. By executing... Figure 4 The thread synchronization method shown uses shared memory in conjunction with a thread bundle scheduler and asynchronous acceleration unit to achieve synchronization between threads and between threads and asynchronous transaction operations. For example... Figure 4 As shown, the thread synchronization method may specifically include the following steps: Step S100: Update the target synchronization barrier in the internal storage based on the barrier update request.

[0055] Specifically, whenever a barrier update request is received, the shared memory can update the target synchronization barrier stored internally based on the barrier update request. The target synchronization barrier can be any synchronization barrier stored within the shared memory. The target synchronization barrier can be used to synchronize the execution of at least one target thread and at least one target asynchronous transaction operation. The barrier update request can include a first update request sent by the thread bundle scheduler and a second update request sent by the asynchronous acceleration unit. The first update request can be used to indicate that the corresponding target thread to be synchronized by the target synchronization barrier is in an arrived state, that is, it has already executed to the corresponding predetermined synchronization position. The second update request can be used to indicate that the corresponding target asynchronous transaction operation to be synchronized by the target synchronization barrier is in a completed state, that is, it has been completed.

[0056] Optionally, the target synchronization barrier may include at least a thread counter and a transaction counter. The thread counter may be updated in response to a first update request. The transaction counter may be updated in response to a second update request. The count values ​​of the thread counter and the transaction counter may be used to characterize the arrival status of the target thread and the completion status of the target asynchronous transaction operation, respectively. In step S100, when updating the target synchronization barrier, the shared memory may specifically update the count values ​​of the thread counter or the transaction counter in the target synchronization barrier based on the received barrier update request.

[0057] Optionally, the update operation performed by the shared memory can differ depending on how the thread counter and transaction counter values ​​are set. Illustratively, as one way to set the thread counter and transaction counter values, the initial values ​​can be set to the expected thread arrival count and the expected transaction completion count, respectively. With this setting, when updating the target synchronization barrier, the shared memory can decrement the corresponding counter value by 1 based on the barrier update request. Alternatively, as another way to set the thread counter and transaction counter values, the initial values ​​can be set to a negative expected thread arrival count and a negative expected transaction completion count, respectively. With this setting, when updating the target synchronization barrier, the shared memory can increment the corresponding counter value by 1 based on the barrier update request.

[0058] Step S200: In response to detecting that the target synchronization barrier meets the release condition, a thread wake-up message is sent to the thread bundle scheduler so that the thread bundle scheduler performs a thread wake-up operation based on the thread wake-up message.

[0059] Specifically, each time the target synchronization barrier is updated, the shared memory can detect that the target synchronization barrier meets the release condition. If the target synchronization barrier meets the release condition, the shared memory can send thread wake-up information to the thread bundle scheduler, so that the thread bundle scheduler can perform thread wake-up operations based on the thread wake-up information.

[0060] Optionally, in step S200, the shared memory can confirm that the target synchronization barrier meets the release condition when it detects that the thread counter's count value represents the number of target threads that have arrived reaching the expected thread arrival count, and the transaction counter's count value represents the number of target asynchronous transaction operations that have been completed reaching the expected transaction completion count. Further optionally, corresponding to the two count value setting and update methods described above, the shared memory can specifically confirm that the synchronization barrier meets the release condition when it detects that both the thread counter and the transaction counter's count values ​​are 0.

[0061] Optionally, to ensure that subsequent threads can successfully read the release status of the corresponding synchronization barrier, the target synchronization barrier also includes a phase counter. The phase counter's count value represents the release status of the synchronization barrier. The shared memory can also change the phase counter's count value after detecting that the target synchronization barrier meets the release condition. Illustratively, as one way to set the phase counter's count value, the phase counter can be set to have two different count values, each representing a different release status of the synchronization barrier. Correspondingly, to ensure that the phase counter's count value correctly represents the release status of the synchronization barrier, the shared memory can also update the phase counter's count value after detecting that the target synchronization barrier meets the release condition.

[0062] Optionally, the target synchronization barrier may further include an expected transaction completion count and / or an expected thread arrival count. The expected transaction completion count may be the number of asynchronous transaction operations expected to be completed, and the expected thread arrival count may be the number of threads expected to arrive. After confirming that the synchronization barrier meets the release condition, to avoid misjudgment leading to repeated triggering of thread wake-up information or the phase counter value being updated again, after detecting that the target synchronization barrier meets the release condition, the shared memory may also initialize the thread counter value using the expected thread arrival count, and / or initialize the transaction counter value using the expected transaction completion count.

[0063] Figure 5 This is a flowchart of a thread synchronization method according to an embodiment of the present invention. It is intended to be noted that... Figure 5 The execution entity of the thread synchronization method shown can specifically be the thread bundle scheduler in the above embodiments. By executing... Figure 5 The thread synchronization method shown uses a thread bundle scheduler in conjunction with shared memory and asynchronous acceleration units to achieve synchronization between threads and between threads and asynchronous transaction operations. For example... Figure 5 As shown, the thread synchronization method may specifically include the following steps: Step S100': Perform a thread sleep operation based on the arrival status of each target thread and trigger a first update request for the target synchronization barrier, so that the shared memory updates the target synchronization barrier in the internal storage based on the first update request.

[0064] Specifically, after determining that the thread bundles containing each target thread are ready, the thread bundle scheduler can continuously monitor the execution progress of each target thread to determine its arrival status. Then, based on the arrival status of each target thread, the thread bundle scheduler can perform a thread sleep operation and trigger a first update request for the target synchronization barrier, so that the shared memory updates the target synchronization barrier in its internal storage based on the first update request.

[0065] Optionally, in step S100', for each target thread, the thread bundle scheduler can specifically confirm that each target thread is in an arrived state when it detects that the target thread has been executed to the corresponding predetermined synchronization position. Then, the thread bundle scheduler can convert the target thread to a sleep state and trigger a first update request for the target synchronization barrier.

[0066] Optionally, when a thread sleep operation needs to be performed, the thread bundle scheduler can issue loop instructions to each thread waiting to sleep, enabling each thread to read the release status of the corresponding synchronization barrier and, based on the release status, either transition to a sleep state or continue execution. Specifically, when a thread sleep operation needs to be performed, the thread bundle scheduler can issue loop instructions to each thread waiting to sleep. By executing the loop instructions, the thread waiting to sleep can read the release status of the corresponding synchronization barrier and, based on the release status, either transition to a sleep state or continue execution. If the read release status indicates that the corresponding synchronization barrier is in a released state, the thread waiting to sleep can remain in the awakened state to continue executing instructions issued by the thread scheduler in subsequent scheduling cycles. If the read release status indicates that the corresponding synchronization barrier is not in a released state, the thread waiting to sleep can transition to a sleep state.

[0067] Step S200': In response to receiving thread wake-up information, perform thread wake-up operation based on the thread wake-up information.

[0068] Specifically, once the target synchronization barrier meets the release condition, the thread bundle scheduler can send a thread wake-up message to the thread bundle scheduler. Upon receiving the thread wake-up message, the thread bundle scheduler can perform a thread wake-up operation based on the message.

[0069] Optionally, the thread wake-up information includes a thread block identifier of the target thread block corresponding to the target synchronization barrier. The target thread block may include the thread blocks of each target thread. In step S200', upon receiving the thread wake-up information, the thread bundle scheduler can wake up each target thread based on the thread block identifier in the thread wake-up information. Specifically, the thread bundle scheduler can determine the corresponding target thread block based on the thread block identifier in the thread wake-up information, and then wake up all sleeping threads (i.e., threads in a sleeping state) within the target thread block, thereby realizing the wake-up operation for each target thread.

[0070] Optionally, to avoid false wake-ups, as one implementation, after waking up each thread, the thread scheduler can also issue loop instructions for each awakened thread, enabling each awakened thread to read the release status of the corresponding synchronization barrier and, based on the release status, either switch to a sleep state or continue execution. Specifically, after waking up each thread (including the target thread and any potentially falsely awakened threads), the thread scheduler can issue loop instructions for each awakened thread. By executing the loop instructions, the awakened thread can read the release status of the corresponding synchronization barrier and, based on the release status, switch to a sleep state or continue execution. If the read release status indicates that the corresponding synchronization barrier is in a released state, the awakened thread can remain in the awakened state and continue executing instructions issued by the thread scheduler in subsequent scheduling cycles. If the read release status indicates that the corresponding synchronization barrier is not in a released state, the awakened thread can switch back to a sleep state. Therefore, this embodiment of the invention ensures that each sleeping thread can switch back to a sleep state after being falsely awakened.

[0071] Alternatively, as another implementation, each synchronization barrier within the shared memory can be configured to have corresponding classification information, and the thread wake-up information can also include the target classification information of the target synchronization barrier. In this case, the thread bundle scheduler can also selectively wake up each target thread based on the thread block identifier and target classification information in the thread wake-up information.

[0072] Figure 6 This is a flowchart of a target thread wake-up method according to an embodiment of the present invention. It is intended to illustrate that by executing... Figure 6 The described target thread wake-up method allows the thread bundle scheduler to selectively wake up each target thread based on the thread block identifier and target classification information in the thread wake-up information. For example... Figure 6 As shown, the target thread wake-up method may specifically include the following steps: Step S210': Determine multiple sleeping threads of the target thread block based on the thread block identifier in the thread wake-up information.

[0073] Specifically, the thread bundle scheduler can determine multiple sleeping threads of the target thread block based on the thread block identifier in the thread wake-up information.

[0074] Step S220': Determine each target thread among the plurality of sleeping threads based on the target classification information.

[0075] Specifically, after identifying multiple sleeping threads of the target thread block, the thread bundle scheduler can determine each target thread among the multiple sleeping threads based on target classification information.

[0076] Step S230': Wake up each of the target threads.

[0077] Specifically, after identifying the target thread, the thread bundle scheduler can wake up each target thread in a targeted manner.

[0078] Optionally, as a method for determining classification information, the storage space within shared memory used to store synchronization barriers can be pre-divided into multiple fixed barrier storage regions. Each barrier storage region can be set with corresponding classification information, and each barrier storage region can be used to store a single synchronization barrier. The classification information corresponding to each synchronization barrier can be determined based on the storage address of the synchronization barrier in shared memory.

[0079] Alternatively, as another method for determining classification information, when each synchronization barrier is created, the classification information corresponding to that synchronization barrier can be uniquely determined, and this classification information can be written into the synchronization barrier. Therefore, embodiments of the present invention can further avoid false wake-ups.

[0080] Optionally, to prevent sleep threads from remaining in a sleep state for extended periods, in this embodiment of the invention, if the duration of a sleep thread's sleep state exceeds a preset duration, the thread scheduler can directly wake the sleep thread to allow it to continue execution. Further, optionally, this automatic wake-up operation for long-term sleep can be implemented by setting and carrying a timeout parameter in the triggered sleep command.

[0081] The processing unit of this invention includes shared memory, a thread beam scheduler, and an asynchronous acceleration unit. The shared memory is configured to update the target synchronization barrier in its internal storage based on a barrier update request, and to send thread wake-up information to the thread beam scheduler in response to detecting that the target synchronization barrier meets a release condition. The thread beam scheduler is configured to perform a thread sleep operation and trigger a first update request for the target synchronization barrier based on the arrival status of each target thread, and to perform a thread wake-up operation based on the thread wake-up information received. The asynchronous acceleration unit is configured to trigger a second update request for the target synchronization barrier based on the completion status of each target asynchronous transaction operation. This invention can achieve synchronization between threads and between threads and asynchronous transaction operations without introducing excessive hardware structures.

[0082] Figure 7 This is a schematic diagram of a thread synchronization device according to an embodiment of the present invention. Figure 7 As shown, the thread synchronization device of this embodiment includes a barrier update execution unit 71 and a wake-up message notification unit 72.

[0083] Specifically, the barrier update execution unit 71 is used to update the target synchronization barrier in the internal storage based on the barrier update request.

[0084] The wake-up message notification unit 72 is used to send thread wake-up information to the thread bundle scheduler in response to detecting that the target synchronization barrier meets the release condition, so that the thread bundle scheduler performs a thread wake-up operation based on the thread wake-up information. The target synchronization barrier is used to synchronize the execution of at least one target thread and at least one target asynchronous transaction operation. The barrier update request includes a first update request and a second update request.

[0085] Figure 8 This is a schematic diagram of a thread synchronization device according to an embodiment of the present invention. Figure 8 As shown, the thread synchronization device of this embodiment includes a barrier update triggering unit 81 and a thread wake-up unit 82.

[0086] Specifically, the barrier update triggering unit 81 is used to perform a thread sleep operation based on the arrival status of each target thread and trigger a first update request for the target synchronization barrier, so that the shared memory updates the target synchronization barrier stored internally based on the first update request, wherein the target synchronization barrier is used to synchronize the execution of at least one target thread and at least one target asynchronous transaction operation.

[0087] The thread wake-up unit 82 is used to perform a thread wake-up operation based on the received thread wake-up information in response to receiving the thread wake-up information.

[0088] The processing unit of this invention includes shared memory, a thread beam scheduler, and an asynchronous acceleration unit. The shared memory is configured to update the target synchronization barrier in its internal storage based on a barrier update request, and to send thread wake-up information to the thread beam scheduler in response to detecting that the target synchronization barrier meets a release condition. The thread beam scheduler is configured to perform a thread sleep operation and trigger a first update request for the target synchronization barrier based on the arrival status of each target thread, and to perform a thread wake-up operation based on the thread wake-up information received. The asynchronous acceleration unit is configured to trigger a second update request for the target synchronization barrier based on the completion status of each target asynchronous transaction operation. This invention can achieve synchronization between threads and between threads and asynchronous transaction operations without introducing excessive hardware structures.

[0089] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A processing unit, characterized in that, The processing unit includes: Shared memory is configured to update a target synchronization barrier in internal storage based on a barrier update request, and to send a thread wake-up message to a thread bundle scheduler in response to detecting that the target synchronization barrier meets a release condition. The target synchronization barrier is used to synchronize the execution of at least one target thread and at least one target asynchronous transaction operation. The barrier update request includes a first update request and a second update request. A thread bundle scheduler is configured to perform a thread sleep operation and trigger a first update request for the target synchronization barrier based on the arrival status of each target thread, and to perform a thread wake-up operation based on the thread wake-up information received. An asynchronous acceleration unit is configured to trigger a second update request against the target synchronization barrier based on the completion status of each of the target asynchronous transaction operations.

2. The processing unit according to claim 1, characterized in that, The synchronization barrier includes at least a thread counter and a transaction counter. The thread counter is updated in response to the first update request, and the transaction counter is updated in response to the second update request. Whether the target synchronization barrier satisfies the release condition is determined based on the count value of the thread counter and the count value of the transaction counter. The shared memory is also configured as follows: In response to detecting that the count value of the thread counter represents that the number of target threads that have arrived has reached the expected thread arrival count, and the count value of the transaction counter represents that the number of target asynchronous transaction operations that have been completed has reached the expected transaction completion count, it is confirmed that the target synchronization barrier satisfies the release condition.

3. The processing unit according to claim 2, characterized in that, The target synchronization barrier also includes the expected transaction completion count and / or the expected thread arrival count; After detecting that the target synchronization barrier meets the release condition, the shared memory is further configured as follows: Initialize the value of the thread counter using the expected thread arrival count; and / or, The value of the transaction counter is initialized using the expected transaction completion count.

4. The processing unit according to claim 1, characterized in that, The target synchronization barrier also includes a phase counter; After detecting that the target synchronization barrier meets the release condition, the shared memory is further configured as follows: Update the count value of the phase counter.

5. The processing unit according to claim 1, characterized in that, The thread wake-up information includes the thread block identifier of the target thread block corresponding to the target synchronization barrier, and the target thread block includes each of the target threads; The thread beam scheduler is specifically configured as follows: In response to receiving the thread wake-up information, each target thread in the target thread block is woken up based on the thread block identifier in the thread wake-up information.

6. The processing unit according to claim 5, characterized in that, The thread beam scheduler is specifically configured as follows: A loop instruction is issued to each thread waiting to sleep, so that each thread can read the release status of the corresponding synchronization barrier and switch itself to a sleep state or continue execution based on the release status of the corresponding synchronization barrier.

7. The processing unit according to claim 6, characterized in that, After performing the thread wake-up operation, the thread beam scheduler is further configured to: A loop instruction is issued to each awakened thread so that each awakened thread reads the release status of the corresponding synchronization barrier and switches itself to a sleep state or continues execution based on the release status of the corresponding synchronization barrier.

8. The processing unit according to claim 5, characterized in that, Each synchronization barrier within the shared memory has corresponding classification information, and the thread wake-up information also includes the target classification information of the target synchronization barrier. The thread beam scheduler is specifically configured as follows: Multiple sleeping threads of the target thread block are determined based on the thread block identifier in the thread wake-up information; Based on the target classification information, each target thread is determined from the plurality of sleep threads; Wake up each of the target threads.

9. The processing unit according to claim 1, characterized in that, The thread beam scheduler is specifically configured as follows: For each target thread, in response to detecting that the target thread has been executed to the corresponding predetermined synchronization position, the target thread is put into a sleep state and a first update request for the target synchronization barrier is triggered.

10. A thread synchronization method, characterized in that, The method is applicable to shared memory, and the method includes: Update the target synchronization barrier in internal storage based on the barrier update request; In response to the detection that the target synchronization barrier meets the release condition, a thread wake-up message is sent to the thread bundle scheduler so that the thread bundle scheduler performs a thread wake-up operation based on the thread wake-up message. The target synchronization barrier is used to synchronize the execution of at least one target thread and at least one target asynchronous transaction operation. The barrier update request includes a first update request and a second update request.

11. A thread synchronization method, characterized in that, The method is applicable to thread bundle schedulers, and the method includes: Based on the arrival status of each target thread, a thread sleep operation is performed and a first update request for the target synchronization barrier is triggered, so that the shared memory updates the target synchronization barrier stored internally based on the first update request, wherein the target synchronization barrier is used to synchronize the execution of at least one target thread and at least one target asynchronous transaction operation. In response to receiving thread wake-up information, a thread wake-up operation is performed based on the thread wake-up information.

12. A shared memory, characterized in that, The shared memory is configured to perform the method as described in claim 10.

13. A thread bundle scheduler, characterized in that, The thread beam scheduler is configured to perform the method as described in claim 11.

14. A processing system, characterized in that, The processing system includes at least one processing unit as described in any one of claims 1-9.

Citation Information

Cited By

  • Interactive methods, apparatuses, processors, devices, media, and program products for graphics processors.

    CN122312364A