Asynchronous release operation in multiprocessor system

By adopting asynchronous release operations in multiprocessor systems, the problem of thread stagnation is solved, more efficient data transmission and performance improvement is achieved, and the waiting time for memory synchronization operations is reduced.

CN120578337APending Publication Date: 2025-09-02NVIDIA CORP
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510086487.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-03-01
Filing Date
2025-01-20
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

In multiprocessor computing systems, data transmission between threads stagnates due to memory synchronization operations, which affects overall performance, especially when extended in remote memory locations, reducing the throughput of the processing unit.

Method used

The asynchronous release operation is adopted, by storing flag storage operations in the queue and merging when the memory synchronization operation is completed, threads are allowed to continue to execute after the asynchronous release operation, reducing the waiting for the memory synchronization operation.

Benefits of technology

Improves the execution performance of threads, reduces the number of memory synchronous operations, and improves the overall performance of multiprocessor systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120578337A_ABST
    Figure CN120578337A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to asynchronous release operations in a multiprocessor system. Various embodiments include techniques for performing memory synchronization operations between processors in a multi-processor computing system. The first processor transfers the data by issuing a memory operation to store the data to the shared memory. And the first processor sends an asynchronous release operation to the loading storage unit. In response, the load memory unit issues a memory synchronization operation to ensure that data associated with the memory operation is visible in the shared memory. When the asynchronous release operation is suspended, the first processor can issue further instructions and perform other operations. When data associated with the memory operation is visible in the shared memory, the memory synchronization operation is completed and the load memory unit writes a flag to a separate memory location. Once it is detected that the flag has been written, the second thread and / or other thread may reliably read data stored in the shared memory.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Various embodiments relate generally to computer system architecture, and more particularly to asynchronous release operations in multi-processor systems. Background Art

[0002] Among other things, a computing system typically includes one or more processing units, such as a central processing unit (CPU) and / or a graphics processing unit (GPU), and one or more memory systems. The processing unit executes a user mode software application that submits and starts a computing task, which is executed on one or more computing engines included in the processing unit. In operation, the processing unit loads data from one or more devices, performs various arithmetic and logical operations on the data, and stores the data back to one or more devices. The one or more devices can be memory devices of one or more memory systems and / or one or more peripheral devices, such as a network interface card (NIC), a solid-state drive (SSD), and / or the like.

[0003] In a multi-processor system, the CPU and / or GPU may have multiple processing units. Different processing units of the CPU and / or GPU, or portions thereof, may concurrently execute different threads, each of which is an instance of a program. Certain tasks involve transferring data between different threads executing on different processing units, between processing units and peripheral devices, and / or similar situations. For example, a first thread executing on a first processing unit (source) may transfer a block of data to a second thread executing on a second processing unit by storing the block of data in a memory system accessible to the second processing unit (destination). More specifically, the first thread transfers the block of data to the second thread by issuing a series of memory operations, each of which stores a portion of the data included in the block of data. When all the data in the block of data is visible to the second thread, the second thread may access the data to perform further operations.

[0004] In many cases, CPUs and / or GPUs may be processing units with a relaxed memory model, where thread instructions can execute out of order. Thus, when a thread executes, for example, ten sequential memory operations, the tenth memory operation may complete while one or more of the first nine memory operations remain outstanding. Therefore, after issuing the series of memory operations, the first thread issues a memory synchronization operation, such as a memory fence or memory barrier. The first thread performs the flag store operation by writing a flag to a memory location different from the memory location where the data block is stored, by indivisibly incrementing a flag, and / or similar operations. A memory synchronization operation is a synchronization mechanism that ensures that the series of memory operations is visible to all participating threads within a given scope (such as system-wide or processing unit-wide). The memory synchronization operation prevents the first thread from issuing further memory operations and sends a synchronization request to the memory system. The memory system flushes the memory operations issued prior to the memory synchronization operation. In response to receiving confirmation that the previous memory operation has been flushed, the first thread completes the memory synchronization operation.

[0005] After the first thread completes the memory synchronization instruction, all data in the data block becomes visible to the second thread, and the first thread can issue memory operations again. The first thread performs a flag store operation to indicate that the data in the data block is ready for use by the second thread. Simultaneously, the second thread performs a series of polling operations to determine whether the first thread has written the flag. Alternatively, the flag store operation can trigger a target operation, such as a peripheral device doorbell, a PowerPC lightweight interrupt, and / or the like, such that the flag store operation is event-driven rather than part of a poll-based operation. When the first thread performs the flag store operation, the next poll operation succeeds, and the second thread detects that the flag has been written. The second thread can then continue to use the data stored in the data block. If the second thread consumes the data before detecting the flag store, one or more memory operations may be in progress, and the second thread may read incorrect data from the data block.

[0006] One problem with this technique for transferring data between threads is that the first thread stalls after issuing a memory synchronization operation and remains stalled until the memory synchronization operation is completed. In some examples, additional threads contained in the same thread group as the first thread are also stalled until the memory synchronization operation is completed. Therefore, when the memory synchronization operation is suspended, the first thread and possibly other threads in the same thread group are unable to execute additional instructions or complete additional work. These stalls can negatively impact the memory traffic patterns of the entire computing system, thereby reducing the overall performance of the various processing units in the computing system. In addition, the memory synchronization operation can take a significant amount of time to complete. This amount of time can be particularly long when the suspended memory operation is directed to a memory location that is remotely located relative to the first processing unit that is executing the first thread. As a result, when using memory synchronization operations, the performance of the first thread can be affected and the throughput of the first processing unit can be reduced.

[0007] As noted above, what is needed in the art are more efficient techniques for synchronizing data between threads executing in a multi-processor computing system. Summary of the Invention

[0008] Various embodiments of the present disclosure describe a computer-implemented method for transferring data between threads executed in a multiprocessor computing system. The method includes receiving a first storage operation from a first processing unit, the first storage operation including a first memory synchronization operation and a first flag storage operation. The method also includes storing a first entry corresponding to the first flag storage operation in a queue. The first flag storage operation, when executed, indicates that data associated with the second storage operation generated by the first processing unit before the first storage operation is visible in memory. While the first memory synchronization operation is suspended, the first processing unit continues to execute one or more load operations or store operations.

[0009] Other embodiments include, but are not limited to, systems that implement one or more aspects of the disclosed technology, and one or more computer-readable media including instructions for performing one or more aspects of the disclosed technology, and methods for performing one or more aspects of the disclosed technology.

[0010] At least one technical advantage of the disclosed technique over the prior art is that, using the disclosed technique, a thread issuing an asynchronous release operation does not stall until a memory synchronization operation completes. Instead, after issuing an asynchronous release operation, the thread can perform further operations without waiting for the data associated with the asynchronous release operation to become visible to the receiving thread. Consequently, the execution performance of the issuing thread is improved compared to the prior art. Furthermore, because multiple asynchronous release operations can be combined into a single memory synchronization operation, fewer asynchronous release operations are issued, further improving performance. These advantages represent one or more technical improvements over prior art approaches. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to provide a detailed understanding of the implementation of the relevant features of the various embodiments described above, the inventive concepts briefly summarized above will be described in more detail with reference to various embodiments (some of which are illustrated in the accompanying drawings). However, it should be noted that the drawings illustrate only typical embodiments of the inventive concepts and are therefore not to be construed as limiting the scope in any way, and that other equally effective embodiments may exist.

[0012] Figure 1 is a block diagram of a computing system configured to implement one or more aspects of various embodiments.

[0013] Figure 2 According to various embodiments, Figure 1 A block diagram of a parallel processing unit (PPU) included in an auxiliary processing subsystem of FIG.

[0014] Figure 3 According to various embodiments, Figure 2 A block diagram of a general processing cluster (GPC) included in a parallel processing unit (PPU);

[0015] Figure 4 According to various embodiments, Figure 3 A block diagram of an asynchronous release subsystem for memory synchronization operations of a streaming multiprocessor included in a GPC;

[0016] Figure 5 According to various embodiments, Figure 4 a sequence diagram of memory synchronization operations performed by one or more streaming multiprocessors;

[0017] Figure 6 According to various embodiments, Figure 4 a sequence diagram of asynchronous release operations performed by one or more streaming multiprocessors; and

[0018] Figure 7 According to various embodiments, Figure 4A flowchart of method steps for performing an asynchronous release operation on a streaming multiprocessor. DETAILED DESCRIPTION

[0019] In the following description, numerous specific details are set forth in order to provide a more thorough understanding of various embodiments. However, it will be apparent to one skilled in the art that the present invention can be practiced without one or more of these specific details.

[0020] System Overview

[0021] Figure 1 1 is a block diagram illustrating a computing system 100 configured to implement one or more aspects of various embodiments. As shown, computing system 100 includes, but is not limited to, a central processing unit (CPU) 102, a system memory 104, which is coupled to an auxiliary processing subsystem 112 via a memory bridge 105 and a communication path 113. Memory bridge 105 is further coupled to an I / O (input / output) bridge 107 via a communication path 106, which is in turn coupled to a switch 116.

[0022] In operation, I / O bridge 107 is configured to receive user input information from input device 108 (such as a keyboard or mouse) and forward the input information to CPU 102 via communication path 106 and memory bridge 105 for processing. In some examples, input device 108 is used to verify the identity of one or more users so as to allow authorized users to access computing system 100 and deny unauthorized users access to computing system 100. Switch 116 is configured to provide connections between I / O bridge 107 and other components of computing system 100, such as network adapter 118 and various add-in cards 120 and 121. In some examples, network adapter 118 serves as a primary input device or a dedicated input device for receiving input data for processing by the disclosed technology.

[0023] As also shown, the I / O bridge 107 is coupled to a system disk 114, which can be configured to store content, applications, and data for use by the CPU 102 and the auxiliary processing subsystems 112. Generally, the system disk 114 provides non-transitory storage for applications and data and can include a fixed or removable hard drive, a flash memory device, and a CD-ROM (Compact Disc Read Only Memory), DVD-ROM (Digital Versatile Disk-ROM), Blu-ray, HD-DVD (High Definition DVD), or other magnetic, optical, or solid-state storage device. Finally, although not explicitly shown, other components (such as a universal serial bus or other port connection, an optical disc drive, a digital versatile disk drive, a film recording device, etc.) can also be connected to the I / O bridge 107.

[0024] In various embodiments, memory bridge 105 may be a northbridge chip, and I / O bridge 107 may be a southbridge chip. Furthermore, communication paths 106 and 113, as well as other communication paths within computing system 100, may be implemented using any technically suitable protocol, including but not limited to Peripheral Component Interconnect Express (PCIe), HyperTransport, or any other bus or point-to-point communication protocol known in the art.

[0025] In some embodiments, auxiliary processing subsystem 112 includes a graphics subsystem that delivers pixels to display device 110, which may be any conventional cathode ray tube, liquid crystal display, light emitting diode display, etc. In such embodiments, auxiliary processing subsystem 112 incorporates circuitry optimized for graphics and video processing, including, for example, video output circuitry. Figure 2 As described in more detail in , such circuitry may be combined across one or more auxiliary processors included in the auxiliary processing subsystem 112. An auxiliary processor includes any one or more processing units that can execute instructions, such as a central processing unit (CPU), Figures 2 to 4 Parallel processing unit (PPU), graphics processing unit (GPU), direct memory access (DMA) unit, intelligent processing unit (IPU), neural processing unit (NPU), tensor processing unit (TPU), neural network processor (NNP), data processing unit (DPU), vision processing unit (VPU), application-specific integrated circuit (ASIC), field programmable gate array (FPGA), etc.

[0026] In some embodiments, the auxiliary processing subsystem 112 incorporates circuitry optimized for general-purpose and / or computational processing. Such circuitry can be combined across one or more auxiliary processors included in the auxiliary processing subsystem 112, which are configured to perform such general-purpose and / or computational operations. In other embodiments, one or more auxiliary processors included in the auxiliary processing subsystem 112 can be configured to perform graphics processing, general-purpose processing, and computational processing operations. System memory 104 includes at least one device driver 103 configured to manage processing operations of one or more auxiliary processors in the auxiliary processing subsystem 112.

[0027] In various embodiments, the auxiliary processing subsystem 112 may be coupled to Figure 1 For example, the auxiliary processing subsystem 112 may be integrated with the CPU 102 and / or other components and connecting circuits on a single chip to form a system on a chip (SoC).

[0028] It should be understood that the system shown herein is illustrative and that variations and modifications are possible. The connection topology (including the number and arrangement of bridges, the number of CPUs 102, and the number of auxiliary processing subsystems 112) can be modified as needed. For example, in some embodiments, the system memory 104 can be directly connected to the CPU 102, rather than being connected to the CPU 102 through the memory bridge 105, and other devices will communicate with the system memory 104 via the memory bridge 105 and the CPU 102. In other alternative topologies, the auxiliary processing subsystem 112 can be connected to the I / O bridge 107 or directly to the CPU 102, rather than being connected to the memory bridge 105. In still other embodiments, the I / O bridge 107 and the memory bridge 105 can be integrated into a single chip, rather than existing as one or more discrete devices. Finally, in some embodiments, Figure 1 One or more of the components shown may not be present. For example, switch 116 may be eliminated, and network adapter 118 and add-in cards 120 , 121 may be connected directly to I / O bridge 107 .

[0029] Figure 2 According to various embodiments, Figure 1 A block diagram of a parallel processing unit (PPU) 202 included in the auxiliary processing subsystem 112 of FIG. Figure 2 While one PPU 202 is depicted, as described above, the auxiliary processing subsystem 112 may include any number of PPUs 202. Figure 2 PPU 202 is Figure 1 is one example of an auxiliary processor included in the auxiliary processing subsystem 112. Alternative auxiliary processors include, but are not limited to, a CPU, a GPU, a DMA unit, an IPU, an NPU, a TPU, an NNP, a DPU, a VPU, an ASIC, an FPGA, and / or the like. Figures 2 to 4 The techniques disclosed in

[0045] regarding PPU 202 are equally applicable to any type of auxiliary processor included in auxiliary processing subsystem 112, in any combination. As shown, PPU 202 is coupled to local parallel processing (PP) memory 204. PPU 202 and PP memory 204 may be implemented using one or more integrated circuit devices, such as programmable processors, application-specific integrated circuits (ASICs), or memory devices, or in any other technically feasible manner.

[0030] In some embodiments, PPU 202 includes a graphics processing unit ("GPU") that can be configured to implement a graphics rendering pipeline to perform various operations related to generating pixel data based on graphics data provided by CPU 102 and / or system memory 104. When processing graphics data, PP memory 204 can be used as graphics memory to store one or more conventional frame buffers (and, if desired, one or more other rendering targets). PP memory 204 can be used to, among other things, store and update pixel data and transmit the final pixel data, or display frame, to display device 110 for display. In some embodiments, PPU 202 can also be configured for general processing and computational operations.

[0031] In operation, CPU 102 is the main processor of computing system 100, controlling and coordinating the operations of other system components. In particular, CPU 102 issues commands that control the operation of PPU 202. In some embodiments, CPU 102 writes a command stream for PPU 202 into a data structure ( Figure 1 or Figure 2 The data structure may be located in system memory 104, PP memory 204, or another storage location accessible to both CPU 102 and PPU 202 (not explicitly shown). Additionally or alternatively, a processor other than CPU 102 and / or an auxiliary processor may write one or more command streams from PPU 202 to the data structure. A pointer to the data structure is written to a push buffer to initiate processing of the command stream in the data structure. PPU 202 reads the command stream from the push buffer and then executes the commands asynchronously with respect to the operation of CPU 102. In embodiments where multiple push buffers are generated, the application may specify an execution priority for each push buffer via device driver 103 to control the scheduling of the different push buffers.

[0032] As also shown, PPU 202 includes an I / O (input / output) unit 205, which communicates with the rest of computing system 100 via communication path 113 and memory bridge 105. I / O unit 205 generates data packets (or other signals) for transmission on communication path 113, and also receives all incoming data packets (or other signals) from communication path 113, directing the incoming data packets to the corresponding components of PPU 202. For example, commands related to processing tasks may be directed to host interface 206, while commands related to memory operations (e.g., reading from or writing to PP memory 204) may be directed to crossbar unit 210. Host interface 206 reads each push buffer and sends the command stream stored in the push buffer to front end 212.

[0033] As above combined Figure 1 As described above, the connection between PPU 202 and the rest of computing system 100 can vary. In some embodiments, auxiliary processing subsystem 112 (which includes at least one PPU 202) is implemented as an add-in card that can be inserted into an expansion slot of computing system 100. In other embodiments, PPU 202 can be integrated on a single chip using a bus bridge, such as memory bridge 105 or I / O bridge 107. Likewise, in other embodiments, some or all elements of PPU 202 can be included with CPU 102 in a single integrated circuit or system on a chip (SoC).

[0034] In operation, the front end 212 sends processing tasks received from the host interface 206 to a work distribution unit (not shown) within the task / work unit 207. The work distribution unit receives pointers to processing tasks, which are encoded as task metadata (TMD) and stored in memory. The pointer to the TMD is included in a command stream, which is stored as a push buffer and received by the front end 212 from the host interface 206. The processing tasks that can be encoded as TMDs include indexes associated with the data to be processed, as well as state parameters and commands that define how to process the data. For example, the state parameters and commands can define the program to be executed on the data. The task / work unit 207 receives tasks from the front end 212 and ensures that the GPC 208 is configured to a valid state before initiating the processing tasks specified by each TMD. A priority can be assigned to each TMD, which is used to schedule the execution of the processing tasks. Processing tasks can also be received from the processing cluster array 230. Optionally, the TMD may include a parameter that controls whether the TMD is added to the head or tail of a list of processing tasks (or a list of pointers to processing tasks), thereby providing another level of control over execution priority.

[0035] PPU 202 advantageously implements a highly parallel processing architecture based on a processing cluster array 230, which includes a set of C general processing clusters (GPCs) 208, where C ≥ 1. Each GPC 208 is capable of executing a large number (e.g., hundreds or thousands) of threads simultaneously, where each thread is an instance of a program. In various applications, different GPCs 208 can be assigned to process different types of programs or perform different types of computations. The assignment of GPCs 208 can vary depending on the workload generated by each type of program or computation.

[0036] The memory interface 214 includes a set of D partition units 215, where D ≥ 1. Each partition unit 215 is coupled to one or more dynamic random access memories (DRAMs) 220 residing within the PP memory 204. In one embodiment, the number of partition units 215 is equal to the number of DRAMs 220, with each partition unit 215 coupled to a different DRAM 220. In other embodiments, the number of partition units 215 may differ from the number of DRAMs 220. Those skilled in the art will recognize that the DRAMs 220 may be replaced with any other technically suitable memory device. In operation, various render targets (such as texture maps and frame buffers) may be stored across the DRAMs 220, allowing the partition units 215 to write portions of each render target in parallel, thereby efficiently utilizing the available bandwidth of the PP memory 204.

[0037] A given GPC 208 can process data to be written to any DRAM 220 in PP memory 204. The crossbar unit 210 is configured to route the output of each GPC 208 to the input of any partition unit 215 or any other GPC 208 for further processing. The GPCs 208 communicate with the memory interface 214 via the crossbar unit 210 to read from or write to the various DRAMs 220. In one embodiment, the crossbar unit 210 is connected to the I / O unit 205 in addition to being connected to the PP memory 204 via the memory interface 214, thereby enabling processing cores in different GPCs 208 to communicate with the system memory 104 or other memory that is not local to the PPU 202. Figure 2 In an embodiment, crossbar unit 210 is directly connected to I / O unit 205. In various embodiments, crossbar unit 210 can separate traffic flows between GPCs 208 and partition units 215 using virtual channels.

[0038] Likewise, the GPCs 208 can be programmed to perform processing tasks associated with various applications, including, but not limited to, linear and nonlinear data transformations, filtering of video and / or audio data, modeling operations (e.g., applying physical laws to determine the position, velocity, and other properties of an object), image rendering operations (e.g., tessellation shading programs, vertex shading programs, geometry shading programs, and / or pixel / fragment shading programs), general-purpose compute operations, etc. In operation, the PPU 202 is configured to transfer data from the system memory 104 and / or the PP memory 204 to one or more on-chip memory units, process the data, and write the resulting data back to the system memory 104 and / or the PP memory 204. The resulting data can then be accessed by other system components (including the CPU 102, another PPU 202 in the auxiliary processing subsystem 112, or another auxiliary processing subsystem 112 in the computing system 100).

[0039] As described above, any number of PPUs 202 may be included in the auxiliary processing subsystem 112. For example, multiple PPUs 202 may be provided on a single add-in card, multiple add-in cards may be connected to the communication path 113, or one or more PPUs 202 may be integrated into a bridge chip. The PPUs 202 in a multi-PPU system may be identical to or different from one another. For example, different PPUs 202 may have different numbers of processing cores and / or different amounts of PP memories 204. In implementations where multiple PPUs 202 are present, these PPUs may operate in parallel to process data at a higher throughput than would be possible with a single PPU 202. Systems including one or more PPUs 202 may be implemented in a variety of configurations and form factors, including but not limited to desktops, laptops, handheld personal computers or other handheld devices, servers, workstations, gaming consoles, embedded systems, and the like.

[0040] Figure 3 According to various embodiments, Figure 2208 . In operation, the GPC 208 can be configured to execute a large number of threads in parallel to perform graphics, general processing and / or computing operations. As used herein, a "thread" refers to an instance of a specific program executed on a specific input data set. In some embodiments, a single instruction multiple data (SIMD) instruction issuance technology is used to support the parallel execution of a large number of threads without providing multiple independent instruction units. In other embodiments, a single instruction multiple thread (SIMT) technology is used to support the parallel execution of a large number of generally synchronized threads using a general instruction unit configured to issue instructions to a group of processing engines in the GPC 208. Unlike a SIMD execution architecture (in which all processing engines generally execute the same instruction), SIMT execution allows different threads to more easily follow different execution paths through a given program. It should be appreciated by those skilled in the art that a SIMD processing architecture represents a functional subset of a SIMT processing architecture.

[0041] The operation of GPC 208 is controlled via pipeline manager 305, which distributes processing tasks received from work distribution units (not shown) within task / work units 207 to one or more streaming multiprocessors (SMs) 310. Pipeline manager 305 can also be configured to control work distribution crossbar 330 by specifying the destination of processed data output by SM 310.

[0042] In one embodiment, the GPC 208 includes a set of M SMs 310, where M ≥ 1. In addition, each SM 310 includes a set of function execution units (not shown), such as execution units and load-store units. The processing operations specific to any function execution unit can be pipelined, allowing new instructions to be issued for execution before previous instructions have completed execution. Any combination of function execution units in a given SM 310 can be provided. In various embodiments, the function execution units can be configured to support a variety of different operations, including integer and floating-point arithmetic (e.g., addition and multiplication), comparison operations, Boolean operations (e.g., AND, OR, XOR), bit shifts, and calculations of various algebraic functions (e.g., planar interpolation and trigonometric functions, exponential functions, logarithmic functions, etc.). Advantageously, the same function execution unit can be configured to perform different operations.

[0043] In operation, each SM 310 is configured to process one or more thread groups. As used herein, a "thread group" or "warp" refers to a group of threads that concurrently execute the same program on different input data, with a thread in the group assigned to a different execution unit in the SM 310. A thread group may include fewer threads than the number of execution units in the SM 310, in which case some execution units may be idle during cycles while the thread group is being processed. A thread group may also include more threads than the number of execution units in the SM 310, in which case processing may occur in consecutive clock cycles. Since each SM 310 can support up to G thread groups simultaneously, up to G*M thread groups may be executing in a GPC 208 at any given time.

[0044] In addition, multiple related thread groups can be active (in different stages of execution) simultaneously in an SM 310. This collection of thread groups is referred to herein as a "cooperative thread array" ("CTA") or "thread array." The size of a particular CTA is equal to m*k, where k is the number of concurrently executing threads in the thread group, which is typically an integer multiple of the number of execution units in the SM 310, and m is the number of concurrently active thread groups within the SM 310. In various embodiments, software applications written in the Compute Unified Device Architecture (CUDA) programming language describe the behavior and operations of threads executing on the GPC 208, including any of the behaviors and operations described above. A given processing task can be specified in a CUDA program so that the SM 310 can be configured to perform and / or manage general-purpose computing operations.

[0045] although Figure 3 Not shown, but each SM 310 includes a level 1 (L1) cache, or uses space in a corresponding L1 cache external to the SM 310 to support load and store operations, etc., performed by the execution units. Each SM 310 also has access to a level 2 (L2) cache (not shown) that is shared between all GPCs 208 in the PPU 202. The L2 cache can be used to transfer data between threads. Finally, the SM 310 also has access to off-chip "global" memory, which can include PP memory 204 and / or system memory 104. It should be understood that any memory external to the PPU 202 can be used as global memory. In addition, as Figure 3As shown, a level 1.5 (L1.5) cache 335 may be included in GPC 208 and is configured to receive and store data requested from memory by SM 310 via memory interface 214. Such data may include, but is not limited to, instructions, uniform data, and constant data. In embodiments having multiple SMs 310 within GPC 208, SMs 310 may advantageously share common instructions and data cached in L1.5 cache 335.

[0046] Each GPC 208 may have an associated memory management unit (MMU) 320 that is configured to map virtual addresses to physical addresses. In various embodiments, the MMU 320 may reside within the GPC 208 or the memory interface 214. The MMU 320 includes a set of page table entries (PTEs) that map virtual addresses to physical addresses of tiles or memory pages, and optionally to cache line indices. The MMU 320 may include a translation lookaside buffer (TLB) or a cache that resides within the SM 310, one or more L1 caches, or the GPC 208.

[0047] In graphics and compute applications, GPC 208 may be configured so that each SM 310 is coupled to a texture unit 315 to perform texture mapping operations, such as determining texture sample locations, reading texture data, and filtering texture data.

[0048] In operation, each SM 310 sends processed tasks to the work distribution crossbar 330 so that the processed tasks can be provided to another GPC 208 for further processing, or the processed tasks can be stored in an L2 cache (not shown), parallel processing memory 204, or in system memory 104 via the crossbar unit 210. In addition, a pre-raster operation ("preROP") unit 325 is configured to receive data from the SM 310, direct the data to one or more raster operation (ROP) units in the partition unit 215, perform optimizations for color blending, organize pixel color data, and perform address translation.

[0049] It should be understood that the core architecture described herein is illustrative and that variations and modifications are possible. Among other things, any number of processing units, such as SM 310, texture unit 315, or preROP unit 325, may be included in GPC 208. Figure 2As described above, the PPU 202 may include any number of GPCs 208 that are configured to be functionally similar to one another so that execution behavior does not depend on which GPC 208 receives a particular processing task. Furthermore, each GPC 208 operates independently of the other GPCs 208 in the PPU 202 to execute tasks of one or more application programs. In view of the foregoing, it will be understood by those skilled in the art that Figure 1-Figure 3 The architecture described in this document in no way limits the scope of the various embodiments of the present disclosure.

[0050] Note that, as used herein, reference to shared memory may include any one or more technically feasible memories, including but not limited to local memory shared by one or more SMs 310 or memory accessible via memory interface 214, such as cache memory, parallel processing memory 204, or system memory 104. Note also that, as used herein, reference to cache memory may include any one or more technically feasible memories, including but not limited to L1 cache, L1.5 cache, and L2 cache.

[0051] Transferring data between threads via asynchronous release operations

[0052] Various embodiments include techniques for synchronizing data between threads executing in a multiprocessor computing system. Using the disclosed techniques, a first thread transfers data to a second thread by writing the data to memory using a series of memory operations and performing an asynchronous release operation. The asynchronous release operation includes a store operation with a built-in memory synchronization operation, referred to herein as a flag store operation or flag operation, which does not stall the issuing thread during the memory synchronization operation. In various embodiments, the flag store operation can be a standard memory write operation. Additionally or alternatively, the flag store operation can be an indivisible memory operation, such as an indivisible increment operation, an indivisible logical OR operation, or the like. A load-store unit receives the asynchronous release operation and enqueues the flag store operation into a side structure (referred to herein as a release queue, or more simply, a queue), allowing the issuing thread to continue executing normally. Memory barrier logic performs the memory synchronization operation. When the memory synchronization operation is complete, the memory barrier logic signals the completion of the memory synchronization operation to the release queue. At this point, the earlier memory operation is visible to the receiving thread, and the release queue can perform the flag store operation.

[0053] In some examples, a release queue can store multiple asynchronous release operations, each of which is waiting for a related earlier memory operation to become visible. When coupled with a synchronization unit capable of merging or combining multiple memory synchronization operations, a single memory synchronization operation is sufficient to meet the visibility requirements of the merged multiple pending asynchronous release operations. Therefore, the asynchronous release operation can further improve performance by minimizing the number of barrier operations sent to the memory subsystem. In some examples, the storage operations, memory synchronization operations, and / or flag storage operations described herein can be issued by processor instructions executed by a processing unit.

[0054] Figure 4 According to various embodiments, Figure 3 2 is a block diagram of an asynchronous release subsystem 400 for memory synchronization operations of a streaming multiprocessor 310 included in a GPC 208 of FIG. 1 . The asynchronous release subsystem 400 is also referred to herein as a memory synchronization subsystem. The asynchronous release subsystem 400 may be included within the GPC 208, external to the GPC 208, or partially within the GPC 208 and partially external to the GPC 208. As shown, the asynchronous release subsystem 400 includes, but is not limited to, an address generation unit (AGU) 410, a level 1 (L1) cache tag memory 412, an L1 cache miss generator 414, a memory barrier (membar) merger 420, a release memory barrier merger 422, a release data merger 424, a release queue 430, and a drain unit 440. The address generation unit 410, the L1 cache tag memory 412, and the L1 cache miss generator 414 are at least part of a load store unit (LOU). Figure 4 More generally, Figure 4 The components shown illustrate portions of a load-store unit that implements the asynchronous release operations disclosed herein. The load-store unit is coupled to one or more memory subsystems via interconnects. In some embodiments, these memory subsystems may include one or more crossbar units, such as a crossbar unit 210, an MMU 320, an L2 cache memory, a frame buffer, a memory controller, an arbiter, and / or the like. The memory subsystem routes and processes received memory requests until the memory requests reach their final destination, such as system memory 104, PP memory, a network interface card (NIC), and / or the like.

[0055] In operation, the address generation unit 410 receives a store operation from the corresponding SM. In some embodiments, each SM 310 and / or each load-store unit includes an address generation unit 410. The store operation can be a standard store instruction, a reduction instruction, an indivisible read-modify-write (RMW) instruction, etc. The address generation unit 410 is a virtual address generation and checking unit. The address generation unit 410 receives a set of inputs from one or more threads executing on the SM 310. These inputs may include data from different registers, immediate values ​​from executed instructions, state variables, and / or the like. The address generation unit 410 processes these inputs, such as by addition, multiplication, and / or the like, to generate an address. In addition, the address generation unit 410 performs misalignment testing, error checking, and / or the like. The address generation unit 410 can also detect whether the calculated address is within the memory range of various apertures, such as the shared memory aperture, the global memory aperture, and the local memory aperture. During a synchronization operation, a memory synchronization operation may pass through the address generation unit 410, while a corresponding store operation of a release operation includes an address processed by the address generation unit 410. The address generation unit 410 sends the store operation with the generated memory address to the L1 cache tag memory 412.

[0056] L1 cache tag memory 412 determines whether the data of the generated memory address is stored in the L1 cache. If L1 cache tag memory 412 includes a cache tag entry for the generated memory address, the data accessed by the store operation resides in the L1 cache memory. If L1 cache tag memory 412 does include a cache tag entry for the generated memory address, the data accessed by the store operation does not reside in the L1 cache memory. This latter situation is referred to as an L1 cache miss. In this case, the data included in the store operation is stored in a memory other than the L1 cache memory, such as L2 cache memory, local memory, system memory, local memory of another processing unit, and / or the like. In this case, L1 cache tag memory 412 detects an L1 cache miss and sends the store operation to L1 cache miss generator 414.

[0057] The L1 cache miss generator 414 receives a store operation from the L1 cache tag memory 412. In the event that the L1 cache miss is detected by the L1 cache tag memory 412, the L1 cache miss generator 414 generates a memory request to store data in memory and stores the data in one or more cache lines of the L1 cache tag memory 412. The L1 cache tag memory 412 prepares the store operation to be stored in the address space identified by the L1 cache tag memory 412. The L1 cache miss generator 414 sends the prepared store operation to the memory subsystem (not shown) via the on-chip network 450 and / or other suitable interface. The on-chip network 450 is an interconnect, bus, and / or other communication interface that connects the SMs 310 to each other and to other on-chip components, such as L2 cache memory, local memory, system memory, local memory of another processing unit, and / or the like. In some embodiments, the on-chip network 450 communicates with the network adapter 118 to connect the computing system 100 to other computing systems in a larger computer network. In such embodiments, the destination of flag store operations and / or other storage operations can be the network adapter 118. More generally, flag store operations through the on-chip network 450 can be directed to any storage location in the computer system 100 and / or in a different computer system.

[0058] In addition to standard store operations, the address generation unit 410 may also receive memory synchronization operations and / or store release operations from the SM 310. If the SM 310 sends a memory synchronization operation to the address generation unit 410 and subsequently issues a load operation or a store operation, the SM 310 stalls, waiting for the completion of the memory synchronization operation. If the address generation unit 410 receives a memory synchronization operation from the SM 310, the address generation unit 410 and the L1 cache tag memory 412 pass the memory synchronization operation until the memory synchronization operation reaches the memory barrier merger 420. When the L1 cache miss generator 414 receives a memory synchronization operation from the L1 cache tag memory 412, the L1 cache tag memory 412 sends the memory synchronization operation to the memory barrier merger 420. The memory barrier merger 420 merges the incoming memory barrier operation with other memory barrier operations received shortly before and / or shortly after the incoming memory barrier operation.

[0059] In some examples, to perform these merge operations, the memory barrier merger 420 initializes a counter to a first predefined value. The memory barrier merger 420 periodically increments or decrements the counter to a second predefined value, where the difference between the first predefined value and the second predefined value represents a duration. In some examples, the memory barrier merger 420 may initialize the counter to a maximum counter value and may decrement the counter until the counter reaches zero. In some examples, the memory barrier merger 420 may initialize the counter to zero and may increment the counter until the counter reaches the maximum counter value. While the counter is still counting and has not yet reached the second predefined value, the memory barrier merger 420 merges the memory synchronization operations. When the counter reaches the second predefined value, the merging of the memory synchronization operations is complete. Other components of the memory subsystem external to the load-store unit perform one or more additional portions of the memory synchronization operations. Subsequently, the one or more other components of the memory subsystem send a signal to the load-store unit indicating that these additional portions of the memory synchronization operations have completed. When the memory synchronization operation is complete, the memory subsystem unstalls the pipeline in the load-store unit, allowing the AGU 410 to receive the next memory request. The memory request may be a flag store operation generated by the corresponding SM 310. The load-store unit processes the flag store operation as a standard store instruction received from the SM 310. For example, as described herein, the flag store operation proceeds through the AGU 410, the L1 cache tag memory 412, the L1 cache miss generator 414, and the on-chip network 450.

[0060] If the SM 310 sends a store-release operation to the address generation unit 410, the SM 310 does not stall while the store-release operation is pending, even if the SM 310 subsequently issues a load or store operation. Store-release operations include store operations such as standard store instructions, compact instructions, indivisible read-modify-write (RMW) instructions, and / or the like. Store-release operations also include embedded memory barrier operations and flag store operations that do not stall the issuing SM thread group 610. If the address generation unit 410 receives a store-release operation from the SM 310, the address generation unit 410 processes the store-release operation similarly to a standard store operation. When the L1 cache tag memory 412 receives the store-release operation from the address generation unit 410, the L1 cache tag memory 412 splits the store-release operation into a memory barrier operation and a flag store operation. The L1 cache tag memory 412 sends the memory barrier operation to the release memory barrier merger 422 and sends the flag store operation to the release data merger 424.

[0061] The release memory barrier merger 422 and the release data merger 424 merge the incoming store release operation with other store release operations received shortly before and / or shortly after the incoming store release operation. In some examples, the release data merger 424 merges pending flag store operations to the same memory address into a single flag store operation. In doing so, the release data merger 424 attempts to identify other pending flag store operations to the same memory address and merges these flag store operations to the same memory address into a single flag store operation. Additionally or alternatively, the release data merger 424 merges pending flag store operations to adjacent (e.g., immediately subsequent and / or immediately previous) memory addresses into a single request. In doing so, the release data merger 424 attempts to identify other pending flag store operations to adjacent memory addresses and merges these flag store operations into a single flag store operation that represents a larger block of data than each of the original flag store operations.

[0062] In some examples, to perform a merge operation, the release memory barrier merger 422 initializes a counter to a first predefined value. The release memory barrier merger 422 periodically increments or decrements the counter to a second predefined value, where the difference between the first predefined value and the second predefined value represents a duration. In some examples, the release memory barrier merger 422 may initialize the counter to a maximum counter value and may decrement the counter until the counter reaches zero. In some examples, the release memory barrier merger 422 may initialize the counter to zero and may increment the counter until the counter reaches the maximum counter value. While the counter is still counting and has not yet reached the second predefined value, the release memory barrier merger 422 merges the membar operation. When the counter reaches the second predefined value, the release memory barrier merger 422 stops merging the membar operation and forwards the membar operation to the main memory barrier merger 420.

[0063] At the same time, a release data merger 424 merges data associated with the store release operations being merged by the release memory barrier merger. If a particular store release operation is not merged with any other release store operations, the release data merger 424 generates metadata 432, a virtual address 434, and / or data 436 associated with a marker store operation that releases the store release operation. The release data merger 424 generates the metadata 432, virtual address 434, and data 436 based on the data included in the original store release operation, the address generated from the address generation unit 410, and / or the like. The release data merger 424 stores the metadata 432, virtual address 434, and data 436 marking the store operation in an entry of the release queue 430.

[0064] If the release data merger 424 determines that the incoming store release operation is directed to the same memory address as another store release operation, the release data merger 424 can merge the two corresponding flag store operations. The release data merger 424 generates metadata 432, a virtual address 434, and data 436 for the flag store operation that releases the two store release operations. The release data merger 424 generates the metadata 432, virtual address 434, and data 436 based on the data included in the two store release operations, the addresses generated from the address generation unit 410, and / or the like. The release data merger 424 stores the metadata 432, virtual address 434, and data 436 for the two flag store operations in a single entry in the release queue 430. In some embodiments, the release data merger 424 merges such pending flag store operations by replacing the first entry in the release queue 430 with the second entry.

[0065] If the release data merger 424 determines that the incoming store-release operation is directed to a memory address that is adjacent to the memory address to which another store-release operation is directed, the release data merger 424 may merge the two corresponding tagged store operations into a single tagged-release operation that covers the combined memory address range of the two store-release operations. The release data merger 424 generates metadata 432, a virtual address 434, and data 436 for the tagged store operation that releases the two store-release operations. The release data merger 424 generates the metadata 432, virtual address 434, and data 436 based on data included in the two store-release operations, addresses generated from the address generation unit 410, and / or the like. The release data merger 424 stores the metadata 432, virtual address 434, and data 436 for the two tagged store operations in a single entry in the release queue 430. In some embodiments, the release data merger 424 merges such pending flag store operations by replacing a first entry corresponding to the first portion of the combined memory address range in the release queue 430 with a second entry corresponding to the first portion of the combined memory address range and the second portion of the combined memory address range. In some embodiments, the first portion of the combined memory address range may overlap with the second portion of the combined memory address range.

[0066] In some examples, to perform the merge operation, the release memory barrier merger 422 initializes a counter to a first predefined value. The release memory barrier merger 422 periodically increments or decrements the counter to a second predefined value, where the difference between the first predefined value and the second predefined value represents a duration. In some examples, the release memory barrier merger 422 may initialize the counter to a maximum counter value and may decrement the counter until the counter reaches zero. In some examples, the release memory barrier merger 422 may initialize the counter to zero and may increment the counter until the counter reaches the maximum counter value. When the counter is still counting and has not yet reached the second predefined value, the release memory barrier merger 422 merges the flag store operation. When the counter reaches the second predefined value, the release memory barrier merger 422 stops merging the flag store operation and forwards the flag store operation to the main memory barrier merger 420.

[0067] Subsequently, when one or more pending store release operations are completed, memory barrier merger 420 receives a memory barrier confirmation indicating that the data corresponding to the one or more pending store release operations is now visible in the associated memory address space. Memory barrier merger 420 sends the memory barrier confirmation to release memory barrier merger 422. In response, release memory barrier merger 422 sends a signal to flush unit 440 indicating that the corresponding entry of the flagged store operation can be processed or flushed from the release queue.

[0068] The flush unit 440 begins processing entries of the flag store operations in the release queue 430, thereby flushing the flag store operation release queue 430. As each entry is flushed from the release queue 430, the release memory barrier merger 422 forwards the corresponding flag store operation to the L1 cache miss generator 414. The L1 cache miss generator 414 then processes the flag store operation.

[0069] The release memory barrier merger 422, the release data merger 424, and / or other components of the asynchronous release subsystem 400 switch the release queue 430 to a different memory barrier stage. While the flush unit 440 is flushing the flag store operations from the current buffer of the release queue 430, the release memory barrier merger 422 continues to collect further flag store operations from one or more other back buffers of the release queue 430. Each buffer of the release queue 430 corresponds to a different memory barrier stage of the release queue 430. When the write store operations of the current buffer are flushed for the current memory barrier stage, the asynchronous release subsystem 400 switches the memory barrier stage of the release queue 430, causing one of the back buffers to become the current buffer and the previously current buffer to become the back buffer. In the new memory barrier stage, the release memory barrier merger 422 flushes the flag store operations of the new current buffer, while the previously current buffer collects additional flag store operations for the subsequent memory barrier stage of the release queue 430.

[0070] Due to the asynchronous nature of the asynchronous release operation mechanism, the issue thread group executing on SM 310 can generate query operations to the asynchronous release subsystem 400 to determine the status of outstanding asynchronous release operations. In some examples, the issue thread group executing on SM 310 can perform these query operations to synchronize the current asynchronous release operation flow with the subsequent memory synchronization operation flow (whether asynchronous or not) and query the completion status of the asynchronous release operation. To this end, existing instructions of the instruction set architecture (ISA) can be extended to implement the query operations. Alternatively, new instructions can be added to the ISA to implement the query operations. When one or more threads in the issue thread group execute the extended or new instructions, the thread group waits for the flag store operation associated with the issue thread group to complete and be cleared from the release queue 430, as well as for the same address sort to be stored. At this time, the load-store unit sends a status message to the issue thread group on SM 310 that generated the query operation. The status message indicates that the asynchronous release operation is complete. The issue thread group can then use existing operations to ensure the completion of the flag store operation (i.e., data visibility).

[0071] It should be understood that the system shown herein is illustrative and that variations and modifications are possible. The technology described herein is in Figure 3 The techniques described herein are performed in the context of one or more SMs 310 included in a GPC 208. Additionally or alternatively, the techniques described herein may be performed by one or more alternative processors including, but not limited to, a CPU, a GPU, a DMA unit, an IPU, an NPU, a TPU, an NNP, a DPU, a VPU, an ASIC, an FPGA, and / or any combination thereof. More generally, the techniques described herein may be applied to any CPU 102, a PPU 202, and / or any other processing unit in any combination.

[0072] Figure 5 According to various embodiments, Figure 43 shows a sequence diagram of memory synchronization operations performed by one or more streaming multiprocessors 310. As shown, an SM thread group 510 executing on an SM 310 generates two store operations 520 and 522. A load-store unit 512 included in the memory subsystem of the SM 310 receives the two store operations annotated as store operations 530 and 532. The load-store unit 512 stores the data included in the store operations annotated as store operations 540 and 542 in memory 514. As shown, store operation 540 completes before store operation 542. Because the SM 310 is part of a relaxed memory model processing unit, store operation 542 can complete before store operation 540. To ensure that both store operations 540 and 542 have completed, the SM thread group 510 performs a memory synchronization operation, or memory barrier (membar) operation 526, and sends the memory barrier operation 526 to the load-store unit 512. The load-store unit 512 detects the memory barrier operation annotated as memory barrier operation 536. While the memory barrier operation annotated as memory barrier operation 546 is pending, SM thread group 510 cannot perform further load operations or store operations, such as load operations or store operations directed to memory 514. In some examples, SM thread group 510 can continue to perform operations other than load operations or store operations. However, if SM thread group 510 attempts to perform a load operation or store operation after performing memory barrier operation 526, SM thread group 510 is stalled until memory barrier operation 526 is completed.

[0073] SM thread group 510 remains stalled until store operations 540 and 542, as well as any other pending store operations executed by the SM thread group before memory barrier operation 526, become visible in memory 514. When the data included in these store operations becomes visible in memory 514, memory 514 issues a memory barrier acknowledgment (membar ack) 550. Memory 514 sends memory barrier acknowledgment 550 to SM thread group 510 via load-store unit 512. When SM thread group 510 receives memory barrier acknowledgment 550, memory barrier operation 526 is complete, and SM thread group 510 resumes execution. SM thread group 510 executes store operation 552 to write a flag to a specified location, such as a location in memory 514. Upon detecting that the flag has been written, threads included in SM thread group 510 and / or threads included in other SM thread groups can reliably read data from memory 514, such as data stored in memory 514 by store operations 540 and 542.

[0074] As described herein, SM thread group 510 is stalled while memory barrier operation 526 is pending. Furthermore, one or more other SM thread groups may be exchanging data with, communicating with, and / or the like, the SM thread group 510. Because SM thread group 510 is stalled, these other thread groups are also stalled until SM thread group 510 resumes execution and exchanges data and / or communicates with these other thread groups. As a result, the stalling of SM thread group 510 can affect multiple thread groups executing on one or more SMs 310. To mitigate the effects of the stall, SM thread group 510 may perform an asynchronous resume operation instead of a memory barrier operation.

[0075] Figure 6 According to various embodiments, Figure 4 6. As shown, an SM thread group 610 executing on an SM 310 generates two store operations 620 and 622. A load-store unit 612 included in the memory subsystem of the SM 310 receives the two store operations, annotated as store operations 630 and 632. The load-store unit 612 stores the data included in the store operations (annotated as store operations 640 and 642) in memory 614. As shown, store operation 640 completes before store operation 642. Because the SM 310 is part of a relaxed memory model processing unit, store operation 642 can complete before store operation 640. To ensure that both store operation 640 and store operation 642 have completed, the SM thread group 610 performs an asynchronous release operation, or store.release (store.release) operation 624, and sends the store.release operation 624 to the load-store unit 612. Store release operations 624 include store operations such as standard store instructions, compact instructions, indivisible read-modify-write instructions, etc. Store release operations 624 also include embedded memory barrier operations and flag store operations that do not stall the issuing SM thread group 610. The load store unit 612 detects the store release operation, annotated as store release operation 634. When the store release operation 634 reaches the load store unit 612, the load store unit 612 expands the store release operation 634 into two operations. The first operation is a memory synchronization operation or memory barrier (membar) operation 636, where the memory barrier operation 636 is similar to Figure 5 6. The second operation is a store operation 652 that is temporarily buffered in the load-store unit 612. The load-store unit 612 performs a memory barrier operation 636. The SM thread group 610 continues to execute instructions and performs further load operations and / or store operations, including load operations and / or store operations directed to memory 614.

[0076] Memory barrier operation 646 remains pending at load-store unit 612 until store operations 640 and 642 become visible in memory 614, along with any other pending store operations executed by SM thread group 610 prior to store-release operation 624. When the data included in these store operations becomes visible in memory 614, memory 614 issues a memory barrier acknowledgment (membar ack) 650. Memory 614 sends memory barrier acknowledgment 650 to load-store unit 612. When load-store unit 612 receives memory barrier acknowledgment 650, memory barrier operation 636 is completed. Load-store unit 612 executes store operation 652 to write a flag to a specified location, such as executing store operation 654 to write a flag to a location in memory 614. Store operation 654 completes store-release operation 634. Upon detecting that the flag has been written, threads included in SM thread group 610 and / or threads included in other SM thread groups may reliably read data from memory 614 , such as data stored in memory 614 via store operations 640 and 642 .

[0077] Memory barrier operation 636 and store operation 652 are asynchronous with respect to the original store-release operation 624. Therefore, store-release operation 624 does not stall execution of issuing SM thread group 610. Load-store unit 612 ensures that releasing store operation 652 (e.g., a flag write) is visible after earlier store operations 640 and 642. Therefore, this process meets performance objectives because issuing SM thread group 610 does not experience the throughput penalty associated with directly issuing memory barrier operations. This process further meets functional objectives because load-store unit 612 ensures proper synchronization of data by issuing memory barrier operation 636 and issuing store operation 652 after receiving memory barrier acknowledgment 650. Because SM thread group 610 executed store-release operation 624, SM thread group 610 is not stalled while memory barrier operation 636 is pending, but rather continues executing threads included in SM thread group 610. Furthermore, one or more other SM thread groups may rely on data from SM thread group 610, handle communications with SM thread group 610, and / or the like. Because SM thread group 610 is not stalled, these other thread groups can continue to exchange data and / or communicate with SM thread group 610 while memory barrier operation 636 is pending. As a result, performing memory release operation 624 can improve the performance of SM thread group 610 and / or the performance of other thread groups executing on one or more SMs 310.

[0078] Figure 7 According to various embodiments, Figure 4Flowchart of method steps for performing an asynchronous release operation on a streaming multiprocessor. Additionally or alternatively, the method steps may be performed by one or more alternative auxiliary processors including but not limited to a CPU, GPU, DMA unit, IPU, NPU, TPU, NNP, DPU, VPU, ASIC, FPGA and / or any combination thereof. Although combined Figures 1 to 6 The method steps are described herein with respect to a system, but one of ordinary skill in the art will understand that any system configured to perform the method steps in any order is within the scope of the present disclosure.

[0079] As shown, method 700 begins at step 702, where a processing unit, such as an SM 310 executing an SM thread group, sends an asynchronous release operation to an asynchronous release subsystem 400 included in the SM 310. The asynchronous release operation proceeds through the initial pipeline stages of the asynchronous release subsystem 400, much like any other load and / or store operation directed to memory. More specifically, the asynchronous release operation follows the same path as a memory synchronization operation, as the memory synchronization operation is passed from the SM 310 and received by the asynchronous release subsystem 400 associated with the SM 310.

[0080] In step 704, the asynchronous release subsystem 400 detects the asynchronous release operation and triggers the release memory barrier merger 422 contained in the asynchronous release subsystem 400. Once the asynchronous release operation is detected, the asynchronous release subsystem 400 performs two operations in parallel: (1) the asynchronous release subsystem 400 triggers the release memory barrier merger 422 (if the release memory barrier merger 422 has not been triggered previously); and (2) the release memory barrier merger 422 inserts a store operation to write a flag associated with the asynchronous release operation to the release queue 430. The store operation that writes the flag indicates that the asynchronous release operation has completed and is therefore released. Once the flag store operation is completed, other threads using the data associated with the asynchronous release operation can detect the flag and reliably access the data. The release queue 430 buffers relevant data to perform the flag store operation, including instruction metadata, virtual addresses, write data, and / or the like. The release queue 430 is partitioned to support multiple memory barrier stages. In some examples, the release queue 430 is partitioned to support two memory barrier stages. In this case, the asynchronous release subsystem 400 collects the data associated with the first set of flag store operations into a first buffer of the release queue 430. This buffer is referred to as the back buffer. In parallel, the asynchronous release subsystem 400 processes or flushes the data associated with the second set of flag store operations that were previously collected into a second buffer of the release queue 430. This second buffer is referred to as the current buffer. The ordered nature of the release queue 430 ensures that the order of the flag store operations stored therein is preserved, thereby preserving the same memory address ordering as the original store operations generated by the SM 310.

[0081] At step 706, the release memory barrier merger 422 merges the pending flag store operations to the same memory address into a single flag store operation. In doing so, the release memory barrier merger 422 attempts to identify other pending flag store operations to the same memory address and merges these flag store operations to the same memory address into a single flag store operation. Merging these flag store operations into a single flag store operation can reduce the number of pending flag store operations in the release queue 430 for a particular memory address. As a result, the memory of the release queue 430 is more efficiently utilized, and the total number of pending flag store operations can exceed the total number of entries in the release queue 430.

[0082] At step 708, the release memory barrier merger 422 merges the pending flag store operations to the adjacent memory addresses into a single request. In doing so, the release memory barrier merger 422 attempts to identify other pending flag store operations to the adjacent memory addresses and merges these flag store operations into a single flag store operation that represents a larger block of data than each of the original flag store operations. Merging these flag store operations for smaller blocks of data into a single flag store operation for a larger block of data can reduce the number of pending flag store operations in the release queue 430 for a particular set of memory addresses. As a result, the memory of the release queue 430 is more efficiently utilized, and the total number of pending flag store operations can exceed the total number of entries in the release queue 430.

[0083] In some examples, to perform the merge operations of step 706 and / or step 708, the release memory barrier merger 422 initializes a counter to a first predefined value. The release memory barrier merger 422 periodically increments or decrements the counter to a second predefined value, where the difference between the first predefined value and the second predefined value represents a duration. In some examples, the release memory barrier merger 422 may initialize the counter to a maximum counter value and may decrement the counter until the counter reaches zero. In some examples, the release memory barrier merger 422 may initialize the counter to zero and may increment the counter until the counter reaches the maximum counter value. While the counter is still counting and has not yet reached the second predefined value, the release memory barrier merger 422 merges flag store operations. When the counter reaches the second predefined value, the release memory barrier merger 422 stops merging flag store operations and forwards the flag store operations to the release queue 430.

[0084] At step 710, the release data consolidator 424 inserts the flag store operation into the release queue 430. Subsequently, the flag store operation is flushed from the release queue 430. In this manner, the release data consolidator 424 optimizes the flag store operation to better utilize the storage space in the release queue 430.

[0085] At step 712, the release memory barrier merger 422 clears the flag store operations from the release queue. When the data associated with the store release operation is visible in memory, the release memory barrier merger 422 receives a memory barrier confirmation. Upon receiving the memory barrier confirmation, the flush unit 440 included in the asynchronous release subsystem 400 begins processing entries of the flag store operations in the release queue 430, thereby clearing the release queue 430 of the flag store operations. As each entry is cleared from the release queue 430, the release memory barrier merger 422 forwards the corresponding flag store operation to the system memory 104, the PP memory 204, the memory-mapped I / O device, and the memory subsystem via the on-chip network 450 and / or other appropriate interfaces.

[0086] At step 714, the release memory barrier merger 422 switches the release queue 430 to a different memory barrier stage. While the flush unit 440 flushes the flag store operation for the current buffer of the release queue 430, the release memory barrier merger 422 continues to collect further flag store operations in one or more other back buffers of the release queue 430. Each buffer of the release queue 430 corresponds to a different memory barrier stage of the release queue 430. When the write store operation for the current buffer of the current memory barrier stage is flushed, the release memory barrier merger 422 switches the memory barrier stage of the release queue 430, causing one of the back buffers to become the current buffer and the previously current buffer to become the back buffer. In the new memory barrier stage, the release memory barrier merger 422 flushes the flag store operation for the newly current buffer, while the previously current buffer collects further flag store operations for the subsequent memory barrier stage of the release queue 430. Method 700 then continues to step 702, as described above, to process additional asynchronous release operations.

[0087] In summary, a thread writes data to memory through a series of memory operations and performs asynchronous release operations to synchronize data being transferred to one or more other threads. An asynchronous release operation is a store operation with built-in memory synchronization operations that does not stall the issuing thread during the memory synchronization operation. Instead, the load store unit receives the asynchronous release operation and queues the flag store operation into the release queue, allowing the issuing thread to continue executing normally. The memory barrier logic performs the memory synchronization operation. When the memory synchronization operation is completed, the memory barrier logic signals the completion of the memory synchronization operation to the release queue. At this point, the earlier memory operation is visible to the receiving thread, and the release queue can perform the flag store operation.

[0088] At least one technical advantage of the disclosed technique over the prior art is that, using the disclosed technique, a thread issuing an asynchronous release operation does not stall until a memory synchronization operation completes. Instead, after issuing an asynchronous release operation, the thread can perform further operations without waiting for the data associated with the asynchronous release operation to become visible to the receiving thread. Consequently, the execution performance of the issuing thread is improved compared to the prior art. Furthermore, because multiple asynchronous release operations can be combined into a single memory synchronization operation, fewer asynchronous release operations are issued, further improving performance. These advantages represent one or more technical improvements over prior art approaches.

[0089] Any and all combinations of any claim elements recited in any claim and / or any elements described in this application, in any manner, are within the scope of this disclosure and protection.

[0090] The description of the various embodiments has been presented for purposes of illustration, but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments.

[0091] Aspects of the present embodiments may be embodied as systems, methods, or computer program products. Thus, aspects of the present disclosure may take the form of entirely hardware embodiments, entirely software embodiments (including firmware, resident software, microcode, etc.), or embodiments combining software and hardware aspects, all of which are collectively referred to herein as "modules" or "systems." Furthermore, aspects of the present disclosure may take the form of a computer program product embodied in one or more computer-readable media containing computer-readable program code.

[0092] Any combination of one or more computer-readable media can be used. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. For example, a computer-readable storage medium includes, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared or semiconductor system, device or equipment, or any suitable combination thereof. More specific examples (non-exhaustive enumeration) of computer-readable storage media can include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the context of this document, a computer-readable storage medium can be any tangible medium that can include or store a program for use by or in conjunction with an instruction execution system, device or device.

[0093] Aspects of the present disclosure are described with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to embodiments of the present disclosure. It should be understood that each box in the flowchart and / or block diagram, and the combination of boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine so that the instructions executed by the processor of the computer or other programmable data processing device enable the function / action specified in one or more boxes of the flowchart and / or block diagram to be implemented. Such processors can be, but are not limited to, general-purpose processors, special-purpose processors, application-specific processors, or field programmable gate arrays.

[0094] The flowcharts and block diagrams in the figures illustrate the possible architectures, functions and operations of the systems, methods and computer program products according to the various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, segment or portion of a code, and the code includes one or more executable instructions for implementing one or more specified logical functions. It should also be noted that in some alternative embodiments, the functions indicated in the box may not occur in the order indicated in the figure. For example, two boxes shown in succession can actually be executed roughly at the same time, or the boxes can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, as well as the combination of boxes in the block diagram and / or flowchart, can be implemented by a system based on dedicated hardware that performs a specified function or action or a combination of dedicated hardware and computer instructions.

[0095] While the foregoing is directed to embodiments of the present disclosure, other and further embodiments of the disclosure may be devised without departing from the basic scope thereof, as determined by the claims that follow.

Claims

1. A computer-implemented method for transferring data between threads executing in a multi-processor computing system, the method comprising: receiving a first storage operation from a first processing unit, the first storage operation comprising a first memory synchronization operation and a first flag operation; as well as storing a first entry corresponding to the first flag operation in a side structure, wherein the first flag operation, when executed, indicates that data associated with a second store operation generated by the first processing unit prior to the first store operation is visible in memory, Wherein, while the first memory synchronization operation is suspended, the first processing unit continues to execute one or more operations.

2. The computer-implemented method of claim 1 , wherein at least one of the first store operation, the first memory synchronization operation, or the first flag operation is issued by a processor instruction executed by the first processing unit.

3. The computer-implemented method of claim 1 , further comprising: determining that the data associated with the second storage operation, generated by the first processing unit prior to the first storage operation, is visible in the memory; as well as The first flag operation is performed.

4. The computer-implemented method of claim 3, wherein: equipment: determining that the first flag operation has been performed; and In response, the data associated with the second storage operation is accessed.

5. The computer-implemented method of claim 4, wherein the device comprises at least one of a second processing unit or a peripheral device. The computer-implemented method of claim 3 , further comprising removing the first entry from the side structure.

7. The computer-implemented method of claim 3 , wherein the first entry is stored in a first buffer included in the side structure, and further comprising: receiving a third storage operation from the first processing unit, the third storage operation including a second memory synchronization operation and a second flag operation; as well as A second entry corresponding to the second flag operation is stored in a second buffer included in the side structure.

8. The computer-implemented method of claim 1 , further comprising: determining that data associated with a third storage operation generated by the first processing unit prior to the first storage operation is visible in memory; determining that data associated with a fourth storage operation generated by the first processing unit prior to the third storage operation is visible in the memory; as well as performing the first flag operation, wherein the data associated with the third storage operation is visible in the memory before the data associated with the fourth storage operation is visible in the memory.

9. The computer-implemented method of claim 1 , further comprising: receiving a third storage operation from the first processing unit, the third storage operation including a second memory synchronization operation and a second flag operation; determining that the second flag operation is directed to the same memory address as the first flag operation is directed to; generating a second entry corresponding to the first flag operation and the second flag operation; as well as The second entry is stored in the side structure in place of the first entry.

10. The computer-implemented method of claim 1 , further comprising: receiving a third storage operation from the first processing unit, the third storage operation including a second memory synchronization operation and a second flag operation; determining that the second flag operation is directed to a memory address adjacent to a memory address of the first flag operation; generating a second entry corresponding to the first flag operation and the second flag operation; as well as The second entry is stored in the side structure in place of the first entry.

11. The computer-implemented method of claim 10 , wherein the memory address of the first flag operation is within a first portion of a memory address range, the memory address of the second flag operation is within a second portion of the memory address range, and the first portion of the memory address range overlaps with the second portion of the memory address range.

12. The computer-implemented method of claim 1 , further comprising: receiving, from the first processing unit, a query operation requesting a status of the first storage operation; determining that the first flag operation is complete and has been cleared from the side structure; as well as A status message is sent to the first processing unit indicating that the first storage operation is complete.

13. The computer-implemented method of claim 1, wherein the first entry comprises one or more of metadata associated with the first flag operation, a virtual address associated with the first flag operation, and data associated with the first flag operation.

14. A system comprising: The first processing unit: generating a first storage operation, the first storage operation comprising a first memory synchronization operation and a first flag operation; as well as A memory synchronization subsystem, wherein the memory synchronization subsystem: receiving the first storage operation from the first processing unit, and A first entry associated with the first flag operation is stored in a side structure, wherein the first flag operation, when executed, indicates that data associated with a second memory operation generated by the first processing unit prior to the first memory operation is visible in memory, wherein the first processing unit continues to execute one or more operations while the first memory synchronization operation is pending.

15. The system of claim 14, wherein the memory synchronization subsystem further: determining that the data associated with the second storage operation, generated by the first processing unit prior to the first storage operation, is visible in the memory; and The first flag operation is performed.

16. The system of claim 15, further comprising a device, wherein: determining that the first flag operation has been performed; and In response, the data associated with the second storage operation is accessed.

17. The system of claim 15, wherein the first entry is stored in a first buffer included in the side structure, and wherein the memory synchronization subsystem further: receiving a third storage operation from the first processing unit, the third storage operation including a second memory synchronization operation and a second flag operation; and A second entry corresponding to the second flag operation is stored in a second buffer included in the side structure.

18. The system of claim 14, wherein the memory synchronization subsystem further: determining that data associated with a third storage operation generated by the first processing unit prior to the first storage operation is visible in memory; determining that data associated with a fourth storage operation generated by the first processing unit prior to the third storage operation is visible in the memory; as well as performing the first flag operation, wherein the data associated with the third storage operation is visible in the memory before the data associated with the fourth storage operation is visible in the memory.

19. The system of claim 14, wherein the memory synchronization subsystem further: receiving a third storage operation from the first processing unit, the third storage operation including a second memory synchronization operation and a second flag operation; determining that the second flag operation is directed to the same memory address as the first flag operation is directed to; generating a second entry corresponding to the first flag operation and the second flag operation; as well as The second entry is stored in the side structure in place of the first entry.

20. The system of claim 14, wherein the memory synchronization subsystem further: receiving a third storage operation from the first processing unit, the third storage operation including a second memory synchronization operation and a second flag operation; determining that the second flag operation is directed to a memory address adjacent to a memory address of the first flag operation; generating a second entry corresponding to the first flag operation and the second flag operation; as well as storing the second entry in the side structure in place of the first entry, The memory address of the first flag operation is within a first portion of a memory address range, the memory address of the second flag operation is within a second portion of the memory address range, and the first portion of the memory address range overlaps with the second portion of the memory address range.

Citation Information

Cited By

  • Data processing device and method, processor and chip

    CN120929142A