Fusion data generation and associated communications

By integrating data generation and communication technology, the target communication module and update convergence unit are used to solve the problem of increased traffic in the multiprocessor system, efficient computing and network resource utilization are achieved, and equipment operation performance is improved.

CN120344964APending Publication Date: 2025-07-18ADVANCED MICRO DEVICES INC
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202380085279.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-03-27
Filing Date
2023-12-07
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In multiprocessor systems, alternating execution between data generation and communication operations leads to an increase in communication volume, affecting the overall device operation efficiency. Especially in large-scale deep neural network training, it is difficult for the prior art to effectively optimize the communication process.

Method used

By integrating data generation and communication technology, the target communication module, the data generation and communication tracking module and the update convergence unit are used to realize the concurrent execution of data generation and communication between multiple processors, reduce redundant memory traffic, and concurrently utilize computing and network resources.

Benefits of technology

It improves the operation performance of equipment, reduces the kernel startup overhead, realizes efficient utilization of computing and network resources, supports concurrent execution of computing and communication, and reduces the number of task startups.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120344964A_ABST
    Figure CN120344964A_ABST
Patent Text Reader

Abstract

Converged data generation and associated communication techniques are described. In one embodiment, a system includes a processing system having a plurality of processors. A data generation and communication tracking module is configured to track programmatically defined data generation and associated communications executed by the plurality of processors. The target communication module is configured to trigger a target communication of data between the plurality of processors based on tracked programmatically defined data generation and associated communications.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related Applications

[0002] This application claims priority to U.S. Patent Application Serial No. 18 / 190,620, filed on March 27, 2023, which claims priority to U.S. Provisional Patent Application No. 63 / 387,434, filed on December 14, 2022, and entitled, the entire disclosure of which is hereby incorporated by reference. Background of the Invention

[0003] Processing systems are configured to include multiple processors to improve computing efficiency, such as by using multiple cores of a central processing unit, a graphics processing unit, etc. For example, computations can be performed using multiple processors by alternating between performing computations and the associated data communication generated by the computations. Thus, scenarios involving increased communication traffic between processors (e.g., machine learning) have a direct impact on overall device operation and computing efficiency. Brief Description of the Drawings

[0004] The detailed description is described with reference to the accompanying drawings.

[0005] Figure 1 is a block diagram of a non-limiting example infrastructure configured to employ fused data generation and associated communication techniques.

[0006] Figure 2 is a block diagram of a non-limiting example of data generation and associated communication in a machine learning training example.

[0007] Figure 3 is Figure 1 a block diagram of a non-limiting example of a data generation and communication tracking module, a target communication module, and an update convergence unit.

[0008] Figure 4 is a block diagram of a non-limiting example of data generation and communication tracking that supports fused reduce-scatter operations.

[0009] Figure 5 is a block diagram of a non-limiting example of unordered data generation.

[0010] Figure 6 is a block diagram of a non-limiting example of ordered data generation.

[0011] Figure 7 is a flowchart of a non-limiting example of fused data generation and communication.

[0012] Figure 8 is a block diagram of a non-limiting example of a baseline system and a collective system based on fine-grained in-memory reduction.

[0013] Figure 9 It is a block diagram of non - limiting examples of fine - grained all - to - all operations and all - gather operations. Detailed implementation

[0014] Overview

[0015] In practical application scenarios, common computing methods in various fields are as follows: computations (e.g., deep learning training) are distributed across multiple processors (e.g., GPUs) for execution, and alternate between computations (e.g., computing weight gradients via GEMM operations) and associated communication operations (e.g., reducing weight gradients between GPUs through reduce - scatter computations). As the multi - dimensionality (e.g., neural network scale, dataset) expands, communication continues to increase, so communication optimization has a direct impact on overall device operation.

[0016] For example, large - scale deep neural networks (DNNs) typically rely on distributed training. This training involves storing parameters and activation values partitioned across nodes, and in conjunction with techniques such as data - parallel training, requires cross - node reduction of these structures in each training iteration. That is, each node independently generates data for these structures, and during each training iteration, the generated data is communicated and reduced among the processors participating in the computation.

[0017] To address these issues, this document describes techniques for fusing data generation and associated communication. These techniques are configured to be implemented by using enhanced components, examples of which include a target communication module, a data generation and communication tracking module, and an update convergence unit. This supports multiple technical advantages, including concurrent utilization of computing / networking, performance improvement, energy - efficiency improvement, avoiding separately launching kernels for computation / communication, etc. Various other examples are also considered, which are described in the following discussions and shown using corresponding figures.

[0018] In some aspects, the techniques described herein relate to a system that includes a processing system, the processing system including multiple processors, at least one of the multiple processors being configured to: track programmatically - defined data generation and associated communication performed by the multiple processors, and trigger target communication of data among the multiple processors based on the tracked programmatically - defined data generation and associated communication.

[0019] In some aspects, the techniques described herein relate to a system, wherein the programmatically - defined data generation and associated communication include: generating data by the at least one processor, and performing a target update by the at least one processor to send the data to another one of the multiple processors.

[0020] In some aspects, the techniques described herein relate to a system in which a target update is triggered when the at least one processor completes data generation.

[0021] In some aspects, the techniques described herein relate to a system in which a target update is triggered based on a remote communication event received at the at least one processor, the remote communication event being implemented by a data mover engine of another processor among the plurality of processors.

[0022] In some aspects, the techniques described herein relate to a system in which the remote communication event is part of a batch operation of communication involving the data.

[0023] In some aspects, the techniques described herein relate to a system in which programmatically defined data generation and associated communication are defined using a single fused data generation and associated communication operation.

[0024] In some aspects, the techniques described herein relate to a system in which the fused data generation and associated communication operation identifies another processor among the plurality of processors for receiving the data.

[0025] In some aspects, the techniques described herein relate to a system in which the fused data generation and associated communication operation identifies an address range as the source of the data or an address range as the destination for sending the data.

[0026] In some aspects, the techniques described herein relate to a system in which the at least one processor is further configured to support concurrent updates to the data in physical memory.

[0027] In some aspects, the techniques described herein relate to a system in which an in-memory processing component of a memory module including physical memory is configured to implement the concurrent update.

[0028] In some aspects, the techniques described herein relate to a system in which programmatically defined data generation and associated communication are configured to control the data generation order through corresponding processors among the plurality of processors.

[0029] In some aspects, the techniques described herein relate to a device that includes a processing system, the processing system including a plurality of processors, at least one of the plurality of processors being configured to: trigger a target communication of data between the at least one processor and another processor among the plurality of processors as part of programmatically defined data generation and associated communication, and resolve concurrent updates to the data in physical memory.

[0030] In some aspects, the techniques described herein relate to a device in which an in-memory processing component of a memory module that includes physical memory is configured to resolve concurrent updates to that data in the physical memory.

[0031] In some aspects, the techniques described herein relate to a device in which the at least one processor is further configured to: track programmatically defined data generation and associated communications performed by the plurality of processors, and trigger a target communication based on the tracked programmatically defined data generation and associated communications.

[0032] In some aspects, the techniques described herein relate to a device in which the target communication is configured to be performed based on a single fused data generation and associated communication operation performed by the at least one processor, and the operation identifies another processor among the plurality of processors as the destination for the data.

[0033] In some aspects, the techniques described herein relate to a method that includes: tracking programmatically defined data generation and associated communications performed among a plurality of processors of a processing system; triggering a target communication of data among the plurality of processors as part of the programmatically defined data generation and associated communications; and resolving concurrent updates to physical memory involving data generated by the plurality of processors.

[0034] In some aspects, the techniques described herein relate to a method in which the programmatically defined data generation and associated communications are configured to control the data generation order by respective processors among the plurality of processors.

[0035] In some aspects, the techniques described herein relate to a method in which the programmatically defined data generation and associated communications are configured to identify a particular processor among the plurality of processors for receiving data.

[0036] In some aspects, the techniques described herein relate to a method in which the programmatically defined data generation and associated communications are configured to identify an address range that is the source of the data or an address range that is the destination for the data.

[0037] In some aspects, the techniques described herein relate to a method in which the programmatically defined data generation and associated communications include: generating data by a first processor among the plurality of processors, and performing a target update by the first processor to send the data to a second processor among the plurality of processors.

[0038] Figure 1is a block diagram of a non-limiting example infrastructure 100 configured to employ fused data generation and associated communication. Infrastructure 100 includes a device 102 having a processing system 104. Processing system 104 includes a data mover engine 106, a memory controller 108, and a memory module 110 having a physical memory 112 and an in-memory processing component 114. Examples of physical memory 112 include random access memory implemented using one or more integrated circuits (e.g., double data rate synchronous dynamic random access memory). The in-memory processing component 114 is configured as an integrated circuit that includes both a processing component and a memory component implemented in hardware. Processing system 104 implements a plurality of processors, examples of which are shown as processors 116. Processor 116 represents at least one processor that implements the functions represented by data mover engine 106 and memory controller 108. In one example, memory module 110 is configured as a printed circuit board on which physical memory 112 and in-memory processing component 114 are mounted. Memory module 110 is communicatively coupled to the processor, e.g., via one or more buses on a motherboard that implements at least a portion of device 102. The processor can be configured as a central processing unit, a co-processing unit (such as a graphics processing unit), etc.

[0039] By way of example and not limitation, examples of configurations of device 102 include computing devices, servers, mobile devices (e.g., wearable devices, mobile phones, tablets, laptops), processors (e.g., graphics processing units, central processing units, and accelerators), digital signal processors, interference accelerators, disk array controllers, hard disk drive host adapters, memory cards, solid state drives, wireless communication hardware connections, Ethernet hardware connections, switches, bridges, network interface controllers, and other device configurations. Other examples include artificial intelligence training accelerators, encryption and compression accelerators, network packet processors, and video encoders and decoders.

[0040] The techniques described herein effectively support the fusion of data generation and associated communication by implementing mechanisms and primitives. In practical application scenarios, a common computing method in various fields is to allocate computations (e.g., deep learning training) to multiple processors (e.g., GPUs) for execution and alternate between computations (e.g., computing weight gradients via GEMM operations) and associated communication operations (e.g., reducing weight gradients between GPUs through reduction-scatter computations). As the multi-dimensions (e.g., neural network scale, dataset) expand, communication continues to increase, so communication optimization has a direct impact on overall device operation.

[0041] To this end, infrastructure 100 includes enhancement components. Examples of enhancement components include a target communication module 118 that is part of the data mover engine 106, and a data generation and communication tracking module 120. The infrastructure 100 also supports data synchronization using an update convergence unit 122 that is configured to utilize near-memory / memory-in offloading as an efficient synchronization infrastructure to support concurrent data generation and associated communication. The data mover engine 106, the target communication module 118, the data generation and communication tracking module 120, and the update convergence unit 122 are implemented in any one of hardware, software, firmware, or a combination thereof. In one example, these modules and units are configured as microcontrollers for performing various operations of the fused data management described below. In another example, these modules and units are implemented using hardware such as an application-specific integrated circuit (ASIC) or other integrated circuit (IC) to perform various operations of the fused data management described below.

[0042] This supports multiple technical advantages, including concurrent utilization of computing / networking, performance improvement, energy efficiency improvement, avoiding separately launching kernels for computing / communication, etc.

[0043] For example, the processing system 104 is configured to support the following scenario: Data locally generated on a first processor (e.g., processor 116) needs to be transmitted to another processor among multiple processors participating in the overall computation. In the following discussion, one such example involves training a machine learning model.

[0044] For example, large-scale deep neural networks (DNNs) typically rely on distributed training. This training involves storing parameters and activation values partitioned across nodes, and combined with techniques such as data parallel training, requires cross-node reduction of these structures via a "reduce-scatter operation" in each training iteration. That is, each node independently generates data for these structures and, during each training iteration, communicates and reduces the generated data among the processors participating in the computation.

[0045] The fusion of data generation and associated communication serves as a single fused operation that can concurrently execute these operations while also reducing redundant memory traffic. This supports several technical advantages, including improving the operational performance and energy efficiency of device 102, concurrent utilization (instead of serialized utilization) of computing and network resources, reducing the number of task / kernel launches, etc. Although the fusion of data generation and associated communication has several advantages, in some scenarios, it is too complex to implement solely using software. An example of such a situation involves generating data via a general matrix-matrix multiplication (GEMM) operation and communicating via a reduce-scatter operation.

[0046] To address these challenges, infrastructure 100 programmatically, i.e., in a programmer-defined manner, fuses data generation and associated communication. In Figure 1 the infrastructure 100, this is achieved by using data generation and communication tracking module 120 to track data generation and communication to enhance memory controller 108. Memory controller 108 is a digital circuit configured to manage data flow to and from physical memory 112. Target communication module 118 is implemented in data mover engine 106 for triggering target communication, e.g., sending to a defined address range (e.g., contiguous or non-contiguous), a defined processor, etc. The address range can be configured as the address range of the source of the data or as the address range of the destination of the data transmission. Additionally, in the illustrated example, update convergence unit 122 is also implemented to support computational operations typically associated with communication (e.g., the reduction operation in a reduce-scatter operation) by using near-memory / in-memory processing with physical memory 112 as a synchronization point. This supports local and remote updates of data with a relatively low synchronization overhead.

[0047] Figure 2 is a block diagram of a non-limiting example 200 of data generation and associated communication in a machine learning training example. This example shows data generation using a GEMM operation, followed by associated communication using a reduce-scatter operation.

[0048] In Figure 2 the reduce-scatter primitive depiction, an array with four partitions is shown, which will be reduced on four nodes (e.g., examples of processors 116 shown as "P0", "P1", "P2", and "P3") connected in a ring topology. In a first example 202 of a baseline system, each node (i.e., processors P0 - P3) first undergoes data generation via a GEMM operation. The GEMM operation used in machine learning typically involves a large amount of data generated in multiple steps, which are shown as four time steps "T1 - 4".

[0049] After data generation, a reduce-scatter operation is invoked. To implement this operation, nodes "P0 - P3" transfer a partition's worth of data and call a reduction kernel to reduce the received partitions using the locally available partitions, which requires two time steps in the steady state. Overall, for four partitions on four nodes, three operations need to be completed, e.g., each node transfers three partitions, performs three local reductions, and then receives three partitions. In the illustrated example, the four-node system consumes ten time steps to complete data generation (GEMM) and associated communication (reduce-scatter). After completing the reduce-scatter primitive operation, the nodes can also be configured to share the reduced partitions among the nodes.

[0050] On the other hand, in the second example 204, the target communication module 118, the data generation and communication tracking module 120, and the update convergence unit 122 are used to implement a mechanism for performing communication and reduction when generating each word (or word set) of data. Thus, generation and communication overlap at what is hereinafter referred to as "fine-grained". The techniques described herein also support the ability to program "coarse-grained" updates during data generation through the data mover engine 106 (e.g., direct memory access "DMA"). Both of these scenarios support programming these updates to achieve multiple communication modes.

[0051] In one example, the techniques described herein use the data generation and communication tracking module 120 to track data generation to opportunistically send the generated data as "fine-grained target updates" (e.g., "push") to other nodes. In a second scenario, the target communication module 118 utilizes the tracking performed by the data generation and communication tracking module 120 to trigger updates, e.g., as target updates coordinated by the data mover engine 106. The specific actions to be invoked and the address ranges to be tracked are fully programmable, e.g., programmed by a programmer or as part of an operating system. By using the techniques described herein, in the second example 204, data generation and the associated communication are completed in four time steps, as opposed to the ten time steps required by the first example 202 of the baseline system. As the amount of data to be processed increases, the advantages of these techniques also increase.

[0052] As the number of devices and the scale of GEMM increase, the advantages of these techniques are further amplified. For example, in the baseline, when the number of devices is "n", it involves "{2(n - 1)+n}" steps, while the techniques described herein only involve "n" steps. When the number of "n" and the GEMM size are large, the technique can reduce the time steps by two-thirds. Additionally, the techniques described herein can also be configured to utilize computing and network resources simultaneously, rather than in a serialized manner as in the baseline scenario.

[0053] Figure 3 It is a block diagram of a non-limiting example 300 of the data generation and communication tracking module 120, the target communication module 118, and the update convergence unit 122. These modules represent three logical parts of the described techniques, which support the fusion and overlap of data generation and the associated computations.

[0054] The data generation and communication tracking module 120 represents a function related to implementing low-overhead tracking operations (e.g., remote storage, DMA transfer, etc.) for local data generation 302 and remote data communication 304 as part of implementing programmable tracking-to-communication mapping 306. This supports the programming ability to effectively regulate data communication based on the progress and / or completion of data generation (e.g., local or remote) and other communications, thereby allowing the integration of data generation and communication. For example, the data generation and communication tracking module 120 supports the structures and mechanisms that allow a programmer to program and map a target communication (e.g., to a defined processor and / or address range) to specific data generation and / or communication events.

[0055] For example, the data generation and communication tracking module 120 is configured to implement operations on a tracked address range and execute the following rules:

[0056] · Forward = Y, DMA = N: Issue a read-modify-update locally and to a defined processor;

[0057] · Forward = N and DMA = Y: If local_counter = remote_counter = threshold, send a signal to the data mover engine;

[0058] · Forward = N and DMA = N: If local_counter = remote_counter = threshold, issue a protocol completion signal.

[0059] The target communication module 118 represents the function of a mechanism that performs operations and implements target data communication based on configurable conditions triggered by the tracking performed by the data generation and communication tracking module 120. This includes the fine-grained remote communication 308 and DMA-initiated bulk communication 310 as described above. An example of direct memory access enhancement implemented by the target communication module 118 includes "reading an address range from local memory on the memory control signal for address range x and initiating a read-modify-update to correct address range y in a defined processor".

[0060] The update convergence unit 122 can implement various operations to support scenarios involving computations associated with communication, e.g., reduction-scatter addition operations. This is performed by using a convergence mechanism to allow concurrent updates of data from local store / update 312 and remote store / update 314.

[0061] The communication of data is initiated in various ways, for example, based on the completion of local data generation, remote communication events, etc. To this end, the data generation and communication tracking module 120 implements lightweight tracking of data generation and communication by using enhanced functions of the memory controller 108 (e.g., using table structures). This tracking is used to trigger targeted fine-grained memory operations (e.g., triggering an update to a pre-programmed remote node when generating a local update) and / or targeted DMA-coordinated memory operations, e.g., programmed into the target communication tracking by the data mover engine 106.

[0062] Supporting "fine-grained" and "bulk" operations through direct memory access brings many technical advantages. In a first example, fine-grained memory operations support the immediate transfer of locally generated data to a remote node in a programmable manner. However, in some cases, this may result in high inter-node traffic. In addition, in addition to local generation, data communication can also be conditionally triggered based on remote communication. To address this issue, DMA-coordinated bulk communication is configured to implement multiple communication events, e.g., to support triggering of communication transmissions across multiple words. In addition, the programming described herein also supports triggering specific communication patterns upon completion of data generation or communication events, as Figure 6 further described.

[0063] The data generation and communication tracking module 120, the target communication module 118, and the update convergence unit 122 are implemented in any one of hardware, software, firmware, or a combination thereof. In the illustrated example, the data generation and communication tracking module 120 can be configured by using a microcontroller 316, which can execute instructions 318 as a dedicated machine to achieve the result of generating a programmable tracking-to-communication map 306. In another example, the data generation and communication tracking module 120 is configured to at least partially use hardware 320 (e.g., an integrated circuit 322 such as an application-specific integrated circuit) to generate a programmable tracking-to-communication map 306.

[0064] Similarly, the target communication and tracking module 118 can also be configured by using a microcontroller 324, which can execute instructions 326 as a dedicated machine to achieve the result of generating local storage / updates 312 and remote storage / updates 314. In another example, the target communication module 118 is configured to at least partially use hardware 328 (e.g., an integrated circuit 330, examples of which include application-specific integrated circuits) to generate local storage / updates 312 and remote storage / updates 314.

[0065] In addition, the update convergence unit 122 can also be configured by using a microcontroller 332, which is configured to execute instructions 334 as a dedicated machine to implement a convergence mechanism that supports concurrent updates using the physical memory 112. In another example, the update convergence unit 122 is configured to implement a convergence mechanism that supports concurrent updates using the physical memory 112 at least in part by using hardware 336 (e.g., an integrated circuit 338 such as an application-specific integrated circuit).

[0066] As described above, in some scenarios, the communication associated with data generation involves computations, e.g., reduction. To this end, the infrastructure 100 allows concurrent local data generation (including updates) while allowing remote updates to the data. In the techniques described herein, this is configured by the update convergence unit 122 using near-memory / in-memory processing. The update convergence unit 122 is configured to utilize the physical memory 112 (e.g., main memory) as a synchronization point. To this end, in one example, data generation is implemented only using updates rather than stores. Additionally, the implementation of the updates is performed by the update convergence unit 122, such that both local store / update 312 and remote store / update 314 can be executed concurrently with a low synchronization overhead.

[0067] Figure 4 is a block diagram of a non-limiting example 400 that supports data generation and communication tracking for fused reduce-scatter operations. In the illustrated example, a data generation and communication tracking table 402 implemented by the data generation and communication tracking module 120 is shown. Similarly, a target communication tracking table 404 implemented by the target communication module 118 is illustrated. In Figure 4 the illustrated example, each entry at each node is included in the respective tables. However, in practice, each node can also be configured to store table entries related to itself, e.g., the node "P0" stores the column labeled "P0".

[0068] The data generation and communication tracking table 402 is configured to track address ranges. For each range, the data generation and communication tracking module 120 implemented by the memory controller 108 uses local and remote counters respectively to track local store / update and remote store / update as Figure 3 shown. The data generation and communication tracking module 120 also tracks whether fine-grained local and / or remote updates are triggered for a given address range. The data generation and communication tracking module 120 tracks whether DMA-coordinated communication (e.g., updates) is triggered based on the completion of local data generation or communication events, e.g., "remote update = threshold, remote update = local update = threshold", etc. This can be used for both contiguous and non-contiguous address ranges in memory, e.g., using strided access, multi-dimensional access, or indirect access modes.

[0069] The techniques described herein integrate data generation and associated communication. To this end, these techniques support the ability to program fine-grained updates or updates coordinated with coarse-grained direct memory access during data generation. This supports the ability to programmatically implement any desired communication pattern.

[0070] In one such scenario, as Figure 2 shown, a reduce-scatter operation is programmatically implemented on a ring network. Specifically, the storage of a particular address range is tracked by the data generation and communication tracking module 120 of the memory controller 108 (e.g., the address range "1" at node P0), and is immediately forwarded to the specified address in the specified node, e.g., P1 in the ring topology, address range "A". At the same time, the storage is issued locally / remotely as a read-modify-update since the reduce-scatter involves a reduction operation. In addition, the programming implementation is such that when the local and remote updates to a particular address range reach a threshold (e.g., address range "3" on P0, threshold = 12), the memory controller 108 is programmed to signal this event to the data mover engine 106. The data mover engine 106 then triggers a pre-programmed target communication event through the target communication module 118, e.g., using the value read from the P0 address range "3" to update the DMA address range "C" on P1.

[0071] The nodes can be programmed in a variety of ways, such as programming a static network topology at startup or programming based on communication events. In addition, while the above description refers to the nodes as communication entities, in an alternative embodiment, other components in the system (e.g., switches, programmable accelerators) can also be configured as communication nodes. Additionally, although not depicted, conditions (e.g., local update = remote update = threshold for reduce-scatter) and / or operations (e.g., read-modify-update for reduce-scatter) can also be programmed for operations, applications, etc.

[0072] While several data generation and associated communication scenarios only involve communication (e.g., all-to-all communication in machine learning training), the present solution also contemplates alternative scenarios where the communication has associated computational operations (e.g., the reduction operation in reduce-scatter). To support these alternative scenarios, a low-overhead synchronization infrastructure is implemented by the update convergence unit 122, which supports concurrent data generation and data remote updates. For example, the update convergence unit 122 is configured to handle both local and remote updates of data.

[0073] In the above example of reduce-scatter, data generation stores are represented as updates. This is achieved in one particular implementation through software-level changes, e.g., directing stores to specific address ranges via page tables and / or cache-level mechanisms to bypass the cache, allowing the memory controller to convert local data generation stores to updates. Additionally, memory operations to remote nodes (due to the "forward" flag) or direct memory access coordination are also converted to updates. Each of these updates is offloaded to the update convergence unit 122 to complete processing.

[0074] The update convergence unit 122 can be configured in a variety of ways. In a first example, the update convergence unit 122 is implemented as a dedicated unit at a single level in the memory hierarchy, such as it can be encapsulated at the memory-side cache, at the memory controller, in the base die layer of a 3D memory stack, near a DRAM bank, etc. In scenarios where the update convergence unit 122 is placed at multiple levels, the update convergence units processing the same address coordinate with each other to ensure the correct application of local / remote updates.

[0075] Figure 5 is a block diagram of a non-limiting example 500 of unordered data generation. Figure 6 is a block diagram of a non-limiting example 600 of ordered data generation. The data generation order directly affects the communication pattern used to transfer data. Thus, by controlling the data generation order, the efficiency of fused data generation and the associated communication can be improved. In the example 500 of unordered data generation where the order is not controlled, the fused data generation and the associated communication are completed in six time steps. However, in Figure 6 the example 600 of, the data generation is sorted to improve efficiency, e.g., the fused data generation and the associated communication are completed in four time steps.

[0076] In a further particular implementation, priority information can be programmed into the Target Communication Tracking (TCT) table at the data mover engine 106 to prioritize the processing of specific communications over others, thereby further shortening the critical path. As Figure 6 shown at node P0, the communication for address range "4" is processed preferentially over the range "3".

[0077] Although the fusion of data generation and the associated communication has performance advantages, this fusion may result in higher concurrent memory traffic compared to serialized data generation and communication. Thus, as part of the infrastructure 100, a mechanism for managing interference is implemented. For example, when data generation is not complete, the communication memory traffic is deprioritized. Although these examples relate to reduce-scatter operations, the techniques described herein can also be used to forward remote communications to specified nodes in a fine-grained manner.

[0078] Figure 7 is a flowchart of a non - limiting example 700 of fused data generation and communication. A data generation and communication tracking module tracks programmed data generation and communication executed between multiple processors of a processing system (block 702). For example, the data generation and communication tracking module 120 implements the data generation and communication tracking table 402.

[0079] As part of the programmed data generation and communication, target communication of data between multiple processors is triggered (block 704). For example, the target communication module 118 triggers communication based on the tracking performed by the data generation and communication tracking module 120.

[0080] Concurrent updates are resolved to physical memory by an update convergence unit that involves data generated by multiple processors (block 706). For example, the update convergence unit 122 resolves local store / update 312 and remote store / update 314 to the physical memory 112.

[0081] In the above example, the techniques described herein support an infrastructure that can effectively fuse data generation and associated communication in a programmable manner. This achieves several technical advantages, including but not limited to: improved performance, concurrent utilization (as opposed to serialized utilization) of computing and network resources, offloading communication from the main processor (CPU / GPU) by programming communication once and implicitly triggering it when data generation and / or communication is complete, reduced kernel startup overhead, etc. In the following discussion, these techniques are used in a specific implementation example for collective operations based on fine - grained in - memory reduction.

[0082] Reduction - based collective operations are used as part of the training of natural language processing applications in a multi - device setup. These collective operations involve data communication and reduction from multiple devices and are used to aggregate gradients (in a data - parallel setup) or activation values (in a model - parallel setup) during training.

[0083] However, these collective operations are typically serialized with the application execution and can become a bottleneck, resulting in sub - linear scaling of performance as the number of devices increases during training. However, the data used by these collective operations is typically not generated simultaneously. For example, data generated through matrix multiplication (GEMM) operations is executed in multiple stages, each stage containing a set of workgroups. Thus, in a specific implementation, data communication and reduction from a single GEMM stage overlap with the execution of the next GEMM stage in a fine - grained manner. This reduces the overhead of the collective operations with the producer kernel.

[0084] There are several challenges in implementing such functionality. For example, producer and collective operations are typically implemented as separate kernels in a graphics processing unit, which involves computationally expensive synchronization if executed in a fine-grained manner. Additionally, contention for computational and memory resources between collective operations and the producer GEMM phase degrades overall performance.

[0085] To overcome these challenges, a hardware / software mechanism is described that transparently executes producer and collective operations in a fine-grained manner. This mechanism automatically initiates fine-grained communication of data based on the producer's store instructions by leveraging the address space, and thus can be executed without modifying the kernel. Additionally, these techniques also utilize near-memory computing units to atomically update memory locations based on store instructions, thereby limiting contention with producer operations. As a result, this mechanism reduces communication overhead and frees up computational resources (e.g., graphics processing units) from performing reduction operations. This enables the training efficiency to scale efficiently in an approximately linear fashion as the number of devices increases. Additionally, this mechanism also accelerates collective operations (via fewer memory accesses) while improving the overall utilization of computational and network resources.

[0086] For example, large network matrix multiplication operations (GEMM) are executed in multiple phases and generate data. Additionally, GEMMs from transformer models typically have large output sizes, which are processed in chunks / slices and involve a large number of workgroups (WGs) or thread blocks (TBs) for computation. In practice, due to the limited number of graphics processors or streaming multiprocessors, these workgroups do not typically execute simultaneously. Instead, they are usually executed in phases, where each phase is a group of workgroups or thread blocks that can be accommodated by the graphics processing unit. The number of phases varies depending on the size, shape, and specific implementation of the kernel used for the GEMM. Therefore, the output of the GEMM, and thus the output of the entire layer, is typically not produced all at once but in multiple phases. This remains the case even when the operation is split across devices via model parallelism. This is because GEMMs that are split across devices and involve "all-reduce" collective operations are typically split in the "K" dimension. Therefore, the amount of work performed by threads or workgroups in each sub-GEMM is usually smaller (e.g., dot product operations on shorter rows and columns), but the size of the output matrix generated by each sub-GEMM is still the same as that of the original GEMM. This means that the number of threads / WGs and the number of phases executed by each sub-GEMM remain similar.

[0087] The mechanism described in this paper leverages this insight to overlap the reduction / communication of data (e.g., "all-reduce" operations) with data generation. For example, the communication of data generated in one phase is overlapped with the data generation (computation) in the next phase, and this communication operation is "hidden" during the data generation (computation) in the next phase.

[0088] For example, the mechanism described herein enables the producer GEMM and the fine-grained execution of collective operations to be implemented in a transparent manner by causing the GEMM writes to automatically trigger the communication / reduction of the generated data. This mechanism is achieved by allocating the output of the GEMM within a specific address space while keeping the GEMM kernel unchanged. In this example, the reduction operation is then fully handled by the hardware.

[0089] In addition, the overlap between GEMM and collective operations also leads to contention for graphics processing unit resources and slows down the overall execution speed. There are two sources of contention between GEMM and collective operations. The first is contention for the compute units of the graphics processing, which slows down the performance of GEMM. The second is that the reduction operation is memory-intensive and contends for memory bandwidth with the producer GEMM operation. In one example, to address this issue, the collective operation is automatically initiated when the GEMM writes to the address space. Therefore, no additional compute units are required to execute this collective operation. In addition, these writes are converted to updates in real time and processed by the near-memory arithmetic logic unit, thus incurring only a minimal additional memory overhead compared to the original GEMM writes.

[0090] Figure 8 is a block diagram of a non-limiting example 800 of a baseline system 802 and a collective system 804 based on fine-grained in-memory reduction. The collective system 804 based on fine-grained in-memory reduction is illustrated as a "full reduction" collective operation in a simple dual-device system.

[0091] In the baseline system 802, the graphics processing unit first executes the corresponding producer GEMM and stores the output in local memory. Subsequently, the graphics processing unit initiates a reduce-scatter operation, where each graphics processing unit reduces a "block" of the output array (i.e., this graphics processing unit is the master node of the block). This requires ensuring that each graphics processing unit has each copy of the block it is responsible for through direct memory access transfers (or point-to-point copies). Subsequently, each graphics processing unit performs a memory load of the copy, a reduction operation, and stores the reduced version locally. Finally, the reduced versions of the blocks are transferred (e.g., broadcast) to the remaining devices to complete the "all-gather" operation. The total number of memory loads / stores depends on the topology, the number of devices, and the algorithm (e.g., ring algorithm versus direct algorithm) used by the baseline system 802.

[0092] On the other hand, in the collective system 804 based on fine-grained in-memory reduction, the collective operations are transparently executed in a fine-grained manner with the producer GEMM operations, and the execution time of the collective operations is "hidden". In this example, to perform a reduce-scatter operation, each GEMM write is directed to a local (if it is the master node of an array element) or remote memory location instead of being directed to local memory. Additionally, in this example, writes to the specified location use a near-memory arithmetic logic unit to atomically update the data. Thus, each main memory location contains the fully reduced version of the data after receiving writes from each relevant device. Subsequently, these blocks can be transferred to other devices to complete the "all-gather" operation of the data.

[0093] Thus, in this example, the reduce-scatter of data overlaps with data generation. In this example, this operation is fully coordinated by hardware, thus reducing software complexity and further reducing the overall memory traffic. As Figure 8 shown, for example, in the baseline system 802, the data corresponding to each element requires nine read / write operations in local / remote memory, while due to the concurrent execution of GEMM and collective operations, the collective system 804 based on fine-grained in-memory reduction only requires four operations.

[0094] The mechanisms described herein include supporting the automatic initiation of communication / reduction of data according to the write instructions of the producer. The mechanisms also utilize near-memory computing to atomically update memory locations. To this end, the mechanisms implement an address space for transparently fusing producer operations and collective operations.

[0095] To avoid the complexity of fine-grained collective operations in software and avoid modifying the implementation of hundreds of GEMM kernels in a large library, in this example, the fine-grained execution of producer GEMM and collective operations is transparently implemented in hardware. To this end, the output of the producer GEMM is allocated in an address space such that writes to this address space automatically perform the required collective operations.

[0096] As Figure 8 shown, writes can be used to trigger three types of actions: local, remote, and direct memory access (DMA). Additionally, for different types of collective operations (e.g., all-reduce, reduce-scatter) and techniques (e.g., ring algorithm, direct algorithm), the memory locations and execution order of these actions may be different. Thus, the systems implementing the mechanisms described herein are configured to support memory-mapped APIs that can be used to configure the allocated memory for different write-triggered actions for different collective operation types and techniques.

[0097] This mechanism can be configured using a library with a predefined memory mapping, and the corresponding application can "call" this library. For example, in a four-GPU all-reduce operation, memory is allocated on each device in the address space by specifying a collective operation and mechanism. This function first allocates an array on each device. For an "all-reduce" operation, a local allocation of the entire array is performed on each device to collect the final reduced version of the entire array on each device. Subsequently, through an API call, the locally allocated sub-arrays are mapped to the remote allocations of the array for remote writes. Thus, the output array on each device in this example is mapped to distributed physical memory. This mapping ensures that writes to the local sub-arrays are redirected as remote writes to the corresponding master nodes. In addition, this mapping also defines the specific operations (e.g., updates) performed by the remote write operations. After the allocation is completed, the GEMM operation is executed, and subsequently, the reduced data is transferred from the remote memory to the local memory through additional direct memory access (or point-to-point copying).

[0098] For the memory allocation in the address space, since writes are not locally read until the reduction is complete, the device does not cache writes to this address space. Thus, in this example, writes are written to physical memory 112, e.g., dynamic random access memory (DRAM). In addition, the storage of these pages, if originating from the master node itself, is directed to local physical memory 112 or directly to remote physical memory 112 to avoid redundant writes and relieve memory bandwidth pressure. This also ensures that there is only one aggregation point for each copy of the data. This mechanism is implemented by extending the translation lookaside buffer and page tables to include the local and remote physical addresses of the pages in memory, or via a separate hardware structure. If stored to local physical memory 112, the stores to these locations are sent to the memory controller 108, while the stores to remote physical memory 112 are directed to the remote graphics processing unit memory controller.

[0099] The physical memory 112 on the master device can be used as the aggregation unit for each copy of the array. Local stores issued from the master device and remote stores from other devices are received and queued by the memory controller 108 and then sent to physical memory 112, such as dynamic random access memory. The loading of these pages is only performed as part of the next graphics processing unit kernel. In one example, after the GEMM is completed, a system-wide fence is inserted as part of the direct memory access function to ensure that all stores and direct memory accesses to the location are completed. Thus, the loading is directed to the local copy through the translation lookaside buffer and page tables.

[0100] In a DRAM architecture that supports near-memory computing, each bank is associated with an arithmetic logic unit (ALU) and registers for storing intermediate values. Thus, stores to these memories can be used to update memory locations. Accordingly, a DRAM bank associated with the address space of the techniques described herein can be programmed to update a memory location in accordance with a store command.

[0101] Such updates first write the stored value to a register associated with the near-memory arithmetic logic unit, activate the corresponding memory row, read the column values in the row buffer and add them to the data in the register, and then write the reduced value back to the buffer. Queuing stores or near-memory updates in the memory controller 108 queue can improve the atomicity of these updates such that only a single instruction is issued and executed to the arithmetic logic unit corresponding to the memory location at a given time. Additionally, converting these stores to atomic updates in real time does not violate the memory consistency guarantees of the graphics processing unit. These updates have the property of commutative atomic operations and thus, similar to store operations, can be reordered with respect to other relaxed atomic operations (in this case, also store operations). In one example, a memory queue merger merges these stores / updates in the queue to improve performance. Merging multiple updates to the same location helps reduce the number of row activations and / or row buffer reads / writes. Overall, these near-memory update-based reductions reduce and (in some cases) eliminate contention for memory resources by the executing GEMM. For direct reduce-scatter operations, the total number of memory operations involved in the reduction is the same as the number of operations when the GEMM is executed alone.

[0102] Figure 9 is a block diagram of a non-limiting example 900 of a fine-grained all-to-all operation 902 and an all-gather operation 904. The traffic pattern of the all-to-all operation 902 matches the traffic pattern of the all-reduce collective operation and thus can utilize the same configuration as shown, except that writes do not update memory. On the other hand, the all-gather operation 904 is implemented by directing GEMM writes to both local and remote memories. Figure 8 On the other hand, the all-gather operation 904 is implemented by directing GEMM writes to both local and remote memories.

[0103] Additionally, collective (e.g., "all-reduce") operations in natural language processing applications are typically followed by other memory-intensive operations on each participating device (e.g., parameter updates in data parallel setups or residual / dropout layers in model parallel setups). However, these operations consume the entire reduced array on each device, and thus there is redundancy in some cases. Therefore, the reduction performance in memory provides an opportunity to limit such redundant operations. Consumer operations can also be performed using a near-memory arithmetic logic unit, operating on the (reduced) data subarray on the master node before being "all-gathered" or broadcast to the remaining devices. This reduces redundant computations and further improves distributed natural language processing performance.

[0104] It should be understood that there may be many variations based on the disclosure herein. Although the above features and elements are described in specific combinations, each feature or element can be used alone without other features and elements, or in various combinations with or without other features or elements.

[0105] The various functional units shown in the figures and / or described herein (including device 102 where appropriate) are implemented in any of a variety of different ways, such as hardware circuits, software or firmware executed on a programmable processor, or any combination of two or more of hardware, software, and firmware. The provided methods are implemented in any of a variety of devices such as a general-purpose computer, a processor, or a processor core. For example, suitable processors include general-purpose processors, dedicated processors, conventional processors, digital signal processors (DSPs), graphics processing units (GPUs), parallel acceleration processors, multiple microprocessors, one or more microprocessors associated with a DSP core, controllers, microcontrollers, application-specific integrated circuits (ASICs), field-programmable gate array (FPGA) circuits, any other type of integrated circuit (IC), and / or state machines.

[0106] In one or more specific implementations, the methods or processes provided herein are implemented in a computer program, software, or firmware incorporated in a non-transitory computer-readable storage medium for execution by a general-purpose computer or processor. Examples of non-transitory computer-readable storage media include read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media (such as internal hard disks and removable disks), magneto-optical media, and optical media (such as CD-ROM disks) and digital versatile disks (DVDs).

[0107] Although systems and techniques have been described in language specific to structural features and / or method acts, it should be understood that the systems and techniques defined in the appended claims are not necessarily limited to the specific features or acts described. These specific features and acts are disclosed as example forms of implementing the claimed subject matter.

Claims

1. A system, the system comprising: A processing system, the processing system including a plurality of processors, at least one of the plurality of processors being configured to: Track programmatically defined data generation and associated communications performed by the plurality of processors; And Trigger target communications of data among the plurality of processors based on the tracked programmatically defined data generation and associated communications.

2. The system according to claim 1, wherein The programmatically defined data generation and associated communications include: generating the data by the at least one processor, and performing a target update by the at least one processor to send the data to another one of the plurality of processors.

3. The system according to claim 2, wherein The target update is triggered when the at least one processor completes the generation of the data.

4. The system according to claim 2, wherein, The target update is triggered based on a remote communication event received at the at least one processor, the remote communication event being implemented by a data mover engine of another one of the plurality of processors.

5. The system according to claim 4, wherein, The remote communication event is part of a batch operation of communications involving the data.

6. The system according to claim 1, wherein, The programmatically defined data generation and associated communications are defined using a single fused data generation and associated communication operation.

7. The system according to claim 6, wherein, The fused data generation and associated communication operation identifies another one of the plurality of processors for receiving the data.

8. The system according to claim 6, wherein The fused data generation and associated communication operation identifies an address range as a source of the data or an address range as a destination for sending the data.

9. The system according to claim 1, wherein The at least one processor is further configured to support concurrent updates to the data in physical memory.

10. The system according to claim 9, wherein, An in-memory processing component of a memory module including the physical memory is configured to implement the concurrent updates.

11. The system according to claim 1, wherein, The programmatically defined data generation and associated communications are configured to control a data generation order through corresponding processors among the plurality of processors.

12. A device, the device comprising: A processing system, the processing system including a plurality of processors, at least one of the plurality of processors being configured to: As part of programmatically defined data generation and associated communications, trigger target communications of data between the at least one processor and another one of the plurality of processors; And Resolve concurrent updates to the data in physical memory.

13. The device according to claim 12, wherein, An in-memory processing component of a memory module including the physical memory is configured to resolve the concurrent updates to the data in the physical memory.

14. The apparatus according to claim 12, wherein, The at least one processor is further configured to: track the programmatically defined data generation and associated communications performed by the plurality of processors, and trigger the target communications based on the tracked programmatically defined data generation and associated communications.

15. The apparatus according to claim 14, wherein, The target communications are configured to be performed based on a single fused data generation and associated communication operation performed by the at least one processor, and the operation identifies another one of the plurality of processors as a destination for sending the data.

Citation Information

Cited By

  • Fused data generation and associated communication

    US12474926B2