Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

32 results about "Collective operation" patented technology

A collective operation is a concept in parallel computing in which data is simultaneously sent to or received from many nodes. Common examples of collective operations are gather (in which data is collected from all nodes), scatter (in which a set of data is broken up into pieces, and a different piece is sent...

Offloading of adaptive all reduce operations

Examples described herein relate to a network interface device that includes: a host interface; a direct memory access (DMA) circuitry; a network interface to receive, in at least one packet, time data associated with at least one of multiple layers, wherein the multiple layers provide inputs to a collective operation associated with a large language model (LLM); and circuitry. The circuitry is to based, at least in part, on the time data associated with the multiple layers, identify a first operation of a first layer of the multiple layers as a late completing process relative to times to completion of multiple first operations of other layers and based on the first operation being identified as a late completing process, perform a remedial action to adjust at least one configuration of a first device to execute a second operation of the first layer.
Owner:INTEL CORP

Network collective offload message chunking management

The disclosed device can perform a collective operation on received datasets, and split the result into chunks in accordance with a chunking scheme. The device can also forward the chunks in accordance with a routing scheme that can direct chunks to appropriate nodes of a collective network. Various other methods, systems, and computer-readable media are also disclosed.
Owner:ADVANCED MICRO DEVICES INC

Collectives-aware load balancing

Systems, methods, and machine-readable media may facilitate programmable data trimming. Metadata associated with a collective operation may be determined by an application of a server. The metadata may specify a job identifier corresponding to a unit of work to be completed in conjunction with the collective operation, a collective type of the collective operation, and / or an ordering mode for packets corresponding to the collective operation. The metadata associated with the collective operation may be sent by the application to a network interface card (NIC). The NIC may be caused by the application to transmit a data packet with the metadata embedded in a cookie of the data packet to a switch of a network fabric to cause the switch to use a selected network path and / or selected load-balancing for the collective operation based on one or more of the job identifier, the collective type, and / or the ordering mode.
Owner:ORACLE INT CORP

Automatic run-time skew-aware optimization of rooted collectives

Automatic run-time skew-aware optimization of rooted collectives includes determining skews amongst compute nodes of a distributed computing system, as an application executes on the compute nodes, and determining implementations for collective operations of the application based at least in part on the skews. Skews may be determined based on timestamps of operations of the application program. Time stamps of one or more of the compute nodes may be estimated or inferred from timestamps of other compute nodes. Timestamps may be aggregated to determine global skews. Collective implementations may be determined for a sequence of collective operations based on skew impacts amongst the sequence of collective operations. Subsequent collective operations may be predicted based on current collective operations and a history of persistent collective operations, and implementations may be determined for the predicted collective operations prior to receipt of calls for the predicted collective operations.
Owner:XILINX INC

MPI collective operations

PendingUS20260203145A1Message Passing InterfaceMessage delivery
A method for performing a message passing interface (MPI) collective operation in a network, wherein the network comprises a plurality of interconnected nodes, the method comprising: receiving, at a node of the plurality of interconnected nodes, MPI collective operation information identifying the MPI collective operation to be performed, and a graph of the network; determining a number of algorithmic steps of the MPI collective operation based on the MPI collective operation and the graph of the network; determining an initialisation process for the algorithmic steps; determining a finalisation process for the algorithmic steps; determining, for each of the algorithmic steps: a subset of nodes of the plurality of interconnected nodes for the node to communicate with; one or more portions of data for the node to send to and receive from the nodes within the subset of nodes; and initialising the MPI collective operation based on the determined subset, initialisation process and finalisation process, and the one or more portions of data.
Owner:UCL BUSINESS LTD

Offloading of adaptive all reduce operations

Examples described herein relate to a network interface device that includes: a host interface; a direct memory access (DMA) circuitry; a network interface to receive, in at least one packet, time data associated with at least one of multiple layers, wherein the multiple layers provide inputs to a collective operation associated with a large language model (LLM); and circuitry. The circuitry is to based, at least in part, on the time data associated with the multiple layers, identify a first operation of a first layer of the multiple layers as a late completing process relative to times to completion of multiple first operations of other layers and based on the first operation being identified as a late completing process, perform a remedial action to adjust at least one configuration of a first device to execute a second operation of the first layer.
Owner:INTEL CORP

Symmetric multicast communication for offload operations

Systems, methods, apparatuses, and computer program products for zero-copy symmetric multicast communication buffers for offload operations. A method may include receiving an instruction for performing a collective operation across a plurality of processing elements. The method may also include determining an offset of the first virtual address from a first base address associated with the first set of contiguous virtual addresses. The method may further include translating the first virtual address to a corresponding multicast virtual address based on the offset. Further, the method may include causing the collective operation to be performed based at least on the multicast virtual address.
Owner:NVIDIA CORP

Compressed Transaction Layer Encoding for Collective Operations in UALink Networks

PendingUS20260252517A1Computer networkCollective operation
Implementations for compressed transaction layer encoding of collective operations in accelerator networks. An accelerator for an accelerator interconnect network may encode both collective operation requests and unicast requests into a compressed request field of a transaction layer flit using a shared multi-bit command field, enabling collective operations to achieve the same bandwidth efficiency as unicast operations at the transaction layer. The accelerator may further process compressed responses by extracting a single-bit indicator that distinguishes collective operation responses from unicast responses without reference to the original request, enabling stateless response differentiation at the receiver. The compressed encoding may overload existing identifier fields to carry group identifiers for collective operations while retaining physical accelerator identifiers for unicast operations. Implementations may further map the response type to protocol layer interface signals and may support dedicated compressed response field types for block collective operations.
Owner:UNIFABRIX LTD

Virtual Partition Lifecycle Management with Collective Operation Resource Allocation in a UALink Network

Implementations for virtual partition lifecycle management with collective operation resource allocation in an accelerator interconnect. In some implementations, a system for a UALink-based network comprises switches with group tables and queues for collective operations, and a centralized controller that, upon creating a virtual partition (which may be a virtual pod), programs forwarding entries and then programs group table entries and allocates queue entries for accelerators within the virtual partition. The centralized controller may deactivate group table entries and deallocate queue entries before removing forwarding entries on teardown. In some implementations, the centralized controller maintains per-virtual-partition resource accounting of group table entries and queue entries allocated to each virtual partition, and enforces per-virtual-partition quotas by rejecting allocation requests that would exceed a configured quota.
Owner:UNIFABRIX LTD

Switch-Offloaded Block Collective Operations with Autonomous Memory Access in a UALink Network

PendingUS20260252419A1Term memoryNetwork switch
Implementations for switch-offloaded block collective operations with autonomous memory access in an accelerator interconnect. In some implementations, a switch for a UALink-based network comprises a queue for storing collective operation parameters and a circuit configured to receive an invocation request carrying a control block specifying collective parameters, issue read requests to a group of destination accelerators, perform a reduction operation on received data, issue write requests to write reduced data, and write a completion status to a status buffer in the accelerator. The switch may constrain generated addresses to a bounded memory region using an address mask register. In some implementations, the switch manages queue resources through dynamic allocation of queue partitions, concurrent processing of collective operations from different partitions, and completion of outstanding operations before releasing allocated entries upon deallocation.
Owner:UNIFABRIX LTD

NETWORK OVERLOAD MARKER

Systems, methods, and devices for performing computational operations and managing network congestion are provided. An example describes a device comprising a processing unit that collects multiple messages and performs an operation as part of a collective operation on the data contained within the multiple messages. It then generates an output message containing the result of the operation performed on the data contained within the multiple messages. The processing unit can further embed a congestion notification in the output message if at least one of the multiple messages also contains a corresponding congestion notification.
Owner:MELLANOX TECHNOLOGIES LTD(IL)

Collective Operation Resilience for Chiplet-Based UALink Networks

PendingUS20260252512A1Networked systemOperating system
Implementations for collective operation resilience in a chiplet-based accelerator interconnect. In some implementations, a system for a UALink-based network comprises a chiplet die with stations coupled to an accelerator die via a die-to-die interface, switches with group tables identifying groups of accelerators for collective operations, and a centralized controller. Upon detection of a station fault on the chiplet die, the centralized controller identifies group table entries on the switches that reference accelerators coupled to the faulted station and updates the entries to exclude those accelerators, enabling collective operations to continue with reduced membership. In some implementations, an accelerator die coupled to chiplet dies redistributes transaction layer flits from a faulted chiplet die to operational chiplet dies, and the centralized controller updates forwarding entries on switches to reflect the changed connectivity.
Owner:UNIFABRIX LTD

Hardware based collective operations profiling

A system includes one or more processors to trace one or more packets transmitted by an application distributed among a plurality of computing nodes. The one or more processors are to generate tracing data based at least in part on tracing the one or more packets. The tracing data includes temporal information associated with transmission of the one or more packets. The one or more processors are to manage a data allocation associated with the application based on the tracing data.
Owner:MELLANOX TECHNOLOGIES LTD(IL)

Efficient iterative collective operations using a network-attached memory

A system receives a first request to perform a collective operation. The system stores a mapping of a first virtual address to a descriptor for a physical location of an allocated memory region. The system performs the collective operation, by writing data to a first segment of the memory region and accessing data from other segments of the memory region. The system receives a second request to perform an update operation, the second request indicating the first virtual address, one or more portions of a memory region segment to be updated, and corresponding data units to write to the portions. The system updates, based on the mapping, only the indicated portions by writing the corresponding data units. The system performs a subsequent iteration of the collective operation, based on the mapping, by bypassing writing any data to the memory region and only accessing data units from the memory region.
Owner:HEWLETT PACKARD ENTERPRISE DEV LP

Collective operation using a network-attached memory

In some examples, a processor receives a first request to allocate a memory region for a collective operation by process entities in a plurality of computer nodes. In response to the first request, the processor creates a virtual address for the memory region and allocates the memory region in a network-attached memory coupled to the plurality of computer nodes over a network. The processor correlates the virtual address to an address of the memory region in mapping information. The processor identifies the memory region in the network-attached memory by obtaining the address of the memory region from the mapping information using the virtual address in a second request. In response to the second request, the processor performs the collective operation.
Owner:HEWLETT PACKARD ENTERPRISE DEV LP

Efficient key management in distributed application

An apparatus facilitating efficient key refresh in a node is provided. During operation, the apparatus can determine a collective operation initiated by the node. The node can include a processor and can be in a distributed system comprising a plurality of nodes. The collective operation can be performed by a subset of the plurality of nodes in conjunction with each other. The apparatus can generate a new key based on a previous key maintained at the apparatus. Here, a respective key can be used for encrypting an inter-node packet in the distributed system. The apparatus can maintain the new and previous keys for the duration of the collective operation. Either of the new and previous keys can be used for decrypting messages received at the apparatus from other nodes of the distributed system. Upon determining a threshold point of the collective operation, the apparatus can discard the previous key.
Owner:HEWLETT PACKARD ENTERPRISE DEV LP

Selective Cryptographic Processing and Trust Elevation for Collective Operations in a UALink Network

PendingUS20260254662A1Trusted ComputingTransaction data
Implementations for selective cryptographic processing of collective and unicast traffic at a switch in an accelerator network. As AI training and inference workloads increasingly rely on in-network collective operations to accelerate gradient synchronization, broadcast, and reduction across large-scale accelerator pods, some implementations include a switch comprising a circuit that determines, for each encrypted transaction, whether it is a collective transaction or a unicast transaction. For collective transactions, the circuit decrypts transaction data for processing such as arithmetic reduction at the switch. For unicast transactions, the circuit bypasses decryption and forwards the transaction with data remaining encrypted between source and destination accelerators. Some implementations further include selective trust elevation of the switch into a trusted computing base for collective operations via a security manager, secure session establishment, attestation verification, and encryption key programming, while maintaining the switch outside the trusted computing base for unicast operations.
Owner:UNIFABRIX LTD

Acyclic architecture for ai processors

PCT designated stageWO2025244946A1Resource allocationBiological modelsCollective operationComputer engineering
An AI-accelerating processor system may include an acyclic subset of hardware processing nodes. The acyclic subset includes a plurality of end nodes that are disconnected from other end nodes in the acyclic subset. The acyclic subset of hardware processing nodes is configured to perform, according to schedules, computations that are part of a collective operation. A first hardware processing node in the subset has a first scheduling pattern and a second hardware processing node in the subset has a second scheduling pattern that is different from the first scheduling pattern to account for the subset being acyclic. The acyclic subset of hardware processing nodes is also configured to transmit computation outputs to neighboring hardware processing nodes among the acyclic subset through the bi-directional links to generate a result that is part of the collective operation. The result is contributed by each of the hardware processing nodes in the acyclic subset.
Owner:MATX INC

Acyclic architecture for ai processors

ActiveUS20250355715A1Resource allocationCollective operationComputer engineering
An AI-accelerating processor system may include an acyclic subset of hardware processing nodes. The acyclic subset includes a plurality of end nodes that are disconnected from other end nodes in the acyclic subset. The acyclic subset of hardware processing nodes is configured to perform, according to schedules, computations that are part of a collective operation. A first hardware processing node in the subset has a first scheduling pattern and a second hardware processing node in the subset has a second scheduling pattern that is different from the first scheduling pattern to account for the subset being acyclic. The acyclic subset of hardware processing nodes is also configured to transmit computation outputs to neighboring hardware processing nodes among the acyclic subset through the bi-directional links to generate a result that is part of the collective operation. The result is contributed by each of the hardware processing nodes in the acyclic subset.
Owner:MATX INC

Adaptive quantization and compression for variable-length collective communication

PendingUS20260180596A1Code conversionComputer hardwareCollective communication
A processing system dynamically and selectively quantizes or compresses variable-length input data for a collective operation to fit within a predetermined limit, executes the collective operation on the compressed data, and converts the data back to variable length results by dequantizing or decompressing the results of the collective operation.
Owner:ADVANCED MICRO DEVICES INC +1

Multi-Switch Group Table Consistency Management for In-Network Collective Operations in a UALink network

PendingUS20260254776A1Networked systemEngineering
Implementations for multi-switch group table consistency management for in-network collective operations in an accelerator interconnect. In some implementations, a system for a UALink-based network comprises a plurality of switches with group tables and a centralized controller that computes group table entries for collective groups spanning multiple switches, programs the entries on each switch, and verifies consistency across the switches. The centralized controller may use a three-phase activation protocol, programming entries with valid indicators cleared, verifying consistency, and then activating. In some implementations, the centralized controller maintains a topology model of the network, detects topology changes such as link or accelerator failures, identifies group table entries affected by the change, and updates the affected entries on each switch to maintain consistency between group membership and network reachability.
Owner:UNIFABRIX LTD

Collective operation and compute operation pipelining

Techniques for collective operation and compute operation pipelining on integrated circuit devices configured to perform operations associated with machine learning models are described herein. A first node or integrated circuit device can generate a first input subtensor and a second input subtensor based on the first input tensor. Several collective operations and compute operations can be performed on the first input subtensor and second input subtensor in combination with input subtensors from other integrated circuit devices to generate output subtensors. The first node can generate an output tensor based on the output subtensors.
Owner:AMAZON TECH INC

Network collective offloading cost management

ActiveUS12627589B2Digital computer detailsTransmissionNetwork topologyCollective operation
The disclosed device includes a collective engine that can select a communication cost model from multiple communication cost models for a collective operation and configure a topology of a collective network for performing the collective operation using the selected communication cost model. Various other methods, systems, and computer-readable media are also disclosed.
Owner:ADVANCED MICRO DEVICES INC

Multiple network routers performing collective operations and accelerator system including same

The disclosure relates to a plurality of network routers for performing collective operations and an accelerator system including the network routers. A network router includes: a receiver configured to receive a first input packet from a first other network router in a first direction, receive a second input packet from a second other network router in a second direction, and output one of the first input packet or the second input packet as a collective packet; a network controller configured to receive the collective packet from the receiver and output the collective packet through a first path or a second path based on a packet type of the collective packet; a buffer circuit configured to store, in a distinguishable manner, the collective packet transferred from the network controller through the second path based on a packet type of the collective packet; and a reduction operation circuit configured to receive the collective data packet stored in the buffer circuit, and perform a reduction operation using the received collective data packet.
Owner:SK HYNIX INC

OUTPATIENT ADAPTIVE ALL-REDUCE OPERATIONS

The examples described herein relate to a network interface device comprising: a host interface; a direct memory access (DMA) circuit arrangement; a network interface for receiving, in at least one packet, time data associated with at least one of several layers, the several layers providing inputs to a collective operation associated with a large language model (LLLM); and a circuit arrangement.The circuit arrangement is designed to identify, at least partially based on the timing data associated with the multiple layers, a first operation of a first layer of the multiple layers as a late-completed process relative to the times until the completion of multiple first operations of other layers, and, based on the first operation being identified as a late-completed process, to perform a remedial action to adapt at least one configuration of a first device to perform a second operation of the first layer.
Owner:INTEL CORP

Unreliable, out-of-order data transfers over memory interconnect architectures

The present disclosure relates to non-reliable, out-of-order data transfers over a memory interconnect architecture. An inter-process communication method divides a total payload of a first memory command among a plurality of packets and encodes a last transmitted packet of the plurality of packets to include a metadata code in the total payload. The memory command can be transmitted between processes that perform a collective operation. Payloads of the packets are written to memory in a sequential address order such that the metadata code is written to a particular address and a second process generates an acknowledgement to the first memory command to the first process on a condition that the metadata code is read back from the particular address.
Owner:NVIDIA CORP

Single-step collective operations

ActiveUS12505002B2Interprogram communicationCollective communicationComputer network
A method for collective communications includes invoking a collective operation over a group of computing processes in which the processes concurrently transmit and receive data to and from other processes in the group via a communication medium. Messages are composed for transmission by source processes including metadata indicating how the data to be transmitted by the source processes in the collective operation are to be handled by destination processes that are to receive the data and also including in at least some of the messages the data to be transmitted by one or more of the source processes to one or more of the destination processes. The composed messages are transmitted concurrently from the source processes to the destination processes in the group over the communication medium. The data are processed by the destination processes in response to the metadata included in the messages received by the destination processes.
Owner:MELLANOX TECHNOLOGIES LTD(IL)

Tree-based network architecture for accelerating machine learning collective operations

Aspects of the disclosure are directed to a tree-based network architecture for serving and / or training machine learning models. The architecture includes one or more multi-chip packages having a plurality of compute-memory stacks connected via an input / output (I / O) die. The I / O die includes an aggregator to aggregate computations from the compute-memory stacks. The architecture can further include a plurality of the multi-chip packages connected on a server via a server level aggregator and a plurality of the servers connected on a rack via a rack level aggregator for further aggregation of the computations from the compute-memory stacks. The tree-based network architecture allows for fewer hops, resulting in lower latency and savings in bandwidth when serving and / or training machine learning models.
Owner:GOOGLE LLC

Efficient key management in distributed application

An apparatus facilitating efficient key refresh in a node is provided. During operation, the apparatus can determine a collective operation initiated by the node. The node can include a processor and can be in a distributed system comprising a plurality of nodes. The collective operation can be performed by a subset of the plurality of nodes in conjunction with each other. The apparatus can generate a new key based on a previous key maintained at the apparatus. Here, a respective key can be used for encrypting an inter-node packet in the distributed system. The apparatus can maintain the new and previous keys for the duration of the collective operation. Either of the new and previous keys can be used for decrypting messages received at the apparatus from other nodes of the distributed system. Upon determining a threshold point of the collective operation, the apparatus can discard the previous key.
Owner:HEWLETT PACKARD ENTERPRISE DEV LP