Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

7254results about "Physical realisation" patented technology

AI Serving Hardware and Software Frontier Enhancements

A computer system implements a unified framework integrating an adaptive elastic funnel (AEF) with a convergent intelligence fabric (CIF) for multi-agent AI collaboration. The system provides a universal multi-modal key-value subsystem for sharing partial computations, implements hybrid placement strategies for dynamic memory management, and incorporates quantum-resistant secure enclaves. The architecture integrates hardware acceleration through GPU-FPGA hybrid caching and neuromorphic processors, applies adaptive energy and thermal management across hardware generations, and implements autonomous flash resource orchestration with multi-dimensional wear management. The system orchestrates tensor workflows using hierarchical scheduling, enables cross-agent collaboration with privacy preservation, and supports continuous learning without catastrophic forgetting. This integration delivers unprecedented computational efficiency and security in high-dimensional decision-making environments while supporting incremental adoption through modular interfaces.
Owner:QOMPLX INC

Dynamic Latent Space Adaptation Based on Spatiotemporal Kernal Context for Multiscale Rendering

A system for dynamic latent space adaptation using spatiotemporal kernel context for multiscale rendering with hierarchical and Lorentzian autoencoders. The Spatiotemporal Kernel Estimator (SKE) analyzes media through motion field, temporal recurrence, frequency band, and scene semantics analyzers to generate adaptive kernel parameters encoding content-specific importance distributions. The system dynamically adapts latent manifold geometry by modifying metric tensor properties according to kernel context, enabling content-aware compression that allocates representational capacity based on visual significance. A multiscale cache implements kernel-adaptive retention policies prioritizing important regions. An adaptive renderer provides intelligent level-of-detail selection based on zoom level and kernel-estimated importance, optimizing processing allocation. The self-optimizing architecture continuously refines kernel context and geometric adaptation based on user interaction and performance feedback, achieving superior compression ratios and perceptual quality. Applications include bandwidth-efficient video streaming, virtual reality, scientific visualization, and cognitive video analytics requiring intelligent context-aware visual processing.
Owner:ATOMBEAM TECH INC

Method and apparatus for efficient access to multidimensional data structures and / or other large data blocks

A parallel processing unit comprises a plurality of processors each being coupled to a memory access hardware circuitry. Each memory access hardware circuitry is configured to receive, from the coupled processor, a memory access request specifying a coordinate of a multidimensional data structure, wherein the memory access hardware circuit is one of a plurality of memory access circuitry each coupled to a respective one of the processors; and, in response to the memory access request, translate the coordinate of the multidimensional data structure into plural memory addresses for the multidimensional data structure and using the plural memory addresses, asynchronously transfer at least a portion of the multidimensional data structure for processing by at least the coupled processor. The memory locations may be in the shared memory of the coupled processor and / or an external memory.
Owner:NVIDIA CORP

End-side multi-mode large model accelerated reasoning method and system

The invention provides an end-side multi-modal large model accelerated reasoning method and system, and the method comprises the steps: carrying out the two-stage screening and rearrangement of visual tokens based on the CLS attention and text-to-visual attention in a visual encoder and pre-filling stage, and constructing a sparse attention and sparse key value cache; in a decoding stage, an important neuron set is judged according to activation gating or historical statistics, only a corresponding feedforward network weight is pulled and calculated, missed weights are loaded on demand through asynchronous I / O, and hot neurons are maintained in a high-speed memory to utilize model sparsity, so that video memory / memory occupancy and calculation overhead are remarkably reduced on an end side; throughput and time delay performance are improved. According to the method, the internal memory and computing resources required by reasoning of the multi-modal large language model are reduced from two dimensions by utilizing the endogenous sparsity of the end-side large language model in input and the model, so that a higher reasoning speed is achieved by utilizing fewer resources on the premise of keeping the size of the model unchanged, and the performance of the whole system is improved.
Owner:SHANGHAI JIAOTONG UNIV

Parameter-efficient large-language fine-tuning federated learning framework

Provided in the present invention is a parameter-efficient large-language fine-tuning federated learning framework, comprising the following steps: performing modeling on LoRA adapters of different edge clouds; since different weights exhibit different average performances on the LoRA adapters, using singular values to quantify the importance of the weights, and therefore, before each round of independent training of the LoRA adapters using N edge clouds, using a matrix singular value to decompose a BA matrix in the LoRA adapter for each trainable weight; configuring heterogeneous LoRA adapters on the basis of the importance of the weights; and using different numbers of quantization bits to quantize a pre-trained model, and performing high-precision inverse quantization on the pre-trained model only when matrix multiplication is executed, wherein the pre-trained model is quantized to the maximum number of quantization bits on the basis of the memory budget of the edge clouds. The present invention has the following beneficial effects: the present invention determines the optimal fine-tuning model structure, thereby improving the performance of LLM fine-tuning, and adapts to heterogeneous and resource-constrained edge clouds.
Owner:FUDAN UNIVERSITY

In-memory computing circuit chip based on magnetic cache and computing device

The embodiment of the invention discloses an in-memory computing circuit based on a magnetic cache, and the circuit comprises at least one magnetic cache unit, at least one in-memory computing unit, and a timer. The magnetic cache unit in the at least one magnetic cache unit is used for caching data output by the corresponding in-memory computing unit as to-be-processed data within the corresponding data retention time; the timer is used for respectively setting data retention time for the at least one magnetic cache unit; and the in-memory computing unit in the at least one in-memory computing unit is used for extracting the data to be processed from the corresponding magnetic cache unit for calculation and outputting the computed data to other magnetic cache units. According to the embodiment of the invention, the invention achieves the flexible adjustment of the data retention time of the magnetic cache unit in various in-memory calculation scenes, and achieves the provision of a high-capacity cache for the data needed by in-memory computing under the lower power consumption.
Owner:NANJING HOUMO TECH CO LTD

Target recognition model reasoning optimization method and device

The invention provides a target recognition model reasoning optimization method and device, and the method comprises the steps: firstly carrying out the structural analysis and sensitivity evaluation of a pre-training model, extracting the structural features of each network layer, activating the distribution features, carrying out the quantitative sensitivity scoring, and constructing a data set reflecting the hierarchical features and fault-tolerant capability; and querying a quantitative configuration knowledge base based on the data set to generate a heterogeneous quantitative strategy. Layered low-bit quantization is executed according to the strategy, and a layered weighted loss function is introduced to carry out quantization perception training, so that precision loss caused by bit width compression is effectively compensated. According to the method, through hierarchical heterogeneous quantification, the model recognition precision is preserved to the maximum extent while high compression ratio and reasoning acceleration are achieved, and particularly, the performance of a high-sensitivity layer is protected. The generated heterogeneous quantitative model remarkably reduces memory occupation and power consumption, is suitable for an edge hardware platform with limited resources, forms a set of complete automatic process from analysis and configuration to training compensation, and has good universality and engineering practical value.
Owner:CHINA WEAPON EQUIP RES INST

Temporal dynamics simulation in matmul-free neural architectures

A method is provided for processing data in a neural network system. The method includes receiving input data; processing the input data through a first set of neural network layers configured to perform data processing using MatMul-free techniques to produce intermediate data; further processing the intermediate data through a second set of neural network layers configured to simulate spiking neural network (SNN) functionalities using MatMul-free techniques; and outputting a result based on the processed data from the second set of neural network layers.
Owner:LEPTUDE INC

Intelligent scheduling method and system for Huaan Atlas heterogeneous computing resources based on dynamic load awareness

The invention belongs to the technical field of computing resource scheduling, and particularly relates to an intelligent scheduling method and system for Huaan Atlas heterogeneous computing resources based on dynamic load awareness. The method comprises the steps that the real-time state of multi-dimensional hardware data is collected, and a basic data source is provided for subsequent steps; dynamically adapting tasks and hardware characteristics through a matching degree matrix, modeling aiming at various basic data, and constructing a state vector required by reinforcement learning; predicting a fault risk score through a lightweight prediction model deployed at each computing node; and a deep Q network is adopted as a model architecture, a state vector and a fault risk score are input, reinforcement learning training is performed through a reward function in a multi-target vector form, a final scheduling model is obtained, and a task allocation decision is output. The problems that in the prior art, the hardware state cannot be sensed in real time, hardware characteristic matching is ignored, consequently, the computing resource utilization rate is insufficient, and fault recovery is passive are solved.
Owner:SHANDONG ZHIYANG ELECTRIC

Persistent Artificial Intelligence Agents

Techniques include generating, by an artificial intelligence agent, a request for artificial intelligence data. The techniques further include presenting the request via a user interface. The techniques further include receiving the artificial intelligence data. The techniques further include providing the artificial intelligence data to the artificial intelligence agent for at least one of training or inference.
Owner:AVAULTI INC

Large language model accelerator architecture based on three-dimensional NAND flash memory

The invention discloses a large language model accelerator architecture based on a three-dimensional NAND flash memory, and belongs to the technical field of calculation, reckoning or counting. The architecture comprises a three-dimensional NAND flash memory used for executing feedforward neural network calculation; the auxiliary calculation unit is used for executing attention mechanism calculation; the DRAM chip is used for storing attention mechanism related weights and KV cache; and the interconnection resource is used for realizing data interaction among the components. Wherein the three-dimensional NAND flash memory comprises a logic chip and an NAND array chip, the logic chip is used for controlling and executing calculation, and the NAND array chip is used for storing weights and participating in calculation. The invention further provides a scheduling method based on KV cache awareness. The scheduling method comprises decomposition and dynamic allocation of calculation tasks. Through collaborative design of the hardware architecture and the scheduling method, the memory wall bottleneck in large language model calculation can be relieved, the energy consumption is reduced, and the overall calculation performance is improved.
Owner:SOUTHEAST UNIV

Dynamic agents with real-time alignment

An example may receive at least one input via at least one device. An example may use the at least one input to determine an objective. An example may use the objective, a multi-agent system, and an automated agent to cause at least one first sub-agent of the multi-agent system to generate and execute a first plan including one or more tasks to achieve the objective. An example may cause at least one second sub-agent of the multi-agent system to execute a second plan to supervise the at least one first sub-agent in accordance with a supervision level that indicates a level of supervision of the automated agent by an entity associated with the objective.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Neural network large model efficient reasoning method based on multiple GPGPUs

The invention belongs to the technical field of artificial intelligence and high-performance computing, and particularly relates to a neural network large model efficient reasoning method based on multiple GPGPUs. The method aims to solve the problems of high communication overhead, non-uniform load, low resource utilization rate, high data transmission delay and the like among multiple processors. Dividing a calculation task into a plurality of sub-graphs through static analysis and mixed granularity partitioning of a model calculation graph; distributing the sub-graphs to the optimal GPGPU based on a weighted cost function in combination with heterogeneous resource perception and a dynamic mapping strategy; a global pipeline scheduling plan is constructed by using communication topology perception, and calculation and communication overlap are maximized; data are loaded in advance through a host side hierarchical caching and asynchronous prefetching mechanism, and transmission delay is hidden; multi-stream concurrent execution and event-based lightweight synchronization are adopted on each GPGPU, so that waiting overhead is reduced. According to the method, the reasoning delay can be remarkably reduced, the throughput and the hardware utilization rate are improved, and the method has good adaptivity and expandability.
Owner:BEIJING TOPMOO TECH

Neural network using dynamically compressed and decompressed weights

A method for training or performing inference using a neural network involves performing per-layer decompression and compression of neural network weights. More particularly, compressed weights are retrieved for a particular layer of the neural network. The weights correspond to neurons in the layer. The compressed weights are decompressed, and input data for that layer is subsequently processed using the decompressed weights. This dynamic decompression and recompression of weights allows memory, and in particular random access memory of graphical processing units, to be efficiently used.
Owner:ROYAL BANK OF CANADA

User Behavior Modeling for Detecting and Containing Malicious Activity in a Storage System

An illustrative method includes monitoring operations performed with respect to a storage system by an entity using an identity associated with a particular role, the particular role providing the entity with a set of permissions associated with the storage system; determining, based on the monitoring, that one or more operations of the operations deviate from an expected activity profile associated with the role by more than a threshold; and performing, based on the determining that the one or more operations deviate from the expected activity by more than the threshold, a remedial action with respect to the entity.
Owner:PURE STORAGE INC

Multi-thread low-power-consumption intelligent monitoring system based on AI processor

The invention relates to the technical field of intelligent monitoring, in particular to a multi-thread low-power-consumption intelligent monitoring system based on an AI processor. The method has the advantages that aiming at the problems of unbalanced computing power and power consumption, high multi-task processing delay and strong hardware dependence in the prior art, the NPU module of the RK3588 processor is combined with the INT8 quantitative model, so that the power consumption is lower than 10W under the 6TOPS computing power; a dynamic multi-thread scheduling mechanism is designed, parallel processing of more than eight paths of video streams is supported through binding of a priority queue and an NPU core, and end-to-end delay is compressed to be within 200 ms; a zero-copy video stream architecture is constructed, data transfer is eliminated through memory mapping, and preprocessing time consumption is reduced by 90%; an energy efficiency control module is integrated, the NPU voltage frequency is dynamically adjusted according to the load, and the energy efficiency ratio reaches 0.83 TOPS / W; space-time alignment of multi-model reasoning results is realized by adopting a frame ID synchronization technology, and the mismatching rate is lower than 0.1%.
Owner:FOCALCREST LTD

Neural network inference circuit with piecewise linear activation circuit

Some embodiments provide a neural network inference circuit for executing a neural network that includes computation nodes. Each respective computation node of a set of the computation nodes includes (i) a respective linear function that includes a respective dot product of input values for the computation node and weight values for the computation node and (ii) a respective non-linear activation function. The neural network inference circuit includes a set of dot product circuits to compute the dot product for a computation node and a post-processing circuit to compute (i) a result of the linear function for the computation node based on the dot product for the computation node and (ii) an output for the computation node by applying a piecewise linear function to the result of the linear function for the computation node to apply the non-linear activation function for the computation node.
Owner:AMAZON COM SERVICES LLC

Machine learning model search method, related apparatus, and device

This application relates to the field of artificial intelligence technologies, and discloses a machine learning model search method, a related apparatus, and a device. In the method, before model search and quantization, a plurality of single bit models are generated based on a to-be-quantized model, and evaluation parameters of layer structures in the plurality of single bit models are obtained. Further, after a candidate model selected from a candidate set is trained and tested, to obtain a target model, a quantization weight of each layer structure in the target model may be determined based on a network structure of the target model and evaluation parameters of all layer structures in the target model, a layer structure with a maximum quantization weight in the target model is quantized, and a model obtained through quantization is added to the candidate set.
Owner:HUAWEI TECH CO LTD

Hybrid neural architecture for data processing combining matmul-free techniques and spiking neural networks

A hybrid neural network architecture is disclosed that integrates matrix multiplication-free (MatMul-free) transformation layers with spiking neural network (SNN) layers for efficient, low-power computation. The system includes an interface module configured to convert intermediate continuous-valued data from MatMul-free layers into a spike-compatible format using encoding techniques such as rate coding, phase coding, or threshold-based conversion. The SNN layers process the spike-encoded data in an event-driven manner, enabling sparse, temporal inference. Training is supported by a hybrid optimization strategy combining backpropagation in MatMul-free components with surrogate gradient descent or spike-timing-dependent plasticity (STDP) in SNN layers. The architecture reduces computational complexity, supports real-time adaptability, and enables deployment in energy-constrained environments such as edge devices and neuromorphic platforms. The system may be implemented in hardware, software, or a co-designed pipeline optimized for dynamic sensor data, control signals, or continuous inference tasks.
Owner:LEPTUDE INC

Body multi-agent self-adaptive cooperation method and system based on large language model

The invention discloses a multi-agent self-adaptive cooperation method and system based on a large language model, and belongs to the technical field of artificial intelligence, and the method comprises the steps: obtaining task description and environment information, and inputting the task description and environment information into the large language model; generating a personalized initial instruction corresponding to the multiple agents; constructing a multi-agent topological state space, and establishing a relation topological structure among the agents; based on a distributed negotiation mechanism, determining an action set through natural language interaction; generating a self-adaptive cooperation strategy according to the action set and the current environment state; the multiple agents execute corresponding actions and obtain feedback, and a cooperation strategy is dynamically adjusted; according to the method, the semantic understanding capability of a large language model is innovatively combined with the technologies of topological state mapping, distributed negotiation and the like, seamless conversion from high-level semantic understanding to specific execution actions is achieved, and the method has the advantages of being high in practicability and high in practicability. And the environmental adaptability, the cooperation efficiency, the safety and the reliability of the multi-agent system are greatly improved.
Owner:GUANGDONG UNIV OF PETROCHEMICAL TECH

Attention mechanism calculation method and device, storage medium and product

The invention discloses an attention mechanism calculation method and device, a storage medium and a product, and the method comprises the steps: carrying out matrix multiplication operation through employing a query matrix block of a first register block and a key matrix block of a shared memory, obtaining a first product matrix block, and writing the first product matrix block into a second register block; performing exponential operation by using the first product matrix blocks to obtain sub-matrix blocks, writing the sub-matrix blocks into a second register block in a covering manner, and writing the sub-matrix blocks into a third register block in the form of a target precision type; performing matrix multiplication operation by using the sub-matrix blocks of the third register block and the value matrix blocks of the shared memory to obtain second product matrix blocks, and writing the second product matrix blocks into a second register block; performing softmax operation by using the second product matrix blocks to obtain attention result matrix blocks, and writing the attention result matrix blocks into a fourth register block; and writing the attention result matrix of the fourth register group into the shared memory in blocks. According to the embodiment of the invention, overflow of the register can be avoided, and the utilization of hardware resources is maximized.
Owner:SHANGHAI BIREN TECH CO LTD

System and method for persistent cognitive machines using a digital thought architecture

A system and method for implementing a Persistent Cognitive Machine (PCMs) that extends beyond the traditional prompt-response paradigm of artificial intelligence are disclosed. A PCM maintains persistent cognitive processes regardless of external interaction, stores and organizes thoughts in a thought cache, retrieves relevant thoughts based on current stimuli, generates new thoughts through reasoning processes, and curates stored thoughts during periods of reduced external interaction. The PCM includes language and reasoning model components, a thought cache, an executive component, and an embedding system. The PCM remains continuously active, remembers previous experiences, learns from these experiences, creates new thought experiences independently, and initiates interactions without waiting for external prompts. The PCM enters sleep-like states during which it curates its thought cache, generalizes experiences, and performs other memory management functions. Applications may include but are not limited to synthetic cognitive colleagues, strategic war gaming platforms, and personal cognitive assistants.
Owner:ATOMBEAM TECH INC

Sub-cell, MAC array and bit-width reconfigurable mixed-signal in-memory computing module

A mixed-signal in-memory computing sub-cell requires only 9 transistors for 1-bit multiplication. In one aspect, a computing cell is constructed from a plurality of such sub-cells that share a common computing capacitor and common transistors. As a result, the average number of transistors in each sub-cell is close to 6. Also proposed is a MAC array for performing MAC operations, which includes a plurality of the computing cells each activating the sub-cells therein in a time-multiplexed manner. Also proposed is an in-memory mixed-signal computing module for digitalizing parallel analog outputs of the MAC array and for performing other tasks in the digital domain.
Owner:REEXEN TECH CO LTD

Attention-based video token generation

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for generating a video output using an autoregressive token generation neural network model In one aspect, a system comprises obtaining a model input, processing the model input to generate an input sequence of embeddings that represents the model input, autoregressively generating a plurality of output sequences of tokens, wherein each output sequence of tokens corresponds to a respective output modality of tokens from a set of a plurality of modalities that includes a video modality and one or more other modalities, and generating a model output that includes a video output of the video modality by decoding the sequence of tokens.
Owner:GOOGLE LLC

Multi-modal open intention recognition method and system based on pellet characterization

The invention discloses a multi-modal open intention recognition method and system based on pellet characterization, and belongs to the technical field of artificial intelligence and multi-modal intention understanding, and the method comprises the steps: carrying out the feature extraction and modal fusion of multi-modal input data; carrying out structural modeling on the feature representation of each mode and the fusion mode through an adaptive particle and ball clustering method, and generating a multi-granularity particle and ball set; the mass centers of the pellets serve as multi-granularity anchor points, and the pellets with the same labels in different modalities are aligned; introducing a weighting mechanism based on purity and sample scale into the fusion mode; generating a boundary-constrained pseudo-distribution outer sample in the fusion modal space; and constructing a self-adaptive decision boundary based on the fusion modal particle ball obtained by training, and carrying out known class classification and unknown class detection. According to the invention, by introducing the multi-granularity anchor point and the structure perception particle-ball representation mode, the joint recognition of the known category and the unknown category in the multi-modal scene is realized, and the accuracy and robustness of intention recognition are remarkably improved.
Owner:SOUTHWESTERN UNIV OF FINANCE & ECONOMICS

Core AI Serving Platform Enhancements

A computer system implements a unified framework integrating an adaptive elastic funnel (AEF) with a convergent intelligence fabric (CIF) for flexible and contextualized multi-agent AI and human collaboration at scale. The system provides a universal multi-modal key-value subsystem for sharing partial computations across agents, implements a hybrid greedy / non-greedy placement strategy for dynamic memory management, orchestrates dynamic computational workflows and tensor workflows using hierarchical tensor-fragment scheduling, enables cross-agent orchestration with policy-based privacy preservation, and incorporates quantum-resistant secure memory enclaves. The architecture supports continuous learning without catastrophic forgetting, compositional reasoning across modalities, and secure task execution in distributed environments. This integration enables unprecedented computational efficiency, secure collaboration, and adaptive intelligence in high-dimensional decision-making environments while supporting incremental adoption through modular interfaces.
Owner:QOMPLX INC

Attention mechanism calculation method and device, medium and product

The invention discloses an attention mechanism calculation method and device, a medium and a product. The method comprises the steps that a second thread bundle group is controlled to load an ith query matrix block; controlling the second thread bundle group to perform matrix multiplication operation and exponential operation by using the ith query matrix block and the transposed jth key matrix block to obtain a jth attention score matrix block; and controlling the first thread bundle group and the second thread bundle group to alternately use different sub-blocks of the jth value matrix block to perform attention mechanism operation of the jth attention score matrix block until the last sub-block of the jth attention result matrix block is obtained. By adopting the embodiment of the invention, sufficient register resources can be provided for the calculation of the attention mechanism, so that the calculation efficiency is improved.
Owner:SHANGHAI BIREN TECH CO LTD

Non-linear data generation and processing method based on light quantum and light quantum computer

The invention provides a nonlinear data generation and processing method based on light quantum and a light quantum computer, and the method comprises the steps: obtaining a nonlinear function and attribute information of a nonlinear layer; according to the frequency spectrum number and the attribute information of the nonlinear function, constructing an initial light quantum circuit; measuring the initial light quantum circuit, and determining a target function corresponding to the initial light quantum circuit according to a measurement result; loss information of the target function is determined, whether the initial light quantum circuit is applied or not is determined according to the loss information, and if yes, the initial light quantum circuit serves as a target light quantum circuit; and performing operation on input data of the nonlinear layer based on the target light quantum circuit. According to the method, the low-energy-consumption characteristic of photon calculation, the nonlinear characteristic of light quantum calculation and the general programmable capability of electronic calculation can be fused, and high-energy-efficiency calculation can be carried out on input data of a nonlinear layer. In addition, the method also has the advantages that the complexity of a quantum circuit is smaller, and the participation of a nonlinear optical device is not needed.
Owner:SHANGHAI TURING INTELLIGENT COMPUTING QUANTUM TECHNOLOGY CO LTD

Address decoding by neural network inference circuit read controller

Some embodiments provide a neural network inference circuit (NNIC) for executing a neural network (NN) that includes computation nodes at multiple layers. The NNIC includes a set of processing circuits for executing the computation nodes of the NN, a set of memories for storing data used by the processing circuits to execute the NN layers, and a read controller for retrieving the data from the memories for use by the processing circuits. The data is stored in the memories as multiple varying-size blocks. The read controller receives read instructions for a requested block of data to be used by the processing circuits for one or more computation nodes. The read instructions include a base memory address for multiple blocks of data, a size of the requested block of data, and a location of the requested block of data within the multiple blocks of data.
Owner:AMAZON COM SERVICES LLC

Neural network training in a distributed system

In one example, a method comprises: performing backward propagation computations for a second layer of a neural network to generate second weight gradients; splitting the second weight gradients into a plurality of subsets each associated with a second exchange operation over a computer network, a number of the second weight gradients included in the each subset being based on at least one of first characteristics of the computer network or second characteristics of the neural network; performing backward propagation computations for a first layer of the neural network to generate first weight gradients in parallel with at least one of the second exchange operations; performing a first exchange operation to exchange the first weight gradients after the at least one of the second exchange operations completes; and after the first exchange operation completes, perform the remaining second exchange operations to exchange the remaining subsets of the second weight gradients.
Owner:AMAZON TECH INC