Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

56 results about "Massively parallel" patented technology

In computing, massively parallel refers to the use of a large number of processors (or separate computers) to perform a set of coordinated computations in parallel (simultaneously). In one approach, e.g., in grid computing the processing power of many computers in distributed, diverse administrative domains, is opportunistically used whenever a computer is available. An example is BOINC, a volunteer-based, opportunistic grid system, whereby the grid provides power only on a best effort basis.

Inspection data processing method and system

The invention discloses an inspection data processing method and system, relates to the technical field of data processing, and ensures the accuracy and stability of authentication scheduling in a multi-interface cluster environment by responding to concurrent tasks and dynamically selecting authentication tokens according to interface features. The acquisition request is constructed based on the interface characteristics, so that the adaptive access to the heterogeneous interface path is realized, and the compatibility of large-scale parallel acquisition is guaranteed; by uniformly converting original data into a structured format and carrying out anomaly detection and linked list storage, the problems that multi-source heterogeneous data is difficult to process and an abnormal state is difficult to systematically track are solved; and finally, a standardized report is automatically generated according to a preset rule, and accurate marking of abnormal information is realized, so that full-process automation from data acquisition to report output is completed, and the reliability, the processing efficiency and the intelligent level of inspection operation are remarkably improved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Action-based graph framework and query method

Disclosed is a network and method for creating and searching a graph in a massively parallel manner. The network includes a plurality of graph storage instances that collectively store a graph formed from a plurality of entities and comprising a vertex for each entity, edges connecting pairs of vertices, and adjacency relations. Each graph storage instance stores a subgraph of the graph and comprises a partition with a non-overlapping vertex set of all the vertices in the respective subgraph, an edge set of all the edges between vertices in the respective vertex set, and the adjacency relations for the edge set and to any edges outside the partition, to which the edge set connects. Each graph storage instance also includes an executor for executing one or more of the actions over the respective subgraph, and an action monitor for supplying actions to the respective executor.
Owner:GRABTAXI HOLDINGS PTE LTD

A sparse attention calculation method, device and medium for a GPU

The present application relates to the technical field of GPU computing optimization, and in particular to a sparse attention computing method, device and medium for GPU, wherein the method realizes the high efficiency of long context reasoning through the geometric perception sparse attention framework of ball hashing, and combines a large-scale parallel hashing optimization algorithm and a load adaptive computing kernel. Compared with the existing sparse attention methods based on heuristics or gradient learning, the present application realizes higher retrieval recall rate, lower preprocessing overhead and efficient hardware adaptation to irregular sparse patterns.
Owner:CENT SOUTH UNIV

A parallel situation gravity calculation chip and method based on in-memory computing

PendingCN122509213AEliminate bottlenecksreduce power consumptionMassively parallelCrossbar switch
This invention discloses a parallel situational gravity calculation chip and method based on in-memory computing, belonging to the field of artificial intelligence chips and analog in-memory computing technology. The chip includes: an event polarity input unit, which converts externally input event information loads (six-dimensional polarity vectors) into analog voltage signals; a situational reference storage array, employing a memristor cross-switch array, with each row storing a six-dimensional binary encoded vector of a reference situational type (such as the 64 hexagrams); a parallel gravity calculation unit, which performs large-scale parallel multiply-accumulate operations simultaneously in the analog domain, following the gravity calculation rule of "like lines cooperate, dissimilar lines repel," generating a multi-dimensional gravity intensity distribution; and a gravity readout and post-processing unit, which converts the gravity distribution into situational awareness results and outputs them. This invention adopts an in-memory computing architecture, eliminating the "memory wall" bottleneck of the traditional von Neumann architecture, achieving nanosecond-level fully parallel gravity calculation and ultra-low power consumption operation. It can be deployed as an independent situational awareness processor at sensor terminals, IoT terminals, and wearable devices.
Owner:PUTIAN ZIXU LIFE TECHNOLOGY CO LTD

Massively parallel on-chip coalescence of microemulsions

Embodiments disclosed herein are directed to microfluidic devices that allow for scalable on-chip screening of combinatorial libraries and methods of use thereof. Droplets comprising individual molecular species to be screened are loaded onto the microfluidic device. The droplets are labeled by methods known in the art, including but not limited to barcoding, such that the molecular species in each droplet can be uniquely identified. The device randomly sorts the droplets into individual microwells of an array of microwells designed to hold a certain number of individual droplets in order to derive combinations of the various molecular species. The paired droplets are then merged in parallel to form merged droplets in each microwell, thereby avoiding issues associated with single stream merging. Each microwell is then scanned, e.g., using microscopy, such as high content imaging microscopy, to detect the optical labels, thereby identifying the combination of molecular species in each microwell.
Owner:MASSACHUSETTS INST OF TECH +1

Efficient coupling parallel method for dense matrix and sparse matrix

The invention discloses a dense matrix and sparse matrix efficient coupling parallel method. The method comprises the following steps: S1, carrying out feature analysis on an input dense matrix and sparse matrix; s2, carrying out adaptive partitioning on the dense matrix and the sparse matrix; s3, creating a uniform matrix descriptor for each matrix block; s4, constructing a calculation dependency graph; s5, executing parallel computing; according to the dense matrix and sparse matrix efficient coupling parallel method, through data rearrangement of self-adaptive hybrid storage and cache perception, the cache hit rate is greatly increased, the bottleneck of memory access mode conflicts in hybrid calculation is overcome, and meanwhile, based on an accurate prediction model and dynamic scheduling, the performance data in the running process are collected. The method has the advantages that combined optimization of computing load and communication is achieved, high load balance and resource utilization rate under large-scale parallel are guaranteed, redundant format conversion overhead is almost eliminated by the aid of unified descriptors and inert conversion strategies, a computing pipeline is smoother, and overall performance is improved remarkably.
Owner:NAT SUPERCOMPUTING WUXI CENT

Large-scale parallel sampling environment construction method and device for reinforcement learning of intelligent computing cloud platform through computing power

The invention provides a large-scale parallel sampling environment construction method and device for reinforcement learning of an intelligent computing cloud platform through computing power, and the method comprises the steps: obtaining a plurality of physical disks of a plurality of computing power nodes of the intelligent computing cloud platform, and constructing the plurality of physical disks into a bottom storage pool through an RAID technology, storing the basic mirror image of the operating system in a qcow2 format, and further generating a plurality of virtual machine instances and running the virtual machine instances; packaging the QEMU process and the running environment of the QEMU process into a container, and performing distributed scheduling on the container according to the disk topological relation among the plurality of bottom storage pools on the plurality of computing power nodes and the real-time I / O load condition. Therefore, the construction of a large-scale parallel sampling environment of reinforcement learning can be realized by utilizing the computing power of an intelligent computing cloud platform and integrating storage optimization, hardware virtualization and container arrangement technologies.
Owner:DATACANVAS LTD

Massively parallel cell analysis and sorting apparatus and method

A massively parallel microfluidic chip is provided having multiple sections stacked or layered along a stacking direction to form multiple microchannels at least partially oriented for flow along the stacking direction. The multiple sections can include a transfer section for introducing a sample fluid containing particles, a particle focusing section configured to focus particles in the sample fluid, and an actuation section including multiple interrogation regions and multiple actuators. Each interrogation region and actuator is associated with at least one microchannel within the multiple microchannels. The arrangement of the microfluidic channels along the stacking direction allows for a very high packing density of channels and interrogation regions on a single chip, providing massively parallel processing of particles.
Owner:SITE NOME ESTÉE LELSEY

Method and device for predicting performance of application program in heterogeneous system

The invention discloses an application program performance prediction method and device under a heterogeneous system, and belongs to the technical field of system performance prediction.The method comprises the steps that S100, small-scale programs are extracted from large-scale parallel programs to be predicted; s200, running the small-scale program on a target heterogeneous system, and collecting performance data related to a hardware architecture; s300, constructing a performance model based on the performance data; s400, training the performance model by adopting a staged migration training strategy; and S500, inputting the program parameters of the large-scale parallel program into the performance model, and outputting a performance trend prediction result on the current heterogeneous system. According to the method, large-scale performance prediction with low cost, high precision and strong extrapolation is realized through a representative small-scale program, a dual-channel neural network and staged migration training.
Owner:COMP NETWORK INFORMATION CENT CHINESE ACADEMY OF SCI

Large-scale parallel simulation method based on in-situ decentration field data aggregation, processing architecture and storage medium

The invention discloses a large-scale parallel simulation method based on in-situ decentralized field data aggregation, a processing architecture and a chip, and relates to the field of high-performance computing and artificial intelligence. The architecture comprises a memory coupling array, a data distribution network and a parallel sampling interface. Wherein the storage and calculation array is integrated with an analog resistance network field generation layer, and global field calculation of O (1) complexity is realized by using the Kirchhoff's law; a hierarchical dynamic bypass bus (VST) is integrated, so that zero-delay jump transmission of sparse data is realized; a congestion sensing type asynchronous clock island is integrated, and the relativistic time expansion effect is simulated in a hardware layer to balance loads. According to the method, complex vector mechanics and interactive calculation are converted into a linear signal processing process through the in-situ polymerization characteristic of a physical storage array, the computing power and bandwidth bottlenecks in large-scale multi-body simulation are solved, and the method can be widely applied to large-model (LLM) reasoning, automatic driving point cloud processing and neuromorphic calculation. And the system has extremely high energy efficiency ratio and universality.
Owner:王俊鹏

Compression of bitstream indexes for wide scale parallel entropy coding in neural-based video codecs

Systems and techniques are described herein for processing video data. For example, an encoding device can obtain a sequence of video data and determine a minimum value in the sequence of video data. The encoding device can, based on the minimum value, identify positions in the sequence of video data associated with entry points for individually entropy codable parcels of a parallel entropy codable sequence of video data. The encoding device can generate the parallel entropy codable sequence of video data. The encoding device can further generate an index for the parallel entropy codable sequence of video data, the index identifying the individually entropy codable parcels within the parallel entropy codable sequence of video data.
Owner:QUALCOMM INC

Error correction code decoding method, test system and computer program product

The invention discloses an error correction code decoding method, a test system and a computer program product, and relates to the field of storage firmware error correction, the method is applied to a large-scale parallel processor, and the method comprises the following steps: receiving a coding bit of an error correction code to be decoded; the coding bits are obtained by carrying out coding modulation on original data used for simulation testing by the host end; distributing the coded bits to a plurality of stream processors, and using the plurality of stream processors to execute decoding operation on the coded bits in parallel; in any stream processor, a computing kernel is distributed for the processing stage of the decoding operation, and the computing kernel is used for executing the processing task of the corresponding processing stage in parallel. A multi-level task parallel processing mode of inter-stream parallel and intra-stream parallel is adopted, and a parallel architecture of a large-scale parallel processor is fully utilized for decoding, so that the decoding performance simulation time consumption of error correction codes is reduced, and the simulation efficiency is improved.
Owner:HANGZHOU CORE POWER SEMICON CO LTD

Method for extracting information from an unstructured data source

A method includes extracting information from an unstructured data source, the method including: scraping, by at least one processor, a plurality of texts from the unstructured data source, extracting, by the at least one processor, from the plurality of texts a chunk of relevant text, summarizing, by the at least one processor, using a pre-trained summarizer, the chunk of relevant text to obtain semi-structured information comprising a set of sentences that summarize the chunk of relevant texts, and postprocessing, by the at least one processor, the semi-structured information to obtain structured information. The method can be executed highly efficiently, in particular on massively parallel hardware.
Owner:INTAPP GERMANY GMBH

Sparse attention calculation method and device for GPU and medium

The invention relates to the technical field of GPU calculation optimization, in particular to a sparse attention calculation method and device for a GPU and a medium, and the method realizes high efficiency of long context reasoning through a geometric perception sparse attention framework of ball hash in combination with a large-scale parallel hash optimization algorithm and a load adaptive calculation kernel. Compared with the existing sparse attention method based on heuristic or gradient learning, the method provided by the invention realizes higher retrieval recall rate, lower preprocessing overhead and efficient hardware adaptation to the irregular sparse mode.
Owner:CENT SOUTH UNIV

A numerical simulation method and device for particle collision retrieval in dense gas-solid flow based on GPU parallel computing

The application discloses a numerical simulation method and device for particle collision retrieval in dense-phase gas-solid flow based on GPU parallel calculation, realizes large-scale parallel solution of a neighborhood particle collision retrieval algorithm in dense-phase gas-solid flow numerical simulation research on a computer graphics processing unit (GPU), realizes a parallel prefix sum algorithm in combination with CUDA programming and OpenACC instructions, writes particle collision retrieval results into an array, establishes a particle collision list, obtains dense-phase particle collision retrieval results, and is used for calculating collision forces between particles. The application overcomes the deficiency that a great amount of calculation is consumed in particle neighborhood collision retrieval by a traditional discrete element method, and fills the blank of the discrete element method in efficient processing of large-scale particle collision information.
Owner:ZHEJIANG UNIV +1

High-precision semiconductor test method and system for multi-channel parallel test

The invention relates to a high-precision semiconductor test method and system for multi-channel parallel test. The method comprises the steps of generating a multi-dimensional orthogonal coding calibration signal containing multi-domain information based on a multi-site test requirement of a to-be-tested chip; generating a corrected multi-domain parameter through inherent deviation elimination processing and multi-domain parameter error correction based on Fourier transform relevance; and combining all the corrected multi-domain parameters into a multi-domain correction parameter set, and correcting the real-time test data based on the parameter set to obtain a high-precision test result. By adopting the method, the problems of positioning error and parameter inconsistency caused by multiple physical contacts in the traditional step-by-step calibration can be effectively avoided, the calibration efficiency and parameter self-consistency are greatly improved, the wafer scratch risk is reduced, and a high-precision and high-consistency calibration basis is provided for large-scale parallel test of high-performance semiconductors.
Owner:山西科技学院

Method and system for parallel processing of a software program consisting of a plurality of software blocks identified by function identifiers

A method for massively parallel processing of a software program, which includes a plurality of software blocks, using execution engines orchestrated by a computing orchestrator. The method includes: configuring the software program which involves incorporating software block start markers being a function identifier; launching on a first machine, using a first a first execution engine, by the computing orchestrator, to have it execute at least two software blocks, launching including transmitting the function identifier of the software blocks and input parameters; executing, by the first execution engine, a first software block; instantiating at least one second execution engine to have it execute at least one second software block, including transmitting the function identifier of the second of the software blocks; executing, by the second execution engine, the second software block, and receiving, by the first execution engine, execution status signals produced by the second execution engine.
Owner:MELODIUM

MASSIVELY PARALLEL NEURONAL INFERENCE DATA PROCESSING ELEMENTS

System, exhibiting: a plurality of multipliers (206; 302; 504; 1202), wherein the plurality of multipliers is arranged in a plurality of equally sized groups, each of the plurality of multipliers being designed to apply a weighting to an input activation in parallel to produce an output; a plurality of adders (204; 304; 506; 1204), each of the plurality of adders being operatively connected to one of the groups of multipliers, each of the plurality of adders being designed to add the outputs of the multipliers within their respective groups in parallel to produce a partial sum (306); a first plurality of function blocks (508; 1206), wherein each from the first plurality of function blocks is operatively connected to one from the plurality of adders, wherein each from the first plurality of function blocks is designed to apply a function in parallel to the partial sum of its associated adder in order to produce an output value; a vector register (116), wherein the vector register is operatively connected to the first plurality of function blocks, wherein the vector register is designed to store the output values ​​of the first plurality of function blocks, wherein the first plurality of function blocks is designed to combine the output values ​​stored in the vector register with subsequently calculated output values ​​of the first plurality of function blocks, wherein output values ​​of this combination are stored in the vector register; a second plurality of function blocks, each of which is operatively connected to the vector register, each of which is designed to apply a function to the stored output values ​​in parallel.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Performing hyperparameter tuning of models in a massively parallel database system

Hyperparameter tuning for a machine learning model is performed in a massively parallel database system. A computer system comprised of a plurality of compute units executes a relational database management system (RDBMS), wherein the RDBMS manages a relational database comprised of one or more tables storing data. One or more of the compute units perform the hyperparameter tuning for the machine learning model, wherein the hyperparameters are control parameters used in construction of the model, and the tuning of the hyperparameters is implemented as an operation in the RDBMS that accepts training and scoring data for the model, constructs the model using the hyperparameters and the training data, and generates goodness metrics for the model using the scoring data.
Owner:TERADATA US INC

Massively parallel amplitude-only optical processing system and methods for machine learning

ActiveUS12718323B2Massively parallelAlgorithm
Amplitude-only Fourier optical processors is capable of processing large-scale matrices in a single time-step and microsecond-short latency. The processors may have a 4f optical system architecture and may employ reprogrammable high-resolution amplitude-only spatial modulators, such as Digital Micromirror Devices (DMD). In addition, methods are provided for obtaining amplitude-only electro-optical convolutions between large matrices displayed by the DMDs. The large matrices on which convolution is performed may be feature maps corresponding to images and kernel matrices used in neural networks classification systems. Analog optical convolutional neural networks are also provided that perform accurate classification tasks on large matrices. In addition, methods are provided for off-chip training the analog optical convolutional neural networks. The training includes building an accurate physical model for the analog optical processor and performing computer simulations of the optical processor according to the physical model. The methods do not need to employ any interferometric scheme.
Owner:GEORGE WASHINGTON UNIVERSITY

Real-time path planning system and method for mechanical arm in unstructured environment

The invention discloses a real-time path planning system and method for a mechanical arm in an unstructured environment, and the method comprises the steps: constructing a large-scale parallel simulation training scene based on a GPU physical engine, and creating a parallel simulation system comprising thousands of mechanical arm examples through an example cloning technology; the method comprises the following steps: realizing extraction and dimension reduction of three-dimensional environment features based on a variational auto-encoder technology, and compressing a high-dimensional depth image into a low-dimensional implicit feature vector; a time sequence deep reinforcement learning algorithm fusing course learning and a self-attention mechanism is adopted, and a smooth and stable real-time path planning strategy is generated through fusion of a historical observation sequence and a mechanical arm dynamic state. According to the method, the path planning success rate of 95.2% is achieved in the unstructured environment containing 3-5 dynamic obstacles, the average planning time is 0.12 s, and the autonomous operation capacity of the mechanical arm in complex scenes such as agricultural picking, mine rescue and electric power overhaul is remarkably improved.
Owner:SUZHOU AMIFULUI ROBOT TECHNOLOGY CO LTD

Lake warehouse integrated data processing system and electronic equipment

The embodiment of the present application provides a lake warehouse integrated data processing system, relates to the field of big data processing, and data is accessed from a plurality of heterogeneous data sources by a data access layer; an object storage unit persistently stores initial data accessed from the data access layer in a raw data format, and constitutes a raw data area; a core storage engine unit adopts a large-scale parallel processing architecture, and manages a refined data area through an incremental data snapshot mechanism; a unified metadata service uniformly manages metadata, and records a derivative relationship and a structure mapping; a data processing layer extracts data from the raw data area, carries out conversion and loading operation, and writes processed data into the refined data area, and is integrated with a data quality checking module; a data service layer provides a unified data access interface, and can query and analyze data in the raw data area and the refined data area. The present application effectively solves problems such as data islands, performance bottlenecks, and management difficulties in the prior art.
Owner:XIAMEN MEIYABAIKE INFORMATION SECURITY RES INST CO LTD

Automated setup and communication coordination for training and utilizing massively parallel neural networks

A method is disclosed for training and utilizing massively parallel neural networks. A distributed computing system may be configured to perform various operations. The distributed computing system may divide a directed acyclic graph (“DAG”) that comprises a plurality of vertices linked in pairwise relationships via a plurality of edges among a plurality of nodes. Each node may comprise a computing device. The distributed computing system may provide a map of the DAG that described a flow of data through the vertices to each of the vertices of the DAG. The distributed computing system may perform a topological sort of the vertices of the DAG and may traverse the DAG.
Owner:FORD GLOBAL TECH LLC

Fast carry-calculation oriented redundancy-tolerated fixed-point number coding for massive parallel ALU circuitry design in GPU, TPU, NPU, AI infer chip, CPU, and other computing devices

A code method, a computer program product, and a system, for implementing a code method of Redundancy-Tolerated symmetric binary Coding (RTC) for massive parallel ALU circuitry design in GPU (Graphics Processing Unit), TPU (Tensor Processing Unit), NPU (Neural Processing Unit), Artificial Intelligence Inference Chip, CPU, and other computing chips and devices. The method can remove redundancy and reduce the dependency on carry bit computing. RTC code method can provide a redundancy-tolerated digital coding method for negative and positive integers, and can guarantee that number “0” have one and only one representation.
Owner:ZHOU JUN

Error correction code decoding method, test system, and computer program product

The application discloses an error correction code decoding method, a test system and a computer program product, relates to the field of storage firmware error correction, and is applied to a large-scale parallel processor and comprises the following steps: receiving encoding bits of error correction codes to be decoded; the encoding bits are obtained by encoding and modulating original data for simulation test on a host side; the encoding bits are distributed to a plurality of stream processors; the decoding operation on the encoding bits is performed in parallel by using the plurality of stream processors; and in any stream processor, a calculation kernel is respectively distributed to a processing stage of the decoding operation, and the processing task of the corresponding processing stage is performed in parallel by using the calculation kernel. By adopting the multi-level task parallel processing mode of inter-stream parallel and intra-stream parallel, the decoding is performed by fully utilizing the parallel architecture of the large-scale parallel processor, the time consumption of the decoding performance simulation of the error correction code is reduced, and the simulation efficiency is improved.
Owner:HANGZHOU CORE POWER SEMICON CO LTD

Multi-radar cooperative detection mutual interference calculation method based on GPU

The invention discloses a multi-radar cooperative detection mutual interference calculation method based on a GPU, and the method comprises the steps: reading transmitted signals of different radars from a radar system and arrival signals received by each radar, dividing the number of thread blocks included in a two-dimensional grid for echo and mutual interference calculation according to the total number of radars, the number of transmitted pulse sequences, the number of targets and the number of mutual interference paths, and calculating the number of the thread blocks included in the two-dimensional grid; and determining the number of threads included in each thread block. A GPU kernel function is called at the host end to start calculation, and each thread in the grid calculates relevant parameters of echo signals or mutual interference signals. According to the method, the powerful parallel computing capacity of the GPU is fully utilized, traditional CPU serial processing is converted into large-scale parallel processing, the processing speed of mass data in multi-radar cooperative detection is remarkably improved, the problems that in a CPU computing method, computing time is long, real-time performance is poor, and hardware resources are not fully utilized are effectively solved, and the method is suitable for mass data processing. And efficient and real-time radar mutual interference analysis and signal processing are realized.
Owner:XIDIAN UNIV

Device for dividing three regions after power failure of power grid based on high-order Isin solver

The invention discloses a device for dividing three regions after power failure of a power grid based on a high-order Isin solver. According to the specific implementation mode, a power grid topology is modeled into an undirected graph, and a partition target is converted into a minimum cut problem of minimizing the sum of line weights between different regions under the condition that engineering constraints are met. According to the method, constraint and targets are divided by constructing an Isin Hamiltonian quantity containing a high-order coupling term and a plurality of regions of native coding. The Isin Hamiltonian can be directly deployed on a high-order Isin solver realized by adopting a complementary metal oxide semiconductor architecture, and an optimal partition scheme is quickly searched by utilizing the characteristics of large-scale parallelism and low power consumption of the Isin solver. According to the method, the problem of rapid and accurate partition of parallel recovery after blackout is effectively solved, and the power grid toughness is remarkably improved.
Owner:GUANGXI UNIV

A simulation evaluation method and apparatus for a chip model

PendingCN122365960AMassively parallelSystem-level simulation
This invention discloses a simulation evaluation method and apparatus for chip models, comprising: generating feasible deployment strategies for the computational graph of an AI model through an AI system-level simulation platform, and sending the operator feature information of each candidate deployment strategy extracted from the candidate deployment strategy set to a cloud base; determining the target deployment strategy requiring accurate simulation based on the user's simulation requirements through the cloud base, and sending the operator-level simulation task to the SOC operator-level simulation platform; the SOC operator-level simulation platform feeding back the obtained accurate simulation results to the cloud base, and the cloud base performing accuracy back-annotation on the target deployment strategy based on the accurate simulation results. By introducing a cloud base and cloud-based elastic computing resources, standardized access and end-to-end data connectivity of heterogeneous simulation platforms are achieved. This solution not only improves the efficiency of operator-level simulation from single-machine serial to large-scale parallel in the cloud, but also significantly improves the accuracy of overall network performance prediction through a dynamic back-annotation mechanism.
Owner:SHANGHAI SUIYUAN TECH CO LTD

Automated setup and communication coordination for training and utilizing massively parallel neural networks

A method is disclosed for training and utilizing massively parallel neural networks. A distributed computing system may be configured to perform various operations. The distributed computing system may divide a directed acyclic graph (“DAG”) that comprises a plurality of vertices linked in pairwise relationships via a plurality of edges among a plurality of nodes. Each node may comprise a computing device. The distributed computing system may provide a map of the DAG that described a flow of data through the vertices to each of the vertices of the DAG. The distributed computing system may perform a topological sort of the vertices of the DAG and may traverse the DAG.
Owner:FORD GLOBAL TECH LLC