Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

82results about "Dataflow computers" patented technology

Graph spatial split

A method for reducing latency and increasing throughput in a reconfigurable computing system includes receiving a compute graph for execution on a reconfigurable dataflow processor comprising a grid of compute units and grid of memory units interconnected with a switching array. The compute graph includes a node specifying an operation on a tensor. The node may be split into multiple nodes that each specify the operation on a distinctive portion of the tensor to produce a first modified compute graph. The first modified compute graph may be executed. In addition, the multiple nodes may be within a single meta-pipeline stage and may be processed in parallel. Furthermore, the compute graph may further comprise a separate node for gathering the distinctive portions of the tensor into a complete tensor, to produce a second modified compute graph.
Owner:SAMBANOVA SYSTEMS INC

Energy-minimal dataflow architecture with programmable on-chip network

Disclosed herein is a co-designed compiler and CGRA architecture that achieves both high programmability and extreme energy efficiency. The architecture includes a rich set of control-flow operators that support arbitrary control flow and memory access on the CGRA fabric. The architecture is able to realize both energy and area savings over prior art implementations by offloading most control operations into a programmable on-chip network where they can re-use existing network switches.
Owner:CARNEGIE MELLON UNIV

Calculation method and device based on data flow diagram

The embodiment of the invention provides a method for carrying out calculation based on a data flow graph (DFG) and related equipment. The method comprises the steps that a first operator obtains a plurality of data values from a second operator, the data flow diagram comprises the first operator and the second operator, and the data values are stored in a plurality of instruction units corresponding to the second operator in the DFG respectively; the first operator calculates an output result of the first operator according to the plurality of data values in the second operator, and the plurality of data values in the second operator are input of the first operator. By the proposed techniques, cycling or recursion may be performed in the DFG.
Owner:HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD

Systems and methods for implementing directional operand broadcast and multiply-accumulate execution using a configurable patch mesh in a multi-core processing array of an integrated circuit

A technique is disclosed for operand propagation and accumulation within a processing array of an integrated circuit using overlapping patch regions. The system includes an interconnecting processing patch defined over a rectilinear subset of processing elements, with an origin processing element broadcasting operand data to the remaining elements in a directionally constrained, time-staggered wavefront pattern. A logical processing patch is separately defined over a second rectilinear subset of processing elements. The interconnecting processing patch and the logical processing patch partially overlap to form an interconnecting patch mesh comprising a common set of processing elements. Operand data is propagated from the origin of the interconnecting patch to the common processing elements within the patch mesh, enabling operand handoff or accumulation across patch boundaries. The architecture supports fine-grained, localized data movement and patch-level execution coordination across a mesh of processing elements to optimize compute reuse, operand locality, and execution throughput.
Owner:QUADRIC IO INC

Configurable wavefront parallel processor

An apparatus comprising: at least one processing element configured to process a data flow in at least one direction of a plurality of directions; a configuration register comprising at least one setting that determines the processing of the data flow with the at least one processing element; and a shift register configured to select data of the at least one processing element from the at least one direction, and to provide at least one shifted data sample to a plurality of slices configured to perform at least one arithmetic operation with the data flow; wherein at least one slice of the plurality of slices is configured with the at least one setting of the configuration register.
Owner:NOKIA SOLUTIONS & NETWORKS OY

Executing a compute graph on multiple reconfigurable dataflow processors

A method for a reconfigurable computing system includes receiving a compute graph for execution on multiple RDPs interconnected with a ring network having R interconnected RDPs. A compute graph with a node specifying a reduction operation for a first and second tensor is detected. Executing the compute graph on the multiple RDPs.
Owner:SAMBANOVA SYSTEMS INC

Special signal generating device based on software defined radio and GPU server

The invention belongs to the technical field of information technology, and discloses a special signal generating device based on software defined radio and a GPU server, comprising: a graphic processing unit used for generating high-speed orthogonal data streams corresponding to various communication systems and modulation modes, the high-speed orthogonal data streams being transmitted to an intermediate frequency signal unit through an optical fiber; the intermediate frequency signal unit is used for receiving the high-speed orthogonal data stream, completing signal modulation and outputting an intermediate frequency signal; the software-defined radio platform is used for carrying out secondary development to realize signal generation and reception, and the secondary development builds a heterogeneous system integrated by a central processing unit and a graphic processing unit based on general radio software; the central processing unit is used for task scheduling and memory configuration; the invention aims to solve the problem that a signal generating device in the prior art cannot realize high-efficiency collaboration of a central processing unit and a graphic processing unit in high-speed data stream processing.
Owner:成都玖锦科技有限公司

Merging Buffer Access Operations of a Compute Graph

A method for merging buffers and associated operations includes receiving a compute graph for a reconfigurable dataflow computing system and conducting a buffer allocation and merging process responsive to determining that a first operation specified by a first operation node is a memory indexing operation and that the first operation node is a producer for exactly one consuming node that specifies a second operation. The buffer allocation and merging process may include replacing the first operation node and the consuming node with a merged buffer node within the graph responsive to determining that the first operation and the second operation can be merged into a merged indexing operation and that the resource cost of the merged node is less than the sum of the resource costs of separate buffer nodes. A corresponding system and computer readable medium are also disclosed herein.
Owner:SAMBANOVA SYSTEMS INC

Computer-readable recording medium, conversion method, and conversion apparatus

To provide a conversion program, a conversion method and a conversion device for improving processing capability.SOLUTION: Acquiring a mapping result including a predetermined DFG and information on assignment of operations to respective computing units and wiring between the computing units determined so as to correspond to the predetermined DFG for a CGRA having a plurality of computing units, and extracting a portion corresponding to a predetermined DFG pattern from the predetermined DFG; A computer is made to execute processing for determining a conversion candidate DFG from among DFGs corresponding to a pattern of an extraction place on the basis of a position to which the extraction place in a mapping result is assigned and the number of transmission paths of data to be used between arithmetic units, and generating a converted DFG by converting a predetermined DFG on the basis of the conversion candidate DFG.SELECTED DRAWING: Figure 6
Owner:FUJITSU LTD

Processor for configurable parallel computations

A flexible processor includes (i) numerous configurable processors interconnected by modular interconnection fabric circuits that are configurable to partition the configurable processors into one or more groups, for parallel execution, and to interconnect the configurable processors in any order for pipelined operations, Each configurable processor may include (i) a control circuit; (ii) numerous configurable arithmetic logic circuits; and (iii) configurable interconnection fabric circuits for interconnecting the configurable arithmetic logic circuits.
Owner:STAR ALLY INT LTD

Schedule-aware dynamically reconfigurable adder tree architecture for partial sum accumulation in machine learning accelerator

Embodiments of the present disclosure are directed toward techniques and configurations enhancing the performance of hardware (HW) accelerators. The present disclosure provides a schedule-aware, dynamically reconfigurable, tree-based partial sum accumulator architecture for HW accelerators, wherein the depth of an adder tree in the HW accelerator is dynamically based on a dataflow schedule generated by a compiler. The adder tree depth is adjusted on a per-layer basis at runtime. Configuration registers, programmed via software, dynamically alter the adder tree depth for partial sum accumulation based on the dataflow schedule. By facilitating a variable depth adder tree during runtime, the compiler can choose a compute optimal dataflow schedule that minimizes the number of compute cycles needed to accumulate partial sums across multiple processing elements (PEs) within a PE array of a HW accelerator. Other embodiments may be described and / or claimed.
Owner:INTEL CORP

All reduce across multiple reconfigurable dataflow processors

A method for a reconfigurable computing system includes receiving a compute graph for execution on multiple RDPs interconnected with a ring network having R interconnected RDPs. A compute graph with a node specifying a reduction operation for a first and second tensor is detected. The detected compute graph node is partitioned into a compute subgraph corresponding to an RDP of the R interconnected RDPs. A first node is inserted into the compute subgraph that specifies a partial reduction operation for producing a partial reduction result corresponding to a shard of the first tensor and a shard of the second tensor. A second node is inserted for communicating the partial reduction result to an adjacent RDP. A third node is inserted that specifies a reduction operation for producing a total reduction result. A fourth node is inserted for communicating the total reduction result to at least one other RDP.
Owner:SAMBANOVA SYSTEMS INC

Hardware accelerator with configurable tensor operation pipeline

A hardware accelerator (10) is disclosed that can flexibly be configured to support differing data types and differing operation flows. The hardware accelerator includes a plurality of fixed tensor operation logic units (16), tensor operation pipeline logic (18) configured to receive from the processor a pipeline command (24) including a software-defined tensor operation pipeline definition (26) defining a plurality of tensor operation stages (30) in a tensor operation pipeline (32) and associated predetermined tensor operations to be performed at each of the defined tensor operation stages. The hardware accelerator is further configured to receive tensor data (28) to be computed by the tensor operation pipeline, and implement the tensor operation pipeline to perform the tensor operations in each of the tensor operation stages on the tensor data, to thereby produce a tensor operation pipeline result (34) for the tensor data, and output the tensor operation pipeline result to the processor.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Processing unit and method for configuring the same

A configurable processing unit comprising a core processing element and a plurality of auxiliary processing elements may be coupled together via one or more networks. The core processing element may include large processing logic, large non-volatile memory, input / output interfaces, and multiple memory channels. Each of the plurality of auxiliary processing elements may include smaller processing logic, smaller non-volatile memory, and multiple memory channels. The computing resources of the core processing element and the memory management of the plurality of auxiliary processing elements may be configured and reconfigured using one or more bitstreams.
Owner:ALIBABA GROUP HOLDING LTD

Semiconductor Apparatus, Semiconductor Device, Method for a Semiconductor Device, and Non-Transitory Computer-Readable Medium, Method, Apparatus and Device for a Computer System

Various examples relate to a semiconductor apparatus, or to a non-transitory computer-readable medium, a method, an apparatus or a device for a computer system, and to a computer system comprising the semiconductor apparatus and the apparatus or device. A semiconductor apparatus comprises interface circuitry for obtaining a dataflow graph comprising a plurality of nodes, and a plurality of processing elements, an interconnect network coupled to the plurality of processing elements and configured to receive an input of the dataflow graph, wherein the dataflow graph is to configure the interconnect network and the plurality of processing elements, wherein the processing elements are to perform a plurality of operations defined by the nodes of dataflow graph, wherein the dataflow graph comprises a first type of node for performing a computation and a second type of node for determining a branching condition, wherein the semiconductor apparatus is configured to, upon determining a result of a branching condition specified by a node having the second type, configure the interconnect network and the processing elements based on the result of the branching condition.
Owner:INTEL CORP

Adaptive and reconfigurable dataflow computing system and method

The present invention relates to an adaptive and reconfigurable dataflow computing system and method. The dataflow computing includes a dataflow-driven core to dynamically reprogram itself based on computational needs during runtime. In this system, an application program is broken down into a dataflow graph with functional units by a compiler, and an instruction dispatcher allocates and runs available processing elements to execute all functional units in the dataflow graph to perform distributed or parallel processing. The dynamic allocation of processing units can be adapted during runtime.
Owner:NATIONAL UNIVERSITY OF SINGAPORE

Tensor data exchange circuit, data stream processing device and method

The invention provides a tensor data exchange circuit and a data stream processing device and method.The tensor data exchange circuit is used for a data stream processor and comprises a first non-blocking full-arrangement assembly and a second non-blocking full-arrangement assembly; performing data format conversion of in-block data exchange and inter-block data exchange on a plurality of sub-data blocks split by input data through the first non-blocking full-arrangement component and the second non-blocking full-arrangement component so as to output data meeting the format requirement of a target component; wherein the first non-blocking full-arrangement component is used for carrying out in-block data exchange on the sub-data blocks; and the second non-blocking full-arrangement component performs inter-block data exchange on the plurality of sub-data blocks.
Owner:BEIJING TSINGMICRO INTELLIGENT TECH CO LTD

Flow control for reconfigurable processors using control counters

The technology disclosed relates to storing a dataflow graph with a plurality of compute nodes that transmit data between the compute nodes, and controlling data transmission between compute nodes in the plurality of compute nodes based on ready-to-read credit counters and write credit counters. For example, systems and methods according to this disclosure may control data transmission between compute nodes along the data connections between the compute nodes by selectively controlling writing of data based on both the ready-to-read credit counter and the write credit counter of a particular compute node of the plurality of compute nodes.
Owner:SAMBANOVA SYSTEMS INC

Data Processing Method, Apparatus, Electronic Device, and Storage Medium

A data processing method, apparatus, electronic device, and storage medium. The data processing method includes: obtaining a computational graph; determining at least one core operator based on the computational graph; determining the arrangement modes respectively corresponding to a plurality of operators through at least one round of iterative operations; wherein, in each round of iterative operation: selecting one or more core operators from the at least one core operator as starting operators; determining the arrangement mode corresponding to each starting operator according to the read-write overheads of each starting operator under multiple preset arrangement modes; starting from the starting operators, combining the arrangement mode corresponding to the starting operators, and propagating forward and backward according to the dependency relationship to sequentially obtain the arrangement modes corresponding to each operator, wherein, in the process of sequentially obtaining the arrangement modes corresponding to each operator, obtaining the arrangement mode corresponding to the current operator by combining the arrangement modes corresponding to adjacent operators, and the adjacent operators are operators that are adjacent to the current operator according to the dependency relationship and whose arrangement modes have been determined.
Owner:SHANGHAI BIREN TECH CO LTD

Data processing devices and methods

A data processor is suggested comprising at least an instruction issue stage issuing instructions, a number of processing elements to which at least some of the instructions are issued and which receive operand data, generate result data in accordance with the instructions received and transmit their result data to other processing elements for use as new operands; and a bus system for these transmissions, wherein the instruction issue stage is adapted to issue instructions to a group of processing elements to operate them in at least two different modes, namely an out-of-order mode wherein instructions may be executed out of order and a loop acceleration mode wherein loops can be executed efficiently, and wherein the bus system comprises an arbiter operative in the out-of-order mode to arbitrate access of the group of processing elements to at least a part of the bus system and inoperative in the loop acceleration mode.
Owner:UBITIUM GMBH

Configurable processor for parallel computing

The flexible processor includes (i) a number of configurable processors interconnected by a modular interconnection fabric circuit that is configurable to divide the configurable processors into one or more groups for parallel execution and to interconnect the configurable processors in any order for pipeline processing, each configurable processor having (i) control circuitry, (ii) a number of configurable arithmetic logic circuits, and (iii) configurable interconnection fabric circuitry for interconnecting the configurable arithmetic logic circuits.
Owner:XINGMENG INT CO LTD

Systems and methods for implementing directional operand broadcast and multiply-accumulate execution using a configurable patch mesh in a multi-core processing array of an integrated circuit

A technique is disclosed for operand propagation and accumulation within a processing array of an integrated circuit using overlapping patch regions. The system includes an interconnecting processing patch defined over a rectilinear subset of processing elements, with an origin processing element broadcasting operand data to the remaining elements in a directionally constrained, time-staggered wavefront pattern. A logical processing patch is separately defined over a second rectilinear subset of processing elements. The interconnecting processing patch and the logical processing patch partially overlap to form an interconnecting patch mesh comprising a common set of processing elements. Operand data is propagated from the origin of the interconnecting patch to the common processing elements within the patch mesh, enabling operand handoff or accumulation across patch boundaries. The architecture supports fine-grained, localized data movement and patch-level execution coordination across a mesh of processing elements to optimize compute reuse, operand locality, and execution throughput.
Owner:QUADRIC IO INC

Polymorphic computing fabric for static dataflow execution of computation operations represented as dataflow graphs (DFGS)

Provided is polymorphic computing fabric (100) for static dataflow execution of computing kernels represented as dataflow graphs (DFGs), wherein the DFGs are realized directly in hardware. The computing fabric (100) includes a plurality of components arranged in a matrix configuration. The plurality of components includes processing elements (PEs 102-1 to 102-n) and switching elements (SEs 102-1 to 102-n). Computational Operations represented as DFGs are realized on the polymorphic computing fabric (100) by mapping the nodes onto the PEs (102-1 to 102-n) and routing the edges via the circuit switched network formed by SEs (102-1 to 102-n) to form a virtual circuit. The polymorphic computing fabric (100) is configured to run in a configure-and-execute scheme allowing for pre-configuration of the plurality of components prior to execution.
Owner:MORPHING MACHINES PVT LTD

Tensor data exchange circuit, data flow processing apparatus and method

The application provides a tensor data exchange circuit, a data stream processing device and a method. The tensor data exchange circuit is used for a data stream processor and comprises a first non-blocking full permutation component and a second non-blocking full permutation component. The first non-blocking full permutation component and the second non-blocking full permutation component are used for performing data format conversion on a plurality of sub-data blocks split from input data, performing intra-block data exchange and inter-block data exchange, and outputting data meeting target component format requirements. The first non-blocking full permutation component performs intra-block data exchange on the sub-data blocks. The second non-blocking full permutation component performs inter-block data exchange on the plurality of sub-data blocks.
Owner:BEIJING TSINGMICRO INTELLIGENT TECH CO LTD

Intelligent graph execution and orchestration engine for a reconfigurable data processor

A data processing system including an array of reconfigurable units and a compiler configured to generate to execute a dataflow graph of a user application is disclosed. The dataflow graph includes a sequence of temporal partitions, each temporal partition including a sequence of graph control operations. Also disclosed is an intelligent graph orchestration and execution engine (IGOEE) configured to receive an optimization objective from the complier. The optimization objective can be for minimizing execution time of the reconfigurable processor or maximizing computing resource utilization of the reconfigurable processor. The IGOEE can reorganize the sequence of temporal partitions and the sequence of graph control operations within each temporal partition to satisfy the optimization objective; and execute the reorganized dataflow graph on the reconfigurable processor. A corresponding method is also disclosed herein.
Owner:SAMBANOVA SYSTEMS INC

Data processing method of processor, electronic device, and storage medium

The present application discloses a data processing method of a processor, an electronic device, and a storage medium. The method is applied to a data processing device, and may comprise: acquiring a first data set and a second data set to be processed by a processor; respectively performing block decomposition on the first data set and the second data set to obtain a plurality of first initial sub-data sets and a plurality of second initial sub-data sets; using permutation rule information to respectively partition the first initial sub-data sets in a first permutation direction, and partition the second initial sub-data sets in a second permutation direction to obtain a plurality of first target sub-data sets and a plurality of second target sub-data sets; in corresponding processing elements, performing computational processing on the first target sub-data sets and the second target sub-data sets to obtain output results of the corresponding processing elements; and determining an output result of the processor on the basis of the output results of the processing elements. The present application solves the technical problem of low data processing efficiency of the processor.
Owner:CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD

Configurable processor for parallel computing

The flexible processor includes (i) a number of configurable processors interconnected by a modular interconnection fabric circuit that is configurable to divide the configurable processors into one or more groups for parallel execution and to interconnect the configurable processors in any order for pipeline processing, each configurable processor having (i) control circuitry, (ii) a number of configurable arithmetic logic circuits, and (iii) configurable interconnection fabric circuitry for interconnecting the configurable arithmetic logic circuits.
Owner:XINGMENG INT CO LTD

Disaggregation of processing pipeline

A method for processing includes receiving a definition of a processing pipeline including multiple sequential processing stages. The processing pipeline is partitioned into a plurality of partitions. The first partition of the processing pipeline is executed on a first computational accelerator, whereby the first computational accelerator writes output data from a final stage of the first partition to an output buffer in a first memory. The output data are copied over a packet communication network to an input buffer in a second memory. The second partition of the processing pipeline is executed on a second computational accelerator using the copied output data in the second memory as input data to a first stage of the second partition.
Owner:NVIDIA CORP