Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

33results about "Dataflow computers" patented technology

Calculation method and device based on data flow diagram

The embodiment of the invention provides a method for carrying out calculation based on a data flow graph (DFG) and related equipment. The method comprises the steps that a first operator obtains a plurality of data values from a second operator, the data flow diagram comprises the first operator and the second operator, and the data values are stored in a plurality of instruction units corresponding to the second operator in the DFG respectively; the first operator calculates an output result of the first operator according to the plurality of data values in the second operator, and the plurality of data values in the second operator are input of the first operator. By the proposed techniques, cycling or recursion may be performed in the DFG.
Owner:HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD

Systems and methods for implementing directional operand broadcast and multiply-accumulate execution using a configurable patch mesh in a multi-core processing array of an integrated circuit

A technique is disclosed for operand propagation and accumulation within a processing array of an integrated circuit using overlapping patch regions. The system includes an interconnecting processing patch defined over a rectilinear subset of processing elements, with an origin processing element broadcasting operand data to the remaining elements in a directionally constrained, time-staggered wavefront pattern. A logical processing patch is separately defined over a second rectilinear subset of processing elements. The interconnecting processing patch and the logical processing patch partially overlap to form an interconnecting patch mesh comprising a common set of processing elements. Operand data is propagated from the origin of the interconnecting patch to the common processing elements within the patch mesh, enabling operand handoff or accumulation across patch boundaries. The architecture supports fine-grained, localized data movement and patch-level execution coordination across a mesh of processing elements to optimize compute reuse, operand locality, and execution throughput.
Owner:QUADRIC IO INC

Configurable wavefront parallel processor

An apparatus comprising: at least one processing element configured to process a data flow in at least one direction of a plurality of directions; a configuration register comprising at least one setting that determines the processing of the data flow with the at least one processing element; and a shift register configured to select data of the at least one processing element from the at least one direction, and to provide at least one shifted data sample to a plurality of slices configured to perform at least one arithmetic operation with the data flow; wherein at least one slice of the plurality of slices is configured with the at least one setting of the configuration register.
Owner:NOKIA SOLUTIONS & NETWORKS OY

Executing a compute graph on multiple reconfigurable dataflow processors

A method for a reconfigurable computing system includes receiving a compute graph for execution on multiple RDPs interconnected with a ring network having R interconnected RDPs. A compute graph with a node specifying a reduction operation for a first and second tensor is detected. Executing the compute graph on the multiple RDPs.
Owner:SAMBANOVA SYSTEMS INC

Special signal generating device based on software defined radio and GPU server

The invention belongs to the technical field of information technology, and discloses a special signal generating device based on software defined radio and a GPU server, comprising: a graphic processing unit used for generating high-speed orthogonal data streams corresponding to various communication systems and modulation modes, the high-speed orthogonal data streams being transmitted to an intermediate frequency signal unit through an optical fiber; the intermediate frequency signal unit is used for receiving the high-speed orthogonal data stream, completing signal modulation and outputting an intermediate frequency signal; the software-defined radio platform is used for carrying out secondary development to realize signal generation and reception, and the secondary development builds a heterogeneous system integrated by a central processing unit and a graphic processing unit based on general radio software; the central processing unit is used for task scheduling and memory configuration; the invention aims to solve the problem that a signal generating device in the prior art cannot realize high-efficiency collaboration of a central processing unit and a graphic processing unit in high-speed data stream processing.
Owner:成都玖锦科技有限公司

Computer-readable recording medium, conversion method, and conversion apparatus

PendingJP2026027914ADataflow computersCAD circuit designPathPingAlgorithm
To provide a conversion program, a conversion method and a conversion device for improving processing capability.SOLUTION: Acquiring a mapping result including a predetermined DFG and information on assignment of operations to respective computing units and wiring between the computing units determined so as to correspond to the predetermined DFG for a CGRA having a plurality of computing units, and extracting a portion corresponding to a predetermined DFG pattern from the predetermined DFG; A computer is made to execute processing for determining a conversion candidate DFG from among DFGs corresponding to a pattern of an extraction place on the basis of a position to which the extraction place in a mapping result is assigned and the number of transmission paths of data to be used between arithmetic units, and generating a converted DFG by converting a predetermined DFG on the basis of the conversion candidate DFG.SELECTED DRAWING: Figure 6
Owner:FUJITSU LTD

Processor for configurable parallel computations

A flexible processor includes (i) numerous configurable processors interconnected by modular interconnection fabric circuits that are configurable to partition the configurable processors into one or more groups, for parallel execution, and to interconnect the configurable processors in any order for pipelined operations, Each configurable processor may include (i) a control circuit; (ii) numerous configurable arithmetic logic circuits; and (iii) configurable interconnection fabric circuits for interconnecting the configurable arithmetic logic circuits.
Owner:STAR ALLY INT LTD

Hardware accelerator with configurable tensor operation pipeline

PCT designated stageWO2026015262A1Dataflow computersPhysical realisationData classComputer architecture
A hardware accelerator (10) is disclosed that can flexibly be configured to support differing data types and differing operation flows. The hardware accelerator includes a plurality of fixed tensor operation logic units (16), tensor operation pipeline logic (18) configured to receive from the processor a pipeline command (24) including a software-defined tensor operation pipeline definition (26) defining a plurality of tensor operation stages (30) in a tensor operation pipeline (32) and associated predetermined tensor operations to be performed at each of the defined tensor operation stages. The hardware accelerator is further configured to receive tensor data (28) to be computed by the tensor operation pipeline, and implement the tensor operation pipeline to perform the tensor operations in each of the tensor operation stages on the tensor data, to thereby produce a tensor operation pipeline result (34) for the tensor data, and output the tensor operation pipeline result to the processor.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Semiconductor Apparatus, Semiconductor Device, Method for a Semiconductor Device, and Non-Transitory Computer-Readable Medium, Method, Apparatus and Device for a Computer System

PendingUS20260119174A1Dataflow computersProgram saving/restoringComputer architectureData stream
Various examples relate to a semiconductor apparatus, or to a non-transitory computer-readable medium, a method, an apparatus or a device for a computer system, and to a computer system comprising the semiconductor apparatus and the apparatus or device. A semiconductor apparatus comprises interface circuitry for obtaining a dataflow graph comprising a plurality of nodes, and a plurality of processing elements, an interconnect network coupled to the plurality of processing elements and configured to receive an input of the dataflow graph, wherein the dataflow graph is to configure the interconnect network and the plurality of processing elements, wherein the processing elements are to perform a plurality of operations defined by the nodes of dataflow graph, wherein the dataflow graph comprises a first type of node for performing a computation and a second type of node for determining a branching condition, wherein the semiconductor apparatus is configured to, upon determining a result of a branching condition specified by a node having the second type, configure the interconnect network and the processing elements based on the result of the branching condition.
Owner:INTEL CORP

Adaptive and reconfigurable dataflow computing system and method

PCT designated stageWO2026015086A1Dataflow computersResource allocationComputer architectureApplication procedure
The present invention relates to an adaptive and reconfigurable dataflow computing system and method. The dataflow computing includes a dataflow-driven core to dynamically reprogram itself based on computational needs during runtime. In this system, an application program is broken down into a dataflow graph with functional units by a compiler, and an instruction dispatcher allocates and runs available processing elements to execute all functional units in the dataflow graph to perform distributed or parallel processing. The dynamic allocation of processing units can be adapted during runtime.
Owner:NATIONAL UNIVERSITY OF SINGAPORE

Systems and methods for implementing directional operand broadcast and multiply-accumulate execution using a configurable patch mesh in a multi-core processing array of an integrated circuit

A technique is disclosed for operand propagation and accumulation within a processing array of an integrated circuit using overlapping patch regions. The system includes an interconnecting processing patch defined over a rectilinear subset of processing elements, with an origin processing element broadcasting operand data to the remaining elements in a directionally constrained, time-staggered wavefront pattern. A logical processing patch is separately defined over a second rectilinear subset of processing elements. The interconnecting processing patch and the logical processing patch partially overlap to form an interconnecting patch mesh comprising a common set of processing elements. Operand data is propagated from the origin of the interconnecting patch to the common processing elements within the patch mesh, enabling operand handoff or accumulation across patch boundaries. The architecture supports fine-grained, localized data movement and patch-level execution coordination across a mesh of processing elements to optimize compute reuse, operand locality, and execution throughput.
Owner:QUADRIC IO INC

Tensor data exchange circuit, data flow processing apparatus and method

ActiveCN121029690BDataflow computersTransmissionAlgorithmData stream processing
The application provides a tensor data exchange circuit, a data stream processing device and a method. The tensor data exchange circuit is used for a data stream processor and comprises a first non-blocking full permutation component and a second non-blocking full permutation component. The first non-blocking full permutation component and the second non-blocking full permutation component are used for performing data format conversion on a plurality of sub-data blocks split from input data, performing intra-block data exchange and inter-block data exchange, and outputting data meeting target component format requirements. The first non-blocking full permutation component performs intra-block data exchange on the sub-data blocks. The second non-blocking full permutation component performs inter-block data exchange on the plurality of sub-data blocks.
Owner:BEIJING TSINGMICRO INTELLIGENT TECH CO LTD

Hardware accelerator with configurable tensor operation pipeline

A hardware accelerator is disclosed that can flexibly be configured to support differing data types and differing operation flows. The hardware accelerator includes a plurality of fixed tensor operation logic units, tensor operation pipeline logic configured to receive from the processor a pipeline command including a software-defined tensor operation pipeline definition defining a plurality of tensor operation stages in a tensor operation pipeline and associated predetermined tensor operations to be performed at each of the defined tensor operation stages. The hardware accelerator is further configured to receive tensor data to be computed by the tensor operation pipeline, and implement the tensor operation pipeline to perform the tensor operations in each of the tensor operation stages on the tensor data, to thereby produce a tensor operation pipeline result for the tensor data, and output the tensor operation pipeline result to the processor.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Systems and methods for implementing directional operand broadcast and multiply-accumulate execution using a configurable patch mesh in a multi-core processing array of an integrated circuit

A technique is disclosed for operand propagation and accumulation within a processing array of an integrated circuit using overlapping patch regions. The system includes an interconnecting processing patch defined over a rectilinear subset of processing elements, with an origin processing element broadcasting operand data to the remaining elements in a directionally constrained, time-staggered wavefront pattern. A logical processing patch is separately defined over a second rectilinear subset of processing elements. The interconnecting processing patch and the logical processing patch partially overlap to form an interconnecting patch mesh comprising a common set of processing elements. Operand data is propagated from the origin of the interconnecting patch to the common processing elements within the patch mesh, enabling operand handoff or accumulation across patch boundaries. The architecture supports fine-grained, localized data movement and patch-level execution coordination across a mesh of processing elements to optimize compute reuse, operand locality, and execution throughput.
Owner:QUADRIC IO INC

Processor for configurable parallel computations

A flexible processor includes (i) numerous configurable processors interconnected by modular interconnection fabric circuits that are configurable to partition the configurable processors into one or more groups, for parallel execution, and to interconnect the configurable processors in any order for pipelined operations, Each configurable processor may include (i) a control circuit; (ii) numerous configurable arithmetic logic circuits; and (iii) configurable interconnection fabric circuits for interconnecting the configurable arithmetic logic circuits.
Owner:STAR ALLY INT LTD

Non-transitory computer-readable recording medium, mapping result verification method, and mapping result verification apparatus

PendingUS20260030200A1Dataflow computersMachine execution arrangementsData streamAlgorithm
A non-transitory computer-readable recording medium stores therein a mapping result verification program that causes a computer to execute a process including acquiring both of a data flow graph that represents predetermined calculation including a plurality of arithmetic operations and a mapping result obtained by mapping the data flow graph into a CGRA that includes a plurality of arithmetic operation units, and verifying, regarding each of first arithmetic operation units to which the arithmetic operations have been respectively allocated in the mapping result, based on an arithmetic operation result that has been obtained by performing the arithmetic operations, whether or not the mapping result matches the data flow graph.
Owner:FUJITSU LTD

Configurable processing architecture

ActiveUS12517862B2Dataflow computersElectric digital data processing
A configurable processing unit including a core processing element and a plurality of assist processing elements can be coupled together by one or more networks. The core processing element can include a large processing logic, large non-volatile memory, input / output interfaces and multiple memory channels. The plurality of assist processing elements can each include smaller processing logic, smaller non-volatile memory and multiple memory channels. One or more bitstreams can be utilized to configure and reconfigure computation resources of the core processing element and memory management of the plurality of assist processing elements.
Owner:ALIBABA GROUP HOLDING LTD

Scheduling tasks for execution by a processor system

A computer-implemented method schedules a plurality of tasks for execution by a processor system. A first execution model for the plurality of tasks is accessed. Data is generated that identifies which tasks in the execution model are not direct-feedthrough tasks. The data is used to determine an order for executing the tasks at least partly in dependence on whether or not each task is a direct-feedthrough task.
Owner:COLLINS AEROSPACE IRELAND LTD

Matrix multiplier for transformer-based model training

This invention provides a matrix multiplier for training Transformer-type models, comprising an M-row, N-column systolic array. The systolic array is two-dimensional and consists of R-row, C-column interconnected processing units (PEs). Each PE includes one multiplier, one adder, two internal registers, one left-side multiplexer, and two right-side multiplexers. The left-side multiplexer can select whether the input to the multiplier comes from outside the PE or retains the input from the previous cycle. When retaining the input from the previous cycle, the PE maintains the WS data stream with weights. This invention designs a reconfigurable processing unit (PE) that can flexibly support multiple data streams at different stages and cycles of training and select the data source according to requirements.
Owner:NANJING UNIV

Variable recalculation method, system and equipment based on data flow analysis and medium

PendingCN121560385ADataflow computersConcurrent instruction executionData controlAlgorithm
The invention relates to a variable recalculation method, a variable recalculation system, variable recalculation equipment and a medium based on data flow analysis. The method comprises the steps of setting Boolean attributes on a ternary lattice of a re-computable seed, and spreading the Boolean attributes according to a preset instruction type, so that the Boolean attributes are unified at a data control flow merging point through union operation of the lattice, and a re-computable judgment result of each variable is obtained. And identifying cross-sub-kernel variables, and classifying the cross-sub-kernel variables according to the re-computable judgment result to obtain uniform variables, re-computable variables and non-re-computable variables. A first data structure set is constructed by assigning on-stack variables to uniform variables, recording a calculation sequence of recalculable variables, and assigning on-stack arrays to non-recalculable variables. According to the first data structure set, code conversion is carried out when the work item loop is created, and variable recalculation and expansion of the variables are completed. By adopting the method, redundant storage and memory access operations can be reduced to the maximum extent.
Owner:NAT UNIV OF DEFENSE TECH

Mapping result verification method and mapping result verification apparatus

PendingJP2026019020ADataflow computersCAD circuit designData streamAlgorithm
To provide a mapping result verification program, a mapping result verification method and a mapping result verification device for improving the accuracy of mapping to a CGRA.SOLUTION: Acquiring a data flow graph representing a predetermined calculation including a plurality of operations and a mapping result obtained by mapping the data flow graph to a CGRA having a plurality of arithmetic elements, and verifying whether the mapping result and the data flow graph match based on an operation result obtained by executing an operation for each of first arithmetic elements to which the operation is assigned in the mapping result.SELECTED DRAWING: Figure 4
Owner:FUJITSU LTD

System and method for efficient processing with a processing element local memory

A system and corresponding method provide efficient processing with a processing element local memory. The method transfers input data from a host processing unit to a processing element local memory. The host processing unit is coupled to a host local memory. The processing element local memory is coupled to a processing element. The method performs, by the processing element, at least one operation on the input data in the processing element local memory. The method stores, in the processing element local memory, at least one intermediate result of the at least one operation performed. The method performs, by the processing element, at least one subsequent operation on the at least one intermediate result stored in the processing element local memory. The method transfers at least one result of the at least one subsequent operation to the host processing unit, enabling efficient retrieval and analysis of data in data storage systems.
Owner:DATAPELAGO INC

Energy-minimal dataflow architecture with programmable on-chip network

Disclosed herein is a co-designed compiler and CGRA architecture that achieves both high programmability and extreme energy efficiency. The architecture includes a rich set of control-flow operators that support arbitrary control flow and memory access on the CGRA fabric. The architecture is able to realize both energy and area savings over prior art implementations by offloading most control operations into a programmable on-chip network where they can re-use existing network switches.
Owner:CARNEGIE MELLON UNIV

Data processing devices and methods

A data processor is suggested comprising at least an instruction issue stage issuing instructions, a number of processing elements to which at least some of the instructions are issued and which receive operand data, generate result data in accordance with the instructions received and transmit their result data to other processing elements for use as new operands; and a bus system for these transmissions, wherein the instruction issue stage is adapted to issue instructions to a group of processing elements to operate them in at least two different modes, namely an out-of-order mode wherein instructions may be executed out of order and a loop acceleration mode wherein loops can be executed efficiently, and wherein the bus system comprises an arbiter operative in the out-of-order mode to arbitrate access of the group of processing elements to at least a part of the bus system and inoperative in the loop acceleration mode.
Owner:UBITIUM GMBH

Intelligent graph execution for a reconfigurable data processor

A data processing system including an array of reconfigurable units and a compiler configured to generate to execute a dataflow graph of a user application is disclosed. The dataflow graph includes a sequence of temporal partitions, each temporal partition including a sequence of graph control operations. Also disclosed is an intelligent graph orchestration and execution engine (IGOEE) configured to receive an optimization objective from the complier. The IGOEE can reorganize the sequence of temporal partitions and the sequence of graph control operations within each temporal partition to satisfy the optimization objective, and execute the reorganized dataflow graph on the reconfigurable processor.
Owner:SAMBANOVA SYSTEMS INC

Near-memory operator and method with accelerator performance improvement

A system configured to perform an operation includes: a hardware device comprising a plurality of computing modules and a plurality of memory modules arranged in a lattice form, each of the computing modules comprising a coarse-grained reconfigurable array and each of the memory modules comprising a static random-access memory and a plurality of functional units connected to the static random-access memory; and a compiler configured to divide a target operation and assign the divided target operation to the computing modules and the memory modules such that the computing modules and the memory modules of the hardware device perform the target operation.
Owner:SAMSUNG ELECTRONICS CO LTD

Array processing entity array placement and routing

PendingUS20260133933A1Dataflow computersElectric digital data processingAlgorithmProcessing element
A method for automatically placing and routing processing elements (PEs), where each PE has a limited connectivity to its neighbors. The method may obtain placement primitive (PP) location window information, the location window information defining possible locations of a group of PPs within a coarse-grained reconfigurable array (CGRA). The method may obtain hardware constraints regarding the array of PEs, the hardware constraints comprise local connectivity constraints and remote connectivity constraints. The method may receive a computation graph (CG) that represents mathematical expressions to be calculated by the array of PEs. The method may determine a location of the PEs of the array of PEs based on the CG, the PP location window information, and hardware constraints.
Owner:MOBILEYE VISION TECH LTD