A hierarchical control stack cuts latency and data-channel load to enable stable closed-loop control of large independent element arrays.
A layered controller cuts latency and data flow for large element arrays while preserving closed-loop precision and independent parallel control.
A hierarchical parallel FSM structure passes only essential state information between levels to enable faster real-time pattern processing.
Compressed state outputs and feedback between cascaded finite state machines cut transfer volume and support real-time pattern recognition.
Selective state output, aggregation, and difference vectors cut inter-engine transfer time in hierarchical parallel machines for real-time operation.
Cascaded finite state machine levels cut state vector transfer through extraction, aggregation, and compression to support real-time pattern recognition.
Using parallel register-based circuit units, this case shows how data statistics can cut hardware usage and power consumption.
Aggregated state vectors between cascaded finite state machine engines reduce processing delays while preserving real-time pattern recognition.
Selective register updates and local counter increments cut memory access, hardware resources, and power in data statistics circuits.
Cascaded finite state machine levels compress shared state information to cut transfer load and support real-time pattern recognition.
Compressed final-state aggregation cuts transfer delay in hierarchical finite state machines while preserving needed state information.
Moving memory mapping out of critical CGRA stages cuts latency and improves dataflow throughput for AI and machine learning workloads.
Sharded tensors and staged reduction nodes cut inter-chip latency when compute graphs run across ring-connected reconfigurable dataflow processors.
An IGOEE reorganizes temporal partitions and graph control operations to cut execution overhead and improve reconfigurable processor utilization.
Pattern-based DFG conversion cuts data transmission columns in CGRA mapping, improving PE utilization and processing capacity.
A co-designed compiler and CGRA offload control flow to a programmable on-chip network, cutting energy overhead while handling irregular code.
Automatic verification checks whether CGRA mapping results match the original data flow graph, improving mapper reliability without manual review.
Modular interconnection fabric lets configurable processors be grouped for parallel execution or reordered into pipelines for repetitive computing tasks.
Software-defined tensor stages let fixed logic units process custom data formats directly, avoiding conversion overhead while sustaining throughput.
Dynamic bitstream reconfiguration balances compute intensity, memory capacity, and bandwidth across core and assist processors.
Partitioning tensor operations into parallel graph nodes cuts latency and boosts throughput in reconfigurable dataflow arrays.
Partitioning sequential AI pipeline stages across GPUs with RDMA cuts transfer overhead and enables real-time high-throughput processing.
Partitioned all-reduce across ring-connected RDPs cuts inter-chip latency while scaling tensor reduction for large neural network models.
Dynamic operand-ready scheduling and broadcast data paths cut routing latency and mapping complexity in reconfigurable PE arrays.
An IGOEE reorganizes temporal partitions and graph control operations to reduce setup overhead and maximize resource utilization during execution.
Cost estimation of logical-edge bandwidth before placement and routing helps compilers allocate reconfigurable-processor resources more effectively.
Traditional compilers waste processor resources when routing dataflow graphs; bandwidth cost estimates guide better placement and routing.
Traditional compilers lack accurate placement costs; this tool predicts logical-edge bandwidth and link congestion to guide routing on reconfigurable processors.
Compiler analysis aligns same-packet operands across variable-latency CGRA stages, supporting correct pipelined results with efficient storage.
Cost estimation guides compiler placement and routing by measuring logical-edge bandwidth and physical-link congestion before final assignments.
Multi-level graph APIs let AI processing units skip unnecessary image-processing subgraphs when runtime conditions are not met.
Moving memory mapping from critical stages to adjacent buffer stages reduces latency and increases throughput in tensor-based CGRA execution.
Compile-time instrumentation and hardware counters help isolate stage-latency bottlenecks in asynchronous dataflow workloads on reconfigurable processors.
This case splits tensor operations across graph nodes for parallel execution, then gathers results to improve throughput.
This case combines shared data flows and switches routing to account for computing-unit capabilities and resource use.
This case maps operations across heterogeneous pipeline stages and uses compiler-planned register delays to align outputs efficiently.
A data processing apparatus generates new inputs with fewer combinations to associate and reuse arithmetic outputs.
A convolution engine accumulates partial results on a processing element array using unidirectional dataflows to reduce memory access.
A parallel processor system combines control and data flow units to minimize memory latency.
A dynamic register machine modifies its own program instructions during execution to perform complex computations.
Segments disconnected graphs into independent components to reduce processing time and memory usage during flow model computation.
System executes data flow graphs using multiple actors to independently calculate results and select the best output based on quality descriptors.
Splitting buffers based on cost analysis reduces memory unit consumption and optimizes resource utilization in coarse-grained reconfigurable processors.
Identifying direct-feedthrough tasks allows a scheduler to resolve contention delays and ensure deterministic execution in multi-core systems.
A dynamic processor architecture uses an arbiter to skip unnecessary processing stages, reducing time delays while maintaining complete data reliability.
A network overlay hardware unit isolates computation from memory management tasks.
Segmenting layer normalization into pipelined stages with checkpoints resolves processing speed bottlenecks in large AI models.