An indirect branch instruction enables function calls in single-instruction multiple-thread processors by accepting address registers as arguments.
A unified shader core balances vertex and pixel loads by executing diverse programs on shared clusters, reducing idle cycles in graphics processors.
Chaining engines via a super-descriptor reduces firmware overhead and enhances throughput by enabling autonomous data flow processing.
Parallel memory banks with a shifted storage scheme resolve silicon area and power bottlenecks while boosting processing throughput for SIMD processors.
Horizontal aggregation SIMD instructions accelerate sorting speed by reducing comparisons in K-heaps, resolving sequential processing bottlenecks.
A parallel processing device implements contention-free routing schedules to optimize data movement across tile arrays.
A multi-core image processor uses a configurable internal network to dynamically couple specific cores for targeted execution.
A random matrix network simplifies backpropagation using fixed weight matrices for faster learning.
Vector merge circuitry determines required lanes and allocates execution threads to optimize resource utilization while managing device complexity.
A compiler inserts acquire and release instructions to dynamically allocate registers from a shared pool based on liveness analysis.
A scalable compute fabric dynamically configures CPU and GPU cores into processing pipelines to match specific workflow instruction sets.
A stream data processor partitions processing into dedicated hardware units and flexible control logic to optimize die size and power consumption.
Dual ALU pipelines process wavefronts via cache-stored operands, resolving idle pipeline bottlenecks and boosting execution throughput.
A min/max collector hardware component tracks extreme values during memory writes to accelerate vector processing unit throughput.
Peripheral sensor implements HID SPB interface to communicate via simple peripheral bus, resolving integration complexity from proprietary drivers.
A scalable array processor uses interconnected micro-engines to execute waveform processing algorithms in software defined radio systems.
Trusted hardware chains and spanning tree protocols ensure memory access integrity by validating address decoding mechanisms.
A decode modifier instruction maps scalar operands to vector registers, reducing encoding space pressure and instruction set complexity.
Daisy chain processing units with demultiplexers reduce routing congestion by assigning dedicated time frames for data packets, lowering die size.
A programmable logic device control program selects logic circuits for replacement based on their state data save and restore times.
Direct interconnections between processing elements reduce data movement power consumption in neural network computations.
Selection logic supplies identical data across multiple pipeline inputs, reducing chip area and propagation delays during gradient calculations.
Modified bitonic merge sorts data into ascending and descending registers, eliminating reverse operations that increase computational complexity.
Segmenting memory address updates from ALU computing operations eliminates waiting time for memory access, reducing overall processing cycles.
A computing device configures vector operation units to execute composite operations defined by specific parameters.
Input DMA circuit transfers commands to state management unit before data reaches the data path unit, eliminating command analysis delays during processing.
A parallel computing architecture uses three-dimensional stacked interconnects to accelerate machine learning workloads.
A tensor processor with a parallel processing element array reduces time for tensor operations by distributing workload across multiple units.
A communications bridge routes data between distinct processing units using separate connections to optimize resource usage.
Dedicated write paths isolate processor outputs from shared read buses, reducing arbitration conflicts and increasing processing throughput.
Control chiplets receive discovery signals to map physical topology, enabling dynamic workload distribution that reduces thermal hotspots and latency.
A parallel processor calculates operand and result addresses based on field of action position and elementary processor topology.
A block-based processor architecture executes instruction blocks atomically to sustain high performance with near in-order power efficiency.
Weight sequencer units shift inputs across the array to reduce memory access latency and improve processing speed.
Segmented crossbars reduce chip area and power consumption in SIMD architectures.
Parallel vector processing with a Benes network accelerates heterogeneous stream handling by eight times compared to scalar methods.
Mini-cores map loop iterations via local schedulers while a global scheduler adjusts mapping relationships to generate loop skew.
A hierarchical multi-core processor uses tree-like computing planes to decompose and map application sub-functions across parallel cores.
Conflict detection instructions identify duplicate values within SIMD registers to enable parallel binary tree reductions.
Processor architecture uses execution lanes coupled to a two-dimensional shift register array for simultaneous image stencil processing.
Balanced binary trees in tensor processors cancel delay dependency between aggregation logics, reducing wiring demands and improving productivity.
Segmenting the graphics pipeline crossbar into three levels reduces fan-out and gate count while maintaining non-blocking data routing.
A configurable control barrier network registers source tokens and status signals to produce enable signals for processing units.
Replicates ALU inputs across SIMD units for parallel execution, detecting faults without increasing hardware overhead.
A digital signal processor compute array executes instructions through successive engines to boost processing speed.
A processor array of resistive processing units performs analog computations via programmable SIMD instruction sets.
A tagged interrupt forwarding mechanism routes I/O responses to the originating processor using request descriptors.