Segmenting branch history into a dedicated taken-branch table reduces storage and power consumption while maintaining misprediction recovery reliability.
Segmented execution lanes enable parallel processing of smaller data widths, reducing cycle counts and power consumption without increasing silicon complexity.
Token-based pipelined architecture enables instruction execution units to fetch subsequent instructions before releasing previous ones.
Constraint solver partitions machine learning computation graphs into ordered execution pipeline stages.
An optimization processing unit uses stochastic computing units and programmable interconnects to generate weighted responses.
The xIMD computing system forks execution lanes at data-dependent branches, resolving SIMD stall bottlenecks by enabling independent instruction streams.
A branch predictor monitors register writes to store values in a table for jump table switch statement prediction.
A distributed streaming system shares incoming data samples with applications upon arrival at the I/O interface for immediate processing.
A branch target array stores predicted addresses using instruction tags, reducing pipeline congestion and improving processor throughput.
Walk instructions iterate SIMD thread subsets to optimize channel processing speed.
A programmable packet processing pipeline diverts data to processor cores for specialized operations.
A peridynamic method adds mirroring nodes to boundary regions, completing horizon interactions through shape tensor calculations.
Semantic ordering manages compute slice tasks by checking address aliasing between load and store instructions to prevent data corruption in parallel execution.
Multiplexer tree indexing executes hashing and row reduction in parallel to eliminate serial processing bottlenecks and reduce read access time.
Threads redistribute multi-sample processing workloads using rasterized coverage information to reduce the number of required processing passes.
A failover processor executes operations from a failed primary unit using instruction set translation to maintain system functionality.
A bus master uses mode information in requests to handle error notifications for speculative accesses.
An instruction fetch unit retrieves neural processor instructions out of order based on dependency information.
A load store unit selectively forwards return addresses based on call instruction tags.
Dynamic quantization parameter selection reduces computational errors in float-type matrix multiplication by adapting constraints to input data ranges.
A central processor decodes instructions and sets a register tag for the coprocessor destination.
An embedded compute engine with arithmetic logic units processes data directly within a memory device.
Scalar core integration assigns workload subsets to dedicated processor complexes, resolving preemptive page fault bottlenecks in parallel graphics pipelines.
Dynamic partitioning isolates branch predictor states across privilege levels, preventing information leakage while maintaining high prediction accuracy.
Dynamic load balancing across multiple decode clusters resolves underutilization bottlenecks to improve processor pipeline throughput.
A branch prediction mechanism segments fetch addresses into index and tag bits to enable one-cycle latency.
A shared prediction learning table stores load value hashes alongside address data to unify hardware resources.
A local compiler generates executable object code from source definitions to configure programmable pipeline devices without external connectivity.
A dispatch stage routes operations to a temporary queue when execution resources are unavailable, allowing subsequent instructions to proceed without delay.
A specialized instruction calculates value differences to handle out-of-range data in vector processing environments.
A relay prioritizes data withdrawal from input buffers using state signals to manage pipeline throughput.
Dynamic width adjustment reduces power consumption by executing only active threads, eliminating wasted energy from predicated-off operations.
A re-triggering wake-up mechanism synchronizes vector micro-operations with scalar pipelines using secondary wakeup signals.
A processor scheduler dispatches interruptible instruction batches to functional units while storing batch-level resources in memory.
A two-level adaptive predictor updates global branch history using bitwise XOR operations on shifted bits and branch signatures.
Allocating store queue entries at dispatch time eliminates delays in store-to-load forwarding caused by waiting for physical address translation.
A multi-way pattern history table indexed by a global path vector determines multiple branch predictions simultaneously.
Dynamic loop partitioning maps iterations to scalable ALU blocks, resolving energy-performance tradeoffs across diverse platforms.
A processor analyzes user input to generate concise issue descriptions and confidence-tagged labels using natural language processing.
A coprocessor architecture with a scheduler circuit and hardware engines that parse and schedule directed acyclic graphs to execute multidimensional vector operations.
Prefetch throttling circuitry manages transaction issuance across interconnect routes using congestion tracking indications.
A branch prediction module creates execution path identifiers by hashing current and previous instruction addresses to search a prediction table for target addresses.
Front-end extensions detect zero-optimizable instructions to bypass execute and writeback stages in processor pipelines.
Dynamically adjusts mini-batch reuse counts via validation error feedback to reduce CPU-GPU bandwidth bottlenecks.
Dynamic thread deactivation based on cache miss detection resolves performance stalls and reduces power consumption in graphics processors.
A load store queue manages instruction entries through dynamic capacity allocation to optimize processing throughput.
Transposed memory stores weight coefficients by output channel order to enable parallel product-sum operations.
Adjusting asynchronous pipeline stage speeds using completion status signals to conserve energy.