A repeat control instruction updates source and destination addresses automatically during hardware accelerator execution.
A processor instruction calculates the difference between a floating point value and its rounded version to determine the exact round-off amount.
A trace suppressor generates suppressed address traces by selectively omitting redundant data from processor execution logs.
Produce and consume commands delay encoding to reduce redundant state settings, resolving graphics processing efficiency bottlenecks.
Distributing registers across multiple banks reduces bank conflicts and pipeline delays without increasing hardware complexity.
A central processor merges original register setting commands with address continuity into fewer merged commands.
A processor unit with vector registers and an M×M comparator matrix performs character-wise comparisons between reference and target strings.
Compiler assigns instructions to processing clusters and configures global virtual registers, reducing access time and energy consumption.
A processor architecture moves data directly between integer and floating-point register files using an integer execution unit.
Machine instructions convert zoned EBCDIC data directly into decimal floating point registers.
A register map table stores narrow produced values directly to reduce physical register file port pressure.
Compacting register tags within superslices rebalances architected resources, preventing synchronization stalls and collisions during mode reduction.
A register file splits validity values from data to detect operation status without complex predicate routing.
Explicitly addressable forwarding registers bypass the VLIW processor register file, reducing detection logic complexity and power consumption.
Dynamic lane coupling and register renaming enable efficient vectorizable loop processing while reducing silicon area and energy dissipation.
A SHIFTINSIDE instruction calculates segment address offsets by multiplying coordinates with dimension sizes.
A matrix operations accelerator circuit uses a two-dimensional grid of fused multiply accumulate units to execute decoded single instructions.
Decoder circuitry merges operand size conversion with data operations into single instructions, reducing energy consumption and improving processing speed.
A graphics processing unit integrates unified processing clusters with a scheduler to distribute work across parallel cores.
A unified graphics processing unit execution unit handles integer and floating-point operations through dynamic instruction selection logic.
Segmented register files and switching circuitry eliminate quadratic port growth while maintaining parallel execution speed.
Checkpoint storage elements transfer register mapping information to capture architectural states without moving actual data.
A dedicated computation engine uses extract instructions to move data between local memories without accessing main memory.
A multiported parity scoreboard circuit segments busy-state data across multiple SRAM modules to enable concurrent register tracking.
An address adjustment unit translates virtual addresses to access non-contiguous memory areas as a continuous linear block.
A vector register bank write interface includes a data rearrangement path for simultaneous element reordering.
Vector execution circuitry processes Keccak permutations within a single lane to eliminate cross-lane access latency while maintaining high data throughput.
Bypassing execution units via a constant register file reduces pipeline conflicts and conserves bandwidth.
A single I2C slave circuit analyzes data flows to generate address messages for multiplexer routing.
A processor moves load instructions to a destination memory to execute data transfers without stalling.
A modified enclave resume instruction with return-to-handler functionality manages trusted execution environment transitions.
A graphics processor execution unit suppresses redundant register file reads by replicating data for multiple source operands.
Shared register pools let GPU threads request accumulators on demand, reducing thread completion time without expanding physical hardware.
A sequential equivalency check method identifies registers where faults cannot propagate to minimize circuit logic duplication.
Specialized instructions reduce data movement and processing time during the Smith-Waterman matrix-filling phase on parallel processors.
PredCount and SegCount instructions determine active elements and segment counts within processor vectors to enable dynamic parallel execution.
Standard PCI interfaces enable cross-platform trusted execution environments, resolving hardware platform specificity limitations.
A dynamic operations registry tracks registered operation counts to enforce execution limits based on user and group account quotas.
Reordering assembly instructions separates dependent operations, preventing pipeline stalls and reducing clock cycles during cross-architecture conversion.
Instruction translator splits macro-instructions into atomic micro-instructions to handle wider data bit widths while maintaining correctness.
A vector instruction generates result vectors with incremental values based on start and increment inputs.
A FPGA device uses a PCIe module in its management logic unit to enable remote configuration and debugging without physical JTAG cables.
A channel convolution processor unit executes matrix operations using vector units and data reuse.
Abort event detection circuitry captures syndrome information for transactional memory exceptions.
Copying state information from a master register to per-group registers lets processing engines read locally, balancing vertex and pixel shader loads.
Per lane predicate information enables segmented operations on irregular data structures to optimize parallel processing utilization.
A preload predictor and data buffer fetch cache data before ALU execution, eliminating stall cycles caused by load-use instruction pairs.
A move-immediate logic circuit allocates physical registers from an immediate physical register file to enable early instruction execution.
Programmable counters track executed instructions by register and element size to count floating-point operations accurately.