Splitting circuitry allocates threads into sets aligned with storage boundaries, reducing redundant memory access requests and improving processing throughput.
An intermediate buffer with priority-based eviction reduces primary storage bandwidth usage and energy consumption by minimizing frequent data transfers.
Parallel selection of shifted data groups reduces memory accesses and power consumption while boosting processing speed for image filters.
A multi-threaded processor records cycle counts to synchronize execution traces across independent threads.
A processing method identifies loop parameters directly from hardware-recognizable instructions to execute loop bodies accurately.
An overlapped immediate register field shares bits between immediate values and register specifiers to expand instruction encoding capacity.
A microprocessor uses a time counter to statically dispatch instructions based on preset execution times.
A compiler program reorders unrolled vector load and store instructions to bypass architectural dependencies.
A CPU die integrates connection determination logic to detect motherboard interface types and adjust execution policies.
An enhanced look-up table uses bank segmentation to enable simultaneous read-modify-write operations across multiple address banks.
Boundary detectors track special symbols to update offset registers, correcting segmentation drift when phase-locked loops lose lock.
Merging separate read and write operations into one cycle reduces communication time while maintaining data transmission reliability.
Decoder writes immediate values directly to registers, bypassing standard read operations that increase arithmetic unit busy rate and reduce processing speed.
A vector register access circuit performs data rearrangement across multiple instructions to distribute computational burden.
A vector exception code embeds the specific element position causing a fault within a vector register during instruction execution.
An accumulator pool caches source operands to retrieve values without accessing the register file directly.
A processor instruction stores sorted indexes of source data elements in packed registers to accelerate sorting operations.
An ARM64 floating point emulator uses an instruction classifier to detect and dispatch specific operations without hardware coprocessors.
Arithmetic operator executes imaginary-number matrix multiplication-addition using extended instruction bits.
A microprocessor uses shared hardware registers to store architectural state for both x86 and ARM instruction sets.
A register usage mask identifies active registers during function calls, reducing processing overhead by saving only commonly used values.
Inhibiting register read access via a setting circuit holding specific values reduces power consumption while maintaining processing capability.
A parallel data processing circuit executes matrix multiplication using source operands accessed only once from a vector register file.
A general purpose register employs a conflict queue to store conflicting operand values, allowing parallel execution without read port bottlenecks.
Variable word length instructions and payload mechanisms expand instruction set capacity while reducing decoding complexity in matrix processing systems.
Wide register design minimizes memory accesses to resolve energy bottlenecks in biomedical signal processing applications.
Segmenting physical registers into parallel banks enables simultaneous instruction execution in out-of-order processors.
Calculating physical distances between master and servant circuit partitions determines the exact number of pipeline registers needed, reducing die area usage.
NDMA core performs hardware pre-processing on data blocks to reduce memory bandwidth pressure caused by input feature padding.
A dual-level branch target buffer system dynamically moves instruction entries between hierarchical storage tiers to optimize access speed.
Segmenting the format decoder and register file reduces latency while maintaining protocol flexibility.
A broadcast control unit writes operation results to multiple destination registers in a single clock cycle.
A vector processor shuffle unit routes input data to SIMD lanes via multiplexers without re-storing values in register files.
Segmented writing ports eliminate timing conflicts between execution units with different delays, boosting pipeline efficiency while reducing resource overhead.
A register file port sharing structure reduces silicon area and manufacturing costs by allowing multiple instruction pipes to share read ports.
Compiler groups partial register writes and hints hardware to combine them, eliminating redundant read-modify-write cycles.
Load circuitry stores excess data from full word loads in a stream buffer to suppress redundant initial operations for subsequent unaligned instructions.
A control apparatus modifies integrated circuit functions via authenticated signals to enable secure dynamic reconfiguration.
Memory controller determines optimal data layouts to supply operands to DRAM banks, eliminating timing overheads and boosting computational performance.
Dynamic speculation width adjustment manages overflow conditions in a reconfigurable buffer to support unpredictable loop iterations.
A move and zero instruction clears array storage elements while moving data to reduce processing overhead.
Split-point broadcast and move instructions flatten loops in SIMD pipelines, reducing instruction counts and energy waste during molecular dynamics simulations.
A processor delays common resource switching using a status flag to resolve the contradiction between resource consistency and switching speed.
A register file uses dedicated circuitry to set entries to predetermined values via control signals.
Software-visible PUF instructions generate platform-unique keys to encrypt memory, excluding CSP software from the trusted computing base.
Iteration level commits manage synchronization and load-store operations to resolve runtime latency mismatches in coarse grained reconfigurable architectures.
A computer processor uses asymmetric dual execution paths to separate control and data processing operations.
An arithmetic processing device stores input and weight data as matrices to perform row portion operations using element data from predetermined rows.
A hardware accelerator stores memory address translations from a multi-threaded core to enable seamless operation.
A neural network controller circuit uses multiple registers and address generation logic to process convolutional instructions concurrently.