A multi-port vector register file stores and retrieves data elements across parallel input lanes, enabling simultaneous CNN and RNN processing with low latency.
A register access method uses separate write and read select inputs to enable direct CPU control.
Dual horizontal pixel line searches reduce computational intensity and power consumption by enabling faster sum of absolute differences operations.
Dynamic reconfiguration of shared register banks reduces computational overhead while maintaining security against quantum attacks.
A processor bypass unit extracts valid processing results from an arithmetic logic unit and directs them to a multiplexer for register file access.
Hardware exception control circuitry saves register state to reduce processing delays during secure domain transitions.
Different gate oxide thicknesses in a CMOS transistor pair create stable unclonable chip IDs through BTI, avoiding e-fuse decoding vulnerabilities.
A computer system overrides function behavior at runtime by copying replacement code to memory addresses.
A VPDELTAENCODE instruction performs vector packed delta encoding on processor registers.
A vector floating-point scale instruction adds scale values to exponent fields within processor lanes.
A neural processing device uses universal elements to handle multiple data formats via conversion signals.
Compiler transforms loops with local variables into pipelined code using buffer storage to overcome FPGA internal capacity limits.
A processor training table system predicts load data values using register storage to bypass memory reads and reduce Level-1 cache latency.
A network switch with a direct memory access controller transfers data between host system memory and network devices.
Shared local registers isolate thread team memory from global bandwidth contention, reducing latency while maintaining high parallel processing capacity.
IOMMU extensions resolve performance overheads by enabling direct memory access to private memory through segmented trusted translation tables.
Segmented integration points manage version dependencies to reduce runtime errors during non-disruptive storage system upgrades.
An arithmetic unit executes immediate multiplication via shift and add operations to reduce logic depth.
A reconfigurable circuit matrix dynamically selects register banks to optimize data transfer paths.
A ZZYX processor architecture streams data through multiple ALU-Blocks to enable scalable processing across varying computational workloads.
A banked physical register data flow architecture assigns age tags to instructions and allocates physical registers to enable efficient execution.
A register allocation method dynamically controls static and rotating registers to optimize file usage.
A processor renames registers using standard memory maps to track data producers and manage physical register availability.
A processor manages register maps using map entries to enable indirect access to an extended set of actual registers via SIMD-oriented updates.
Segmenting the register file into slices reduces hardware cost and fabrication difficulty while enabling automatic port selection for n consecutive entries.
Segmenting a register file into banks mapped to thread identifiers reduces processor die area and power consumption while maintaining computational throughput.
A chaining table links command buffers for GPU processing without explicit copying, reducing memory overhead and CPU operations.
Segmented storage reduces circuit area and power consumption while maintaining architectural register compliance.
Pre-startup register free list excludes defective registers, eliminating complex repair logic and reducing register file size.
Segments inserted into intermediate code verify runtime execution against valid paths, reducing resource consumption on low-end embedded devices.
Partitioning the vector register file into disjoint portions enables thread migration between cores with different maximum vector lengths.
A processor executes a vector reverse instruction to reorder data elements within a single operation cycle.
Enqueue submission model carries quality of service information within command data packets to enable granular rate control.
A generation tag memory tracks register loads against global counters to block speculative execution paths.
Consolidating separate register files into a single physical register file reduces power consumption and die-area by eliminating unnecessary data copying.
A stateless capture mechanism records data linear addresses during precise event-based sampling using a register and counter.
Bridge chipset detects cable resistance to determine USB type and version for adaptive power management.
Segmented holding registers allow concurrent register access by buffering partial updates, eliminating semaphore overhead and improving throughput.
Speculative writes to special-purpose registers allow younger instructions to proceed without stalling.
Dynamic register remapping optimizes access to incomplete physical connections, preventing micro-operation splitting and reducing power consumption.
Grouping read port logic by entry simplifies interconnection routing, reducing wire congestion and improving area efficiency in high-density designs.
XOR-based control circuitry coordinates switching between dual-port banks to emulate multi-port memory access.
Detection logic merges spatially close load and store operations into one instruction, reducing memory subsystem strain and increasing effective bandwidth.
Separate push and pull request buses eliminate collision scheduling complexity while a turn-back bus maintains high data transfer throughput.
Vector checksum instruction accumulates operand elements using end-around carry add operations within vector registers.
A VSHPOP instruction shuffles vector data elements and performs arithmetic operations in a single execution step.
An instruction controller routes commands through dedicated pipelines to parallel function units.