A data path block circuit uses operand multiplexers to interface with specialized register files having distinct read and write port configurations.
Masking circuitry skips irrelevant rows and columns in matrix processing, eliminating data remapping overhead and reducing latency.
A Vector Load Immediate Decimal instruction generates signed packed decimal values directly within a register using sign control and shift operations.
A predecoder merges legacy instructions into single executable units for wide datapath processing.
Non-memory mapped bank select register allows contiguous access to larger memory blocks without special function registers interrupting data areas.
A vector friendly instruction format executes simultaneous multiplication, summation, and accumulation of packed signed bytes within a single operation cycle.
Segmenting the branch target address across the fixed-width instruction and separate memory extends the address range without increasing decoding complexity.
Processor randomizes register mappings at subroutine calls to disrupt return-oriented programming attacks without adding hardware complexity.
New instructions merge multiplication, rounding, and saturation into single operations to resolve precision versus versatility trade-offs.