Enabling high-performance Scalable Matrix Extension (SME) instruction issue in processor devices is disclosed herein. In some aspects, a processor device comprises a
reservation station circuit configured to perform, during a first phase, a reduced-precision vector accumulator (ZA) tracking operation on micro-ops for which corresponding vector (Z) registers and corresponding predicate (P) registers are ready. Based on the reduced-precision ZA tracking operation, the
reservation station circuit selects a first micro-op and a second micro-op having no Read-After-Write (RAW)
hazard with respect to the ZA registers. During a subsequent second phase, the
reservation station circuit performs a full-precision ZA tracking operation on the first micro-op and the second micro-op, and selects one as a micro-op for issue for which the full-precision ZA tracking operation indicates no RAW
hazard exists with respect to the ZA registers. The reservation
station circuit then issues the selected micro-op for execution.