An indirect branch instruction enables function calls in single-instruction multiple-thread processors by accepting address registers as arguments.
A unified shader core balances vertex and pixel loads by executing diverse programs on shared clusters, reducing idle cycles in graphics processors.
Chaining engines via a super-descriptor reduces firmware overhead and enhances throughput by enabling autonomous data flow processing.
Parallel memory banks with a shifted storage scheme resolve silicon area and power bottlenecks while boosting processing throughput for SIMD processors.
Horizontal aggregation SIMD instructions accelerate sorting speed by reducing comparisons in K-heaps, resolving sequential processing bottlenecks.
A parallel processing device implements contention-free routing schedules to optimize data movement across tile arrays.
A multi-core image processor uses a configurable internal network to dynamically couple specific cores for targeted execution.
A random matrix network simplifies backpropagation using fixed weight matrices for faster learning.