The application provides a storage-computation integrated device, an end-side AI
inference method, a medium and a terminal. Through direct
interconnection of a storage group and an AI computation module,
register transfer and parallel data carrying mechanism,
system performance and energy efficiency are improved, storage
delay and bandwidth are optimized, serial-parallel conversion overhead is saved,
cache management overhead is avoided,
power consumption and area of prefetch logic are saved, complexity of the storage
system itself and data carrying is reduced, and storage carrying is lightweighted. In the process of executing a distributed AI
inference operation, data carrying and shaping can be synchronously performed,
data synchronization parallelism is improved, data supply rate is matched with computation rate, idle waiting of an AI computation unit is reduced, and
system throughput and energy efficiency ratio are improved.