The invention discloses an efficient and parallel autoregressive
diffusion model
hardware acceleration system based on an FPGA (
Field Programmable Gate Array), and relates to the technical field of FPGA and
machine learning. The
system comprises a host end and an FPGA hardware end, the host end is responsible for preprocessing and post-
processing of data, and the FPGA hardware end comprises a
system control scheduling layer, a three-level
storage structure layer and an autoregressive
diffusion parallel architecture. According to the method, a collaborative pipeline design is adopted, deep parallelism of an autoregression Token
generation process and a
diffusion picture
generation process is realized through a Token grouping adding mechanism, and a diffusion denoising process can be started without waiting for the completion of
complete sequence generation. Besides, a universal attention
processing hardware module is designed in the system, QKV projection and
matrix multiplication are realized by using a
systolic array, and calculation resources are optimized in cooperation with a pipelined Softmax calculation and
delay normalization strategy. Through
collaborative design of
software and hardware, the problems of high reasoning
delay and large resource overhead of the autoregression diffusion model in an
edge computing scene are effectively solved, and the
throughput and the energy efficiency ratio are remarkably improved while the generation quality is ensured.