The invention discloses an auto-regression and
mask diffusion double-normal-form fusion
language model hardware acceleration system based on an FPGA (
Field Programmable Gate Array), and relates to the field of FPGA and
machine learning. The invention provides a model core calculation process-oriented
hardware acceleration architecture by aiming at a
language model fusing double normal forms of an autoregression model and a
mask diffusion model and adopting the characteristics of revising an attention mechanism and
parallel generation of a KV cache. The method disclosed by the invention is implemented by taking an Eso-LMs (Esolic Language Models) model as an example. The
system comprises a layer normalization and adaptive layer normalization modulation module, a QKV projection module, a rotation position coding
application module, a KV
cache management module, a multi-head attention calculation module, an output projection and residual connection module and a multi-layer
perceptron module. By optimizing the calculation sequence and the data flow of each module in the model calculation process and adopting the pipeline parallel and resource reuse technology, the efficient reasoning acceleration of the
language model fusing the autoregression and
mask diffusion double normal forms on the FPGA platform is realized, and the model reasoning speed is obviously improved. According to the
system, the
advantage of FPGA customizable
hardware acceleration is fully exerted, the utilization efficiency of hardware resources is improved, data pipeline blockage is eliminated, calculation
delay and storage overhead are reduced by reconstructing the data flow direction and constructing a whole-process pipeline
processing architecture, and the system is suitable for the deployment requirement of an edge calculation scene.