The invention discloses a hardware-friendly Transform column balance
pruning model compression and efficient deployment method. A
model compression algorithm, a lightweight parameter storage format, an operation
data buffer, a
systolic array operation block, a vector operation unit, a nonlinear operator unit, a data flow controller and a DMA unit are included. A
model compression algorithm and an efficient deployment architecture are explored according to Transform
network software and hardware collaborative reasoning requirements: in a
software level, the scale calculation complexity of
model parameter quantities is reduced through a fine-grained column balance structured
pruning strategy, and parameters are stored in a single-instruction multi-data-
stream format and the parameter
storage efficiency is optimized through
mask code storage; according to the hardware level, an
edge computing-oriented Transform special accelerator architecture is designed, so that the architecture can support column balance structured
pruning characteristics and a lightweight parameter storage scheme in an original manner. According to the Transform model compression and efficient deployment method, the parameter sparsity after structured pruning is fully utilized, so that the parameter storage pressure of a
hardware architecture is reduced, the complex balance of an arithmetic unit is ensured, the operation efficiency of an accelerator is improved, and load balance and efficient reasoning during
software and hardware collaborative optimization are realized; the method is widely applicable to efficient deployment scenes of Transform models for edge calculation.