The invention relates to the technical field of
artificial intelligence, can be applied to business scenes such as
medical health and financial science and technology, and discloses a model quantitative reasoning acceleration method, device, equipment and medium. The method comprises the steps that an input text is divided into a plurality of
processing blocks, importance scoring is conducted on the non-first
processing block, and calculation precision formats are distributed according to scoring results; determining a unified quantization configuration of each
processing block; dividing the network modules into configuration sharing groups, and sharing quantitative configurations of corresponding processing blocks in the groups; and executing block-level quantization
inference according to the unified quantization configuration, and generating a
model inference result. The quantitative configuration of each processing block is uniformly determined on the basis of the token importance
score, and the configuration is multiplexed in the
network module group, so that block-level precision distribution and parallel quantitative reasoning are realized, the
video memory overhead and the configuration
time overhead are greatly reduced while the reasoning precision is guaranteed, and the reasoning efficiency is improved. And the execution efficiency and the
video memory utilization rate in the long text reasoning task are effectively improved.