The invention discloses a storage-efficient
large model reasoning method for many-core supercomputing, and relates to the field of
machine learning. The overall architecture of the method is based on
video memory optimization of Attention and FFN modules in a Transform model, and high
storage efficiency in a reasoning stage is realized through matrix block calculation and dynamic
video memory management. The method comprises the following steps: firstly, carrying out parameter partitioning and serial calculation: in an Attention module, vertically
cutting a Q and K parameter matrix into a plurality of sub-blocks along a column direction, and keeping the input complete; serially calculating the product of each sub-block and the input to obtain a local
Q matrix and a local
K matrix, and immediately performing QKT block multiplication to obtain a partial attention
score; aggregating calculation results of all the sub-blocks to obtain a complete attention
score matrix; afterwards, normalizing the attention
score by using Softmax, carrying out serial multiplication on the normalized attention score and a V block subjected to
delay calculation, and splicing a result to obtain Attention output; optimizing an FFN module, vertically
cutting parameters of a full connection layer into sub-blocks, inputting
complete data, sequentially carrying out serial calculation on the
complete data and the sub-blocks, and splicing
nonlinear transformation results of all the sub-blocks in real time.