The invention discloses a method for accelerating reasoning of a multi-
modal large
language model based on speculative decoding, the input of the multi-
modal large
language model is composed of a
system lexical element, a visual lexical element, an instruction lexical element and an output lexical element, and the method comprises the following steps: S1, extracting a predetermined number of samples from a ShareGPT
data set which is finely adjusted by using an LLaMA model instruction, pre-training an
Eagle small model, and obtaining a pre-trained
Eagle small model; enabling the
Eagle small model to establish a
basic language generation capability in a
plain text environment; s2, extracting a predetermined number of samples in a ShareGPT
data set finely adjusted by an LLaVA model instruction, and migrating the capability of an Eagle small model by using a transition mode; s3, after the draft model is subjected to two-stage training, on the basis of the improved draft mechanism design, differential
processing is carried out according to the characteristic difference between the text and the visual mode; according to the method, the performance of speculative decoding in a multi-
modal task can be improved, the acceptance length of each
verification of the target model is improved, and respective
adaptation of the text lexical elements and the visual lexical elements in the draft model is realized.