The application discloses a
model inference method and device, equipment and storage medium, relates to the technical field of
large model application, and comprises the following steps: obtaining regular training data suitable for an adaptive
data synthesis scene, a target model to be deployed and a draft model; performing parallel training on the draft model according to the regular training data and a preset parallel decoding strategy to obtain a parallel draft model; modifying the
source code of an
inference framework suitable for a domestic
graphics card, inserting parallel draft decoding logic, and obtaining a customized
inference framework supporting parallel draft
inference; and based on the customized inference framework, jointly inferring the target model and the parallel draft model, accelerating the
processing of an inference request to be sent, and completing effect
verification. The application jointly infers by constructing a parallel draft model and adapting a domestic
graphics card inference framework, effectively alleviates the problem of shortage of domestic
graphics card computing resources, overcomes the defects of insufficient
adaptation of an existing inference framework and high inference
delay, and significantly improves the
model inference efficiency.