The application provides a
large model inference acceleration method, device, equipment, storage medium and program product, and relates to the technical field of
artificial intelligence, and the method comprises the following steps: inputting an interactive prompt word and historical dialogue into a large
language model for self-recurrence decoding to obtain a current word element; it is detected whether the current word element is a text operation semantic word element; if the current word element is a text operation semantic word element, a
target text segment is copied from existing text information based on the text operation semantic word element; a plurality of target word elements are updated to the end of the output text; based on the large
language model, self-recurrence decoding is continued to generate a next word element, the next word element is taken as the current word element, and the step of detecting whether the current word element is a text operation semantic word element is returned to until the large
language model generates an end symbol, and an updated output text is obtained. Through the above manner, the text
copying efficiency is improved, the
inference acceleration effect of the large language model is optimized, and the accuracy and stability of the
copying operation are beneficially ensured.