The invention belongs to the technical field of
natural language processing, and provides a text history-based reasoning method and device in a multi-
modal dialogue scene, a medium and a program product, and the method comprises the steps: obtaining a historical dialogue of voice and text alternation, and carrying out the automatic voice recognition of a voice part in the historical dialogue to obtain a recognition text; extracting non-text feature information at least containing
mood, emotion or emphasis when the recognition text is generated, compressing and encoding the non-text feature information into a
voice tag, and adding the
voice tag to the corresponding recognition text to form historical text information with a semantic compensation mark; and then combining the historical text information and the
current user input voice into model input content, inputting the model input content into a multi-
modal dialogue model for reasoning, and outputting a corresponding text reply. According to the method, the length of a model input sequence can be shortened, the reasoning efficiency is improved, meanwhile, key non-text information is reserved, the dialogue accuracy is improved, and the multi-
modal dialogue interaction experience is optimized.