This invention discloses a method, apparatus, device, and medium for
digital human live streaming interactive processing that integrates a large
language model. It relates to the technical field of
digital human live streaming interactive processing, including
multimodal data acquisition and preprocessing steps,
sentiment analysis and semantic understanding steps, response
content generation and
personalization enhancement steps, and multimodal interactive output steps. By constructing a
multimodal data acquisition and preprocessing mechanism, it integrates input information from three modalities: voice, text, and video / image, achieving comprehensive
perception of the audience's interactive intent in the
live streaming scenario. Based on this, it utilizes a large
language model for deep semantic understanding and
contextual awareness, and combines a multimodal
sentiment analysis module to comprehensively infer the audience's tone,
speech rate, facial micro-expressions, and bullet screen text, enabling the
digital human to accurately capture the audience's fine-grained emotional state changes, thereby generating highly human-like, context-appropriate response content.