The invention provides an accompanying watching method and device of a virtualized image,
electronic equipment and a medium, and the method comprises the steps: taking a preset time length as an interval, intercepting an
image frame of a real-time video
stream, recognizing subtitles in the
image frame, obtaining line subtitles of the
image frame, inputting the image frame and the line subtitles into a target multi-mode
large model, and carrying out the accompanying watching of the virtualized image. And guiding the target multi-
modal large model to perform fusion
sentiment analysis on the visual content of the image frame and the text
semantics of the line subtitles through a structured preset template constructed based on prompt word
engineering, so that the target multi-
modal large model outputs at least one of interactive actions and interactive expressions corresponding to the image frame and the line subtitles, and performing mapping
processing on the interactive action and / or the interactive expression, obtaining an action
image sequence and an expression
image sequence, performing
animation fusion, generating a target
interactive animation, and outputting the target
interactive animation. For an online drama tracking scene, anthropomorphic response based on a dynamic situation can be provided.