The invention belongs to the technical field of
computer graphics and human-computer interaction, and particularly discloses a multi-
modal expression generation
system and a dynamic optimization method in a virtual-real fusion scene, and the method comprises the steps: analyzing an audio
rhythm to generate a phase reference
signal, driving visual collection and calculation, and achieving the sound and picture synchronization of a microscopic
time domain; calculating a dense
optical flow field, and decoupling a facial
muscle movement change vector through a vector dot product by using a
muscle movement direction template; and mapping the audio acoustic feature into an audio driving energy value representing the
sound production intensity. An
energy conservation attenuation or compensation enhancement correction is performed on the
expression vector by comparing the total amount of visual displacement with the audio drive energy to generate a corrected
expression vector. And finally, superposing the corrected expression change vector to the
state parameter of the previous frame, and generating a control instruction to drive the virtual avatar to render in real time. According to the method, the problems of sound and picture
time sequence dislocation and physical constraint missing are solved, and the sense of reality and expressive force of virtual expressions are improved.