The invention relates to a
simulation digital human real-time intelligent voice interaction
system and a
simulation digital human real-time intelligent voice
interaction method based on vision and a
large model, and aims to solve the problems of inaccurate target speaker recognition, high
response delay and the like in
digital human voice interaction in a complex scene. The
system circles an effective recognition range through a camera, triggers audio collection in combination with
face detection, locks a target speaker and reduces
noise by using lip
movement recognition and
sound image fusion technologies, converts the target speaker into a text through voice wake-up, generates an answer by means of a large
language model (LLM) and
knowledge retrieval enhancement (RAG) technologies, generates low-
delay voice through a voice synthesis technology accelerated by the vLLM, and performs voice recognition on the target speaker. And driving the preloaded digital human image to synthesize a video
stream and pushing the video
stream to a front end for rendering in real time. Accurate
pickup, low-
delay interaction and rapid digital human image switching in a complex environment are realized, the accuracy and real-time performance of intelligent voice
question answering are improved, and the method is suitable for government affair halls, exhibition halls and other scenes.