The application relates to the technical field of
digital human voice interaction, and discloses a
digital human live broadcast voice interaction
system fusing emotional computing. The
system acquires user voice input data streams, constructs an emotional
response time window, the starting point of which is the end time of the user voice input data streams, and the ending point of which is the preset maximum response cutoff time minus the necessary duration of voice response synthesis. The
system acquires the emotional
state vector of the user in real time, and judges whether the emotional
state vector reaches an emotional intensity threshold value; if the emotional
state vector reaches the threshold value, the total generation duration of the
digital human voice response is predicted. When the residual duration of the emotional
response time window is equal to the total generation duration, the starting point of the window is taken as the starting time of the voice response, and the digital human is controlled to start voice
response generation. Through fine time window management and emotional state
perception, the system generates voice responses that are adapted to the emotional state of the user, reduces invalid responses and interaction delays, and improves the naturalness of digital human live broadcast voice interaction and user experience.