The invention discloses a voice
signal processing method for an education
robot, relates to the technical field of voice
signal processing, and aims to solve the problems of multi-person overlapping language and
environmental noise interference in a classroom, multi-
channel data is acquired by relying on a
microphone array and a camera, and an
environmental model is constructed in the step 1 to determine a
noise baseline and student distribution; in the second step, overlapped voices are detected, and sound
source localization is carried out in combination with the
time difference of arrival and
mouth shape data; in the third step, directional
gain is executed in the target direction, and a deep network is used for separating
aliasing voice; and in the step 4, the separated voice is input into a children customized recognition engine to complete high-precision recognition and interaction in combination with confidence evaluation. The recognition accuracy and the interaction efficiency can be remarkably improved under the complex scenes of classroom
reverberation and simultaneous speaking of multiple persons, meanwhile, the
noise change is tracked through the global environment model so that the education
robot can keep stable recognition performance in diversified teaching interaction, and the teaching effect is remarkably enhanced.