The invention relates to the technical field of sound equipment, discloses a sound equipment playing control method and
system based on a voice sensor, and aims to solve the problem that audio experience is disjointed due to the fact that an existing sound equipment
system cannot sense a continuously changing acoustic environment and a
user state. According to the method, acoustic signals are collected in real time through a multi-channel voice
sensor array, a three-dimensional position track, an emotional state
label and a room
impulse response function of an audience are extracted, a multi-dimensional dynamic scene state descriptor is constructed in a fusion mode, and
sound field rendering parameters such as a direct sound / reflected
sound energy ratio,
reverberation time, an equilibrium curve and
sound image diffusion degree are generated according to the descriptor. And dynamically synthesizing a personalized three-dimensional
sound field matched with the current scene. The
system comprises an acoustic
signal acquisition module, a scene
feature extraction module, a state fusion modeling module, a rendering parameter generation module and a dynamic
sound field synthesis module. According to the method and the device, the conversion from passive instruction response to
active environment adaptation is realized, and the audio immersion and the man-
machine interaction naturalness are remarkably improved.