The invention discloses a voiceprint continuous
tracking system based on multi-mode space-time fusion and SE (3) manifold optimization and a tracking method thereof. The tracking method comprises the following steps: S1, acquiring an acoustic
signal from a
microphone array deployed on a user wearing device, inertial data from an
inertial measurement unit, and optionally visual data from a camera; s2, estimating the real-time
pose of the equipment worn by the user, wherein the
pose is expressed on the SE (3) manifold; s3, defining a joint
state vector of the sound source, wherein the
state vector comprises a three-dimensional position and a voiceprint
feature vector of the sound source in a world coordinate
system; and S4, constructing a
Bayesian filtering framework running on the SE (3) manifold, and utilizing the real-time
pose. A
wave field synthesis theory and a neural
radiation field technology are combined, an extended acoustic NeRF model is constructed, a dynamic holographic
sound field and spatially distributed voiceprint features are reconstructed in real time by using an estimated sound source state and a glasses pose, and
visualization is carried out in XR equipment.