The invention relates to a digital conference voice
processing method,
system and device and a storage medium, and the method comprises the following steps: carrying out the
pickup of a conference voice, obtaining a mixed voice
signal, carrying out the framing sampling, and forming a voice sampling sequence; carrying out sound source direction
estimation based on the sequence to obtain multi-sound-source position information, and carrying out beam forming and spatial filtering on the voice sampling sequence according to the multi-sound-source position information to obtain a
sound source separation signal; performing voice segment segmentation on the
signal to obtain a voice segment sequence, extracting voiceprint features of each segment, and generating a speaker feature mark; and performing
time sequence recombination on the voice fragment sequence by using the mark, constructing a speaking
time sequence table, selectively outputting the voice fragment sequence according to the table, and finally generating clear and ordered conference voice. The technical problems that due to the fact that a traditional voice
processing method lacks effective space-voiceprint joint constraint, voice separation is not thorough, identities of speakers are confused, and the speaking
time sequence is disordered are solved.