The invention discloses an end-to-end speech
separation algorithm based on a speech
language model, and the
algorithm comprises the steps: discretizing a continuous audio into a 32-order discrete
codebook sequence through a residual
vector quantization coder-decoder, and introducing a
transcription start symbol lt through an SOT strategy; sOSgt, SOSgt; a special separator is lt; sCgt; and a termination symbol lt; eOSgt, EOSgt; splicing a multi-person voice sequence; extracting audio depth features by using a pre-trained WavLM model, and guiding an autoregression decoder to output a separated zero-order
codebook sequence in combination with a cross attention mechanism; predicting a high-order
codebook sequence step by step through a non-autoregression model, configuring an independent embedding layer to fuse low-order information, and introducing a task embedding mechanism to optimize modeling; based on a special separator lt; sCgt; and
slicing the multi-order discrete codebook sequence, and outputting an independent
voice source through an Encodec decoder. According to the method, the intelligibility of voice separation and the
audio restoration quality can be effectively improved, the decoding speed is high, the subjective hearing experiment result and the downstream task performance are excellent, the scene that the number of speakers is unknown is supported, and the industrialization application prospect is wide.