The invention discloses a zero sample ASMR generation method and
system based on a large
language model, and aims to solve the problems that in the prior art, personalized ASMR voice cannot be generated in a zero
sample mode, and a high-quality ASMR special
data set is lacked. The method comprises the following steps that speaker prompt and task signals of texts to be synthesized and normal or ASMR style voice are obtained; retrieving a matching task cue from the virtual
pool of speakers based on the
speaker verification system; pre-training the large
language model to generate a voice token sequence containing a target style; the
stream matching acoustic decoder fuses the voice token sequence and speaker acoustic information to generate a target Mel spectrum; the vocoder synthesizes a target audio; the
system comprises an input module, a virtual speaker
pool module, a large
language model style coding module, a
stream matching decoding module and a vocoder module, and further comprises a
data set module for storing a DeepASMR-DB
data set covering 9 types of themes, Chinese and English and over 670-hour voices. According to the method, zero-sample ASMR generation is realized, the tone and real
breath sound of a speaker are reserved, and the audio fidelity is high.