The invention is suitable for the field of audio
processing, and discloses a real-time voice inflexion method,
terminal equipment and a storage medium. The real-time voice inflexion method comprises the steps of generating original
voice data according to a real-time dialogue audio, and determining condition features and diversity features and filling data masks according to the original
voice data; determining first
tensor information according to the condition features, the diversity features and the filling data
mask, and determining a speaker embedding vector according to the original
voice data; determining second
tensor information according to the first
tensor information, the speaker embedding vector and the filling data
mask; and generating a target
timbre audio according to the second tensor information, the speaker embedding vector and the
pitch frequency of the original voice data. According to the method, the reconstruction precision of the original
timbre characteristics in the voice changing process is remarkably improved, the generated voice reaches the real person-like level in
perception dimensions such as
timbre similarity and intonation naturalness, and the voice changing authenticity of the real-time voice is improved.