A Voice Style Transfer Method, Device, Electronic Device and Storage Medium

By extracting and adjusting audio features, generating synthetic audio with the target audio style, the problem of insufficient accuracy of speech style transfer in the prior art is solved, and a more efficient speech style transfer effect is achieved.

CN113963679BActive Publication Date: 2025-06-17BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111262784.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-28
Publication Date
2025-06-17
Estimated Expiration
2041-10-28

AI Technical Summary

Technical Problem

Existing voice style transfer technologies are difficult to effectively improve the accuracy of voice style transfer, especially in maintaining the speaker's speech speed and style characteristics.

Method used

By extracting the spectral characteristics and phoneme duration characteristics of the target audio, the synthesized phoneme sequence is extracted and phoneme duration prediction is performed, the phoneme duration is adjusted to match the speech speed of the target audio, and synthetic audio with the target audio style is generated based on the spectral characteristics and content characteristics.

Benefits of technology

Improve the accuracy and effect of speech style transfer, making the generated synthetic audio closer to the speech speed and style characteristics of the target audio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113963679B_ABST
    Figure CN113963679B_ABST
Patent Text Reader

Abstract

The present disclosure provides a voice style transfer method, apparatus, electronic device, and storage medium, which relate to the field of artificial intelligence technologies, and particularly to the field of voice technologies. The specific implementation solution is as follows: extracting the spectrogram feature and phoneme duration feature of the target audio to be transferred, performing content feature extraction and phoneme duration prediction on the phoneme sequence to be synthesized, obtaining the content feature and predicted basic duration of each phoneme to be synthesized, adjusting the predicted basic duration of the phoneme to be synthesized based on the phoneme duration feature of the target audio to obtain the target duration of the phoneme to be synthesized, obtaining the target spectrogram with the style of the target audio based on the spectrogram feature of the target audio, the content feature of each phoneme to be synthesized, and the target duration, and performing audio conversion on the target spectrogram to obtain the synthesized audio. Applying the present disclosure enables better audio transfer effect and improves the accuracy of audio transfer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to the fields of deep learning, speech synthesis, speech recognition, and speech transcription, and specifically relates to a speech style transfer method, apparatus, electronic device, and storage medium. Background Art

[0002] Currently, speech style transfer technology is used in many fields, such as voice conversion systems, voice chats, etc. Speech style transfer refers to given the audio of a certain speaker, generating speech with the characteristics of that speaker for any text sequence. Summary of the Invention

[0003] The present disclosure provides a speech style transfer method, apparatus, electronic device, and storage medium for improving the accuracy of speech style transfer.

[0004] According to one aspect of the present disclosure, there is provided a speech style transfer method, including:

[0005] Obtaining a target audio to be transferred and a phoneme sequence to be synthesized;

[0006] Performing spectrogram feature extraction and phoneme duration feature extraction on the target audio to be transferred to obtain the spectrogram feature and phoneme duration feature of the target audio;

[0007] Performing content feature extraction and phoneme duration prediction on the phoneme sequence to be synthesized to obtain the content feature of the phoneme sequence to be synthesized and the predicted basic duration of each phoneme to be synthesized;

[0008] Based on the phoneme duration feature of the target audio, adjusting the predicted basic duration of each phoneme to be synthesized to obtain the target duration of each phoneme to be synthesized;

[0009] Based on the spectrogram feature of the target audio, the content feature of the phoneme sequence to be synthesized, and the target duration of each phoneme to be synthesized, obtaining a target spectrogram corresponding to the phoneme sequence to be synthesized with the style of the target audio;

[0010] Converting the target spectrogram into audio to obtain a synthesized audio corresponding to the phoneme sequence to be synthesized with the style of the target audio.

[0011] According to another aspect of the present disclosure, there is provided a speech style transfer apparatus, including:

[0012] An audio and phoneme sequence acquisition module, configured to obtain a target audio to be transferred and a phoneme sequence to be synthesized;

[0013] A target audio feature acquisition module, configured to extract spectrogram features and phoneme duration features from the target audio to be migrated, so as to obtain the spectrogram features and phoneme duration features of the target audio;

[0014] A phoneme sequence to be synthesized feature extraction module, configured to extract content features and predict phoneme durations for the phoneme sequence to be synthesized, so as to obtain the content features of the phoneme sequence to be synthesized and the predicted basic duration of each phoneme to be synthesized;

[0015] A phoneme duration adjustment module, configured to adjust the predicted basic duration of each phoneme to be synthesized based on the phoneme duration features of the target audio, so as to obtain the target duration of each phoneme to be synthesized;

[0016] A target spectrogram acquisition module, configured to obtain a target spectrogram with the style of the target audio corresponding to the phoneme sequence to be synthesized based on the spectrogram features of the target audio, the content features of the phoneme sequence to be synthesized, and the target duration of each phoneme to be synthesized;

[0017] A synthesized audio acquisition module, configured to convert the target spectrogram into audio to obtain a synthesized audio with the style of the target audio corresponding to the phoneme sequence to be synthesized.

[0018] According to another aspect of the present disclosure, there is provided an electronic device, including:

[0019] At least one processor; and

[0020] A memory communicatively connected to the at least one processor; wherein,

[0021] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute any one of the above-mentioned speech style transfer methods.

[0022] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute any one of the above-mentioned speech style transfer methods.

[0023] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, where the computer program, when executed by a processor, implements any one of the above-mentioned speech style transfer methods.

[0024] In the present disclosure, spectral features and phoneme duration features of the target audio to be migrated are extracted to obtain its spectral features and phoneme duration features. Content features of the phoneme sequence to be synthesized are extracted and phoneme duration prediction is performed to obtain the content features and predicted basic durations of each phoneme. Then, based on the phoneme duration features of the target audio, the predicted basic durations of each phoneme in the phoneme sequence to be synthesized are adjusted to obtain the target durations of each phoneme to be synthesized. Based on the spectral features of the target audio, as well as the content features and target durations of each phoneme to be synthesized, a target spectrum with the style of the target audio corresponding to the phoneme sequence to be synthesized is obtained, and audio conversion is performed on it to obtain a synthesized audio with the style of the target audio corresponding to the phoneme sequence to be synthesized. Applying the present disclosure, by combining the speech rate that has a greater impact on the style of the target audio to be migrated, speech style migration is performed, making the audio migration effect better and improving the audio migration accuracy.

[0025] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0027] Figure 1 is a schematic diagram of the first embodiment of the speech style migration method provided by the present disclosure;

[0028] Figure 2 is a schematic diagram of the second embodiment of the speech style migration method provided by the present disclosure;

[0029] Figure 3 is a schematic diagram of the third embodiment of the speech style migration method provided by the present disclosure;

[0030] Figure 4 is a schematic diagram of the process of training the content encoding model in the present disclosure;

[0031] Figure 5 is a schematic diagram of a process of the speech style migration method provided by the present disclosure;

[0032] Figure 6 is a schematic diagram of the process of training the style encoding model and the spectral decoding model in the present disclosure;

[0033] Figure 7 is a schematic diagram of the process of training and testing each model in the present disclosure;

[0034] Figure 8 is a schematic diagram of the first embodiment of the speech style migration device provided by the present disclosure;

[0035] Figure 9 It is a block diagram of an electronic device for implementing the voice style transfer method of the embodiments of the present disclosure. Specific Embodiments

[0036] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.

[0037] To better transfer the voice style, the present disclosure provides a voice style transfer method, apparatus, electronic device, and storage medium. First, the voice style transfer method provided by the present disclosure will be described below.

[0038] As Figure 1 shown, Figure 1 is a schematic diagram of a first embodiment of the voice style transfer method provided by the present disclosure. The method may include the following steps:

[0039] Step S110, obtain a target audio to be transferred and a phoneme sequence to be synthesized.

[0040] In the embodiments of the present disclosure, the target audio to be transferred may be a sentence of any content input by a speaker through a sound input device such as a microphone. The speaker may be the user himself / herself or a virtual character (such as an anime character, etc.).

[0041] In linguistics, a phoneme refers to the smallest speech unit divided according to the natural attributes of speech. One pronunciation action in a syllable constitutes a phoneme. For example, the word "Mandarin" can be divided into eight phonemes: "p, u, t, o, n, g, h, u, a". In the embodiments of the present disclosure, the phoneme sequence to be synthesized may be a text sequence of any content. The text sequence may be a text sequence input by other users that needs to generate a text sequence with the speaker's style, or it may be a phoneme sequence obtained by an electronic device through operations such as feature extraction on the corresponding audio input by other users through voice. The present disclosure does not make specific limitations in this regard.

[0042] Step S120, perform spectrogram feature extraction and phoneme duration feature extraction on the target audio to be transferred, and obtain the spectrogram feature and phoneme duration feature of the target audio.

[0043] In the embodiments of the present disclosure, the spectrogram feature of the target audio to be transferred may reflect the voice feature of the speaker, and the phoneme duration feature may reflect the speaking speed of the speaker.

[0044] In the embodiments of the present disclosure, the target audio signal can be framed and windowed, and FFT (Fast Fourier Transform) is performed on each frame. After stacking the frequency-domain signals (spectrograms) of each frame after FFT in time to obtain a spectrogram, the spectral features of the spectrogram are extracted.

[0045] In the embodiments of the present disclosure, the above phoneme duration features can be extracted for each phoneme inside the target audio, that is, the obtained can be the duration features of each internal phoneme.

[0046] Step S130: Extract content features and predict phoneme durations for the to-be-synthesized phoneme sequence to obtain the content features of the to-be-synthesized phoneme sequence and the predicted basic duration of each to-be-synthesized phoneme.

[0047] In this embodiment, a preset content feature extraction algorithm can be used to extract content features from the to-be-synthesized phoneme sequence. In the embodiments of the present disclosure, the above content features can represent the association information and semantic information between the content information represented by the phoneme and its context. As a specific implementation, the above content features can be expressed in vector form.

[0048] In the embodiments of the present disclosure, the predicted basic duration of each to-be-synthesized phoneme can be obtained by inputting the to-be-synthesized phoneme sequence into a pre-trained duration prediction model. The duration prediction model can be a speech recognition model such as LSTM-CTC, CNN-RNN-T, LAS, Chain, GMM-HM, etc. pre-trained with phoneme annotation and duration annotation of the audio in the open-source data Aishell3. The prediction result can be characterized by the number of unit durations. For example: a unit duration can be set to 10 ms. If the prediction result of a to-be-synthesized phoneme is 3 unit durations, its actual duration is 30 ms.

[0049] Step S140: Adjust the predicted basic duration of each to-be-synthesized phoneme based on the phoneme duration features of the target audio to obtain the target duration of each to-be-synthesized phoneme;

[0050] This step can make the durations of the to-be-synthesized phonemes close to the phoneme durations in the target audio, that is, make the speech rate of the synthesized audio close to the speech rate of the target audio.

[0051] Step S150: Based on the spectral features of the target audio, the content features of the to-be-synthesized phoneme sequence, and the target duration of each to-be-synthesized phoneme, obtain the target spectrogram with the style of the target audio corresponding to the to-be-synthesized phoneme sequence;

[0052] Step S160: Convert the target spectrogram into audio to obtain the synthesized audio with the target audio style corresponding to the to-be-synthesized phoneme sequence.

[0053] In the embodiments of the present disclosure, the spectrogram signals frame by frame can be transformed into time-domain signals segment by segment through IFFT (Inverse Fast Fourier Transform), and then they are spliced together to obtain the audio. The specific method is not limited in the present disclosure.

[0054] In the speech style transfer method provided by the present disclosure, spectral feature extraction and phoneme duration feature extraction are performed on the target audio to be transferred to obtain its spectral feature and phoneme duration feature. Content feature extraction and phoneme duration prediction are performed on the to-be-synthesized phoneme sequence to obtain the content feature and predicted basic duration of each phoneme. Then, based on the phoneme duration feature of the target audio, the predicted basic duration of each phoneme in the to-be-synthesized phoneme sequence is adjusted to obtain the target duration of each to-be-synthesized phoneme. Based on the spectral feature of the target audio, the content feature and target duration of each to-be-synthesized phoneme, the target spectrogram corresponding to the to-be-synthesized phoneme sequence with the target audio style is obtained, and it is converted into audio to obtain the synthesized audio with the target audio style corresponding to the to-be-synthesized phoneme sequence. Applying the embodiments of the present disclosure, by combining the speech rate that has a greater impact on the style of the target audio to be transferred for speech style transfer, the audio transfer effect is better, and the audio transfer accuracy is improved.

[0055] In an embodiment of the present disclosure, the phoneme duration feature of the above target audio may specifically be the mean and variance of the phoneme duration. Therefore, referring to Figure 2 , Figure 1 the step S120 shown in

[0056] can be specifically refined into the following steps:

[0057] As mentioned above, the phoneme is the smallest speech unit. In the embodiments of the present disclosure, the phoneme duration calculation can be performed for each phoneme in the target phoneme, and the mean and variance of the target phoneme duration obtained can better reflect the speech rate change of the speaker, making the speech style transfer effect better.

[0058] In this way, as Figure 2 shown, Figure 1 the step S140 in

[0059] Step S141: Adjust the predicted basic duration of each phoneme to be synthesized according to the mean and variance of the phoneme durations of the target audio, so as to obtain the target duration of each phoneme to be synthesized that conforms to the target audio speech rate.

[0060] As a specific implementation manner of the embodiments of the present disclosure, the target duration of each phoneme to be synthesized may be the predicted basic duration of each phoneme to be synthesized * variance + mean, and this variance and mean are the mean and variance of the phoneme durations obtained by calculating the phoneme durations for each phoneme in the target audio. For example: the predicted basic duration is 2 units, the variance is 0.5 unit, and the mean is 2 units, then the adjusted duration is 2 * 0.5 + 2 = 3 unit durations.

[0061] As Figure 3 shown, in an embodiment of the present disclosure, Figure 1 before step S150 in

[0062] Step S350: Extract the style features of the target audio based on the spectral features of the target audio.

[0063] Generally, a speaker's behavior consists of conscious and unconscious activities, and generally has strong variability. From the perspective of the sound spectrum, it can include three types of features. One is the inherent, stable, and coarse-grained characteristics, which are determined by the physiological structure characteristics of the speaker. Different speakers are different, and they are felt differently subjectively, manifested in the low-frequency region of the sound spectrum, such as the average fundamental frequency (pitch), the spectral envelope reflecting the vocal tract impulse response, the relative amplitude and position of the formant, etc. Another is the unstable short-term acoustic characteristics, such as the sharp and rapid jitter of the pronunciation speed, the fundamental frequency, the intensity, the refined structure of the spectrum, etc., which can reflect the psychological or mental changes of the speaker and can be used to express different emotions and intentions by changing these during a conversation. Secondly, stress, pause, etc. are not related to the speaker himself, but also affect the subjective hearing.

[0064] Therefore, in the embodiments of the present disclosure, the above-mentioned style features of the speaker can be extracted, so that the style features and content features in the spectral features of the target audio to be migrated by the speaker can be separated, reducing the influence of the content of the target audio on its style, and making the speech style migration more accurate.

[0065] See Figure 3 As Figure 3 shown, Figure 1 step S150 in

[0066] Step S151: Based on the target duration of each to-be-synthesized phoneme in the to-be-synthesized phoneme sequence, copy and combine the content features corresponding to each phoneme in the to-be-synthesized phoneme sequence to obtain the target content features of the to-be-synthesized phoneme sequence.

[0067] As described above, in the embodiments of the present disclosure, the content feature vectors of each to-be-synthesized phoneme are extracted. Therefore, after obtaining the target durations of each to-be-synthesized phoneme, for each to-be-synthesized phoneme, its content feature vector can be copied in combination with its target duration to obtain the target content features of the to-be-synthesized phoneme, and then the target content features of each to-be-synthesized phoneme are combined to obtain the target content features of the entire to-be-synthesized phoneme sequence.

[0068] As described above, the target durations of each to-be-synthesized phoneme can be represented in the form of unit durations. Therefore, as a specific implementation manner of the embodiments of the present disclosure, the content feature vectors can be copied according to the number of unit durations included in the target duration of the to-be-synthesized phoneme. For example: If the to-be-synthesized phoneme sequence includes 3 phonemes, namely phoneme A, phoneme B, and phoneme C, and their content feature vectors are denoted as A, B, and C respectively. It is calculated that the target duration of phoneme A is 3 unit durations, the target duration of phoneme B is 2 unit durations, and the target duration of phoneme C is 2 unit durations. Then, the content features of the to-be-synthesized phoneme sequence include 3 content feature vectors of phoneme A, 2 content feature vectors of phoneme B, and 2 content feature vectors of phoneme C. That is, for the to-be-synthesized phoneme sequence, the target content features obtained by copying and combining the content feature vectors of each to-be-synthesized phoneme are AAABBCC.

[0069] Step S152: Based on the target content features of the to-be-synthesized phoneme sequence and the style features of the target audio, decode the to-be-synthesized phoneme sequence to obtain the target spectrogram with the style of the target audio corresponding to the to-be-synthesized phoneme sequence.

[0070] In the embodiments of the present disclosure, the above steps of spectrogram feature extraction, style feature extraction, content feature extraction of the to-be-synthesized phoneme sequence, phoneme duration prediction, and obtaining the target spectrogram and target audio for the target audio can all be executed by models such as a pre-trained content encoding model, duration prediction model, style encoding model, and spectrogram decoding model.

[0071] As a specific implementation manner, in the embodiments of the present disclosure, the above step of spectrogram feature extraction for the target audio to be migrated may include:

[0072] Input the target audio to be migrated into a preset spectrogram feature extraction model to obtain the spectrogram features of the target audio.

[0073] In the embodiments of the present disclosure, the above acoustic spectral features may be MFCC (Mel-Frequency Cepstral Coefficients), PLP (Perceptual Linear Prediction), or Fbank (Filter bank) features, etc. Taking the MFCC feature as an example, its extraction process may include: pre-emphasizing, framing, and windowing the speech signal; for each short-time analysis window (i.e., each frame divided), obtaining the corresponding spectrum through FFT; passing the calculated spectrum through a Mel filter bank to obtain the Mel spectrum; performing cepstral analysis (taking the logarithm and inverse transformation) on the Mel spectrum to obtain the Mel spectrum cepstral coefficients MFCC.

[0074] The above speaker features can be extracted through a pre-trained speaker recognition model. As a specific implementation manner of the embodiments of the present disclosure, the above speaker recognition model may be composed of multiple layers of TDNN (time delay neural network). When training this model, it can be trained based on the open-source dataset Aishelll-3. Taking the acoustic spectral features of each frame of audio as the input of the model, the acoustic spectral features of each frame of audio may also be MFCC, PLP, or Fbank, etc. The dimension of the acoustic spectral features of each frame of audio may be 20, the duration of each frame may be 25 ms, and the frame shift may be 10 ms. What the model outputs may be the probabilities of the predicted speakers. In this embodiment, the loss function for training may be CE (cross entropy), that is, the loss function value may be obtained based on the probabilities of the predicted speakers output by the model and the true speaker corresponding to the audio, and the parameters of each layer of the network in the speaker recognition model may be updated based on this loss function value. A certain amount of context (for example, 2 frames before and after) may be used during calculation to improve the calculation accuracy. After training, it can be extracted from the penultimate layer of this speaker recognition model, and the extracted data is used as the speaker features. Of course, it can also be extracted from other layers, and the present disclosure does not make specific limitations.

[0075] In the embodiments of the present disclosure, the steps of extracting content features and predicting phoneme durations for the to-be-synthesized phoneme sequence may include:

[0076] Inputting the to-be-synthesized phoneme sequence into a preset content encoding model to obtain the content features of the to-be-synthesized phoneme sequence; and inputting the to-be-synthesized phoneme sequence into a preset duration prediction model to obtain the predicted basic duration of each to-be-synthesized phoneme.

[0077] Next, the above content encoding model and duration prediction model will be described separately.

[0078] In the embodiments of the present disclosure, the above content encoding model may be composed of a multi-layer {Conv1D + ReLU (activation function) layer + IN layer + Dropout layer} network. The above Conv1D is a type of convolutional neural network. The ReLU layer is used to introduce non-linearity and enhance the model's expressive power. The IN (Instance Normalization) layer can be used to normalize the data to ensure that the input data distribution of each layer of the network is the same, thereby accelerating the convergence speed of the neural network. The Dropout layer can avoid the problem of overfitting of the model.

[0079] The IN layer starts from the assumption that each dimension in space is independent, and performs normalization on each dimension separately. For each layer of the content encoding model, for the feature x input to this layer, the mean and variance are calculated from different dimensions, and then normalized. After normalization, the distribution mean is 0 and the variance is 1. The specific calculation formula is as follows:

[0080]

[0081] In this formula, γ and β represent the style of the target audio to be migrated, σ and μ are the mean and standard deviation respectively, x is the feature of the phoneme sequence to be synthesized input to this layer, and IN(x) is the output of the IN layer.

[0082]

[0083] In this formula, H and W respectively represent the two dimensions of the input feature, x is the corresponding input feature, and n and c can respectively represent the sample index and feature channel index of x.

[0084]

[0085] In this formula, ε is a very small constant (for example, 0.1) to avoid the situation where the variance is 0.

[0086] In the embodiments of the present disclosure, the content encoding model is equivalent to mapping phonemes to a high-dimensional space through a neural network, representing the content information represented by the phonemes themselves, their associated information with the context, and semantic information, so that the extracted phoneme features are more complete. As described above, the specific content feature vector (content embedding) of the phoneme sequence to be synthesized can be extracted by the content encoding model.

[0087] In the embodiments of the present disclosure, the content encoding model to be trained can be trained based on the open-source data Aishell3 (multi-speaker Mandarin data) to obtain the above content encoding model.

[0088] As a specific implementation manner, such as Figure 4As shown, the above content encoding model can be trained through the following steps:

[0089] Step S410: Input the first sample audio into a pre-trained speaker recognition model to obtain the sample speaker feature corresponding to the first sample audio.

[0090] The above first sample audio can be the audio data in the above open-source data Aishell3. By using the above speaker recognition model to process this first sample audio, the speaker feature corresponding to this sample audio can be obtained.

[0091] Step S420: Input the first phoneme sequence of the first sample audio into the content encoding model to be trained to obtain the first sample content feature of each phoneme in the first sample phoneme sequence.

[0092] Step S430: Based on the duration of each phoneme in the first sample phoneme sequence, copy and combine the first sample content features of each phoneme to obtain the first sample target content feature of each phoneme in the first sample phoneme sequence.

[0093] In this embodiment, models such as GMM-HMM, CNN, RNN-T, Chain, and LAS can be used to convert audio into text with time information (including: the time point and duration of each phoneme). Based on the above time information, the duration of internal phonemes, as well as the mean and variance of phoneme durations, can be calculated.

[0094] As a specific implementation manner of this embodiment, the above-mentioned spectral decoding model can copy and combine the first sample content features of each phoneme in combination with the duration of the phoneme to obtain the above-mentioned first sample target content feature of each phoneme.

[0095] Step S440: Input the speaker feature and each first sample target content feature into the spectral decoding model to be trained to obtain the first sample spectral feature;

[0096] Step S450: Based on the error between the first sample spectral feature and the true spectral feature of the first sample audio, update the parameters of the content encoding model to be trained until the content encoding model to be trained converges.

[0097] In the embodiment of the present disclosure, the obtained first sample spectral feature and the true spectral feature of the audio can be used to calculate the loss through MSE (mean square error), and then the entire network can be updated.

[0098] Through the trained content encoding model, the content features of the phoneme sequence to be synthesized are extracted. When extracting, the previous and subsequent phonemes and semantic information of the phoneme are associated, so that the extracted content features are more accurate.

[0099] In the embodiments of the present disclosure, the above-mentioned duration prediction model may be composed of 1 layer of self-attention (self-attention mechanism network layer), 2 layers of {Conv1D + ReLU}, and 1 layer of Linear (linear layer). The input may be each phoneme to be synthesized, and the output is the predicted basic duration of the corresponding phoneme. As described above, the duration prediction model may be a speech recognition model pre-trained through phoneme annotation and duration annotation of the audio in the open-source data Aishell3. The speech recognition model may be LSTM-CTC, CNN-RNN-T, LAS, Chain, GMM-HMM, etc.

[0100] The above-mentioned duration prediction model may also adjust the predicted basic duration based on the duration mean and variance of the target audio to be migrated, so as to obtain the target duration of each phoneme to be synthesized.

[0101] Generally, the speech rate is a crucial factor in the subjective experience of the speaking style. In the embodiments of the present disclosure, a separate model is used to adjust the duration (speech rate) of the phonemes to be synthesized, which improves the convenience of adjusting the duration of each phoneme to be synthesized.

[0102] In the embodiments of the present disclosure, the step of extracting the style feature of the target audio based on the spectrogram feature of the target audio may include:

[0103] Input the spectrogram feature of the target audio into a preset style encoding model to obtain the style feature of the target audio.

[0104] The step of decoding the phoneme sequence to be synthesized based on the target content feature of the phoneme sequence to be synthesized and the style feature of the target audio, and obtaining the target spectrogram with the style of the target audio corresponding to the phoneme sequence to be synthesized may include:

[0105] Input the target content feature of the phoneme sequence to be synthesized and the style feature of the target audio into a preset spectrogram decoding model to obtain the target spectrogram with the style of the target audio corresponding to the phoneme sequence to be synthesized.

[0106] See Figure 5 , Figure 5 is a schematic flowchart of a voice migration method provided by the present disclosure, which may specifically include the following steps:

[0107] ①. Extract the spectrogram feature of the target audio to be migrated and calculate the phoneme duration to obtain the target spectrogram feature of the target audio and the mean and variance of its phoneme duration.

[0108] ②. Pass the synthetic phoneme sequence through the duration prediction model to obtain the predicted basic duration of each synthetic phoneme, and adjust it in combination with the mean and variance of the phoneme duration of the target audio to obtain the target duration of each phoneme.

[0109] ③. Pass the synthetic phoneme sequence through the content encoding model to obtain the content feature vector of each synthetic phoneme, and copy and combine the corresponding content feature vectors in combination with the target duration of each phoneme to obtain the target content feature vector of the synthetic phoneme sequence.

[0110] ④. Input the spectrogram features of the target audio into the style encoding model to obtain the style features of the target audio (including the mean and variance of the target audio spectrogram, etc.).

[0111] ⑤. Input the target content feature vector of the synthetic phoneme sequence obtained in step ③ into the spectrogram decoding model, and combine the target audio style features obtained in step ④ to obtain the synthetic spectrogram.

[0112] ⑥. Perform audio conversion on the synthetic spectrogram to obtain the synthetic audio corresponding to the synthetic phoneme sequence with the style of the target audio to be migrated.

[0113] In the embodiments of the present disclosure, the above style encoding model and spectrogram decoding model can form a U-shaped network. Among them, the above style encoding model can be the first U-shaped network model, and the above spectrogram decoding model can be the second U-shaped network model.

[0114] Therefore, the step of inputting the spectrogram features of the target audio into the preset style encoding model to obtain the style features of the target audio may include:

[0115] Input the spectrogram features of the target audio into the first U-shaped network model for content feature extraction, and use the features output by the middle layer of the first U-shaped network model as the style features of the target audio.

[0116] By using an independent style encoding model to extract the style of the audio to be migrated, the influence of content on the audio style can be reduced, making the extracted style information of the audio to be migrated more accurate.

[0117] The step of inputting the target content features of the synthetic phoneme sequence and the style features of the target audio into the preset spectrogram decoding model to obtain the target spectrogram corresponding to the synthetic phoneme sequence with the style of the target audio may include:

[0118] Input the target content features of the synthetic phoneme sequence and the style features of the target audio into the second U-shaped network model to obtain the target spectrogram corresponding to the synthetic phoneme sequence output by the second U-shaped network model with the style of the target audio.

[0119] By using the above U-shaped network for voice style transfer, the speaking style and content in the audio can be decoupled, modeled separately, reducing mutual influence, and improving the extraction accuracy of the content features and style features of the audio to be transferred.

[0120] As described above, from the perspective of spectrogram, the audio features of a speaker can be divided into inherently stable coarse-grained features, unstable short-term acoustic features, and features such as stress and pauses. Therefore, in the embodiments of the present disclosure, a multi-layer convolutional network can be used to extract the style features and content features of the target audio, obtaining the style features and content features of the target audio at different levels. The style features can include the mean, variance, and outputs of each intermediate layer, etc.

[0121] When training the style encoding model, the input can be the true spectrogram features of the sample audio, and the output is the predicted content feature vector. Then, based on the predicted content feature vector and the content feature vector of the audio output by the above content encoding model, the error loss is calculated through MSE (Mean Square Error), and the style encoding model is updated backward.

[0122] As a specific implementation, in the embodiments of the present disclosure, the above style encoding model can be composed of multiple {ResCon1D+IN} layers, and each ResCon1D is composed of 2 layers of {Conv1D+ReLU}.

[0123] In the embodiments of the present disclosure, the above spectrogram decoding model can be composed of multiple {AdaIN+ResCon1D} layers, and each ResCon1D is composed of 2 layers of {Conv1D+ReLU}. The AdaIN (Adaptive Instance Normalization) layer is similar to the IN, except that the coefficients on both sides have changed, as follows:

[0124]

[0125] In this formula, x and y are the content feature and style feature respectively, σ and μ are the mean and variance respectively, and the mean and variance are the mean and variance output by the above style encoding model. This formula aligns the mean and variance of the content feature with the style feature. The spectrogram decoding model is equivalent to the process of applying the style of the target audio to the given content to obtain the target spectrogram.

[0126] When training the spectrogram coding model, the input can be the content feature vector of the phoneme, combined with part of the output (style feature) of the style coding model. The output is the predicted spectrogram feature. Then, based on the predicted spectrogram feature and the actual spectrogram feature of the audio, the error loss is calculated by MSE (Mean Square Error) to reversely update the spectrogram decoding model.

[0127] In this way, combining partial information from the style encoding model to form a complete U-shaped network can more accurately predict detailed acoustic information (e.g., pitch, harmony, spectral envelope, loudness, etc.) and sentence-level speaking styles with strong randomness (e.g., pauses, stress, etc.).

[0128] As described above, the style encoding model and the spectrogram decoding model together constitute a U-shaped network. In the embodiment of the present disclosure, the style encoding model and the spectrogram decoding model can be trained simultaneously. Figure 6 As shown, the above U-shaped network can be trained by the following steps:

[0129] Step S610, calculating the mean and variance of the phoneme duration and the true spectral features of the second sample audio.

[0130] As described above, in this embodiment, GMM-HMM, CNN, RNN-T, Chain, LAS and other models can be used to convert the second sample audio into text with time information (including: the time point and duration of each phoneme), and the duration mean and variance of the internal phonemes can be calculated based on the time information.

[0131] Step S620, using the trained content coding model and the trained duration prediction model, extracting content features and predicting phoneme duration for the second sample phoneme sequence of the second sample audio, to obtain the second sample content features and the sample predicted basic duration of each sample phoneme;

[0132] Step S630: Based on the mean and variance of the phoneme duration of the second sample audio, the predicted basic duration of each phoneme in the second sample phoneme sequence is adjusted to obtain the sample target duration of each phoneme in the second sample phoneme sequence.

[0133] The above steps S620 and S630 have been described in detail above and will not be repeated here.

[0134] Step S640: Based on each sample target duration, copy and combine the second sample content features corresponding to each phoneme in the second sample phoneme sequence to obtain the second sample target content features of the second sample phoneme sequence, which are used as the true content features of the second sample audio.

[0135] In the embodiments of the present disclosure, for each phoneme in the second sample sequence, the content features of the phoneme can be copied and combined according to the number of unit durations included in the target duration to obtain the second sample target content features.

[0136] Step S650: Input the true spectrogram features of the second sample audio into the style encoding model to be trained to obtain the sample style features and sample audio content features of the second sample audio.

[0137] As described above, the style encoding model decouples style and content and models them separately. Therefore, the style encoding model to be trained can output the content features and style features of the second audio.

[0138] Step S660: Based on the error between the sample audio content features and the true content features, update the parameters of the style encoding model to be trained until the style encoding model to be trained converges to obtain the to-be-determined style encoding model.

[0139] As described above, the error loss between the sample audio content features and the true content features can be calculated by MSE (Mean Square Error).

[0140] In the embodiments of the present disclosure, the style encoding model and the spectrogram decoding model together form a U-shaped network. Therefore, when training the spectrogram decoding model subsequently, the parameters of the entire U-shaped network will be updated. So, the style encoding model obtained here is not the trained one.

[0141] Step S670: Input the true content features and the sample style features output by the to-be-determined style encoding model into the spectrogram decoding model to be trained to obtain the sample spectrogram features.

[0142] As described above, the sample style features output by the to-be-determined style encoding model may include the mean and variance of the sample audio, as well as the outputs of each layer in the multi-layer convolutional network.

[0143] Step S680: Based on the error between the sample spectrogram features and the true spectrogram features, update the parameters of the spectrogram decoding model to be trained and the to-be-determined style encoding model until the spectrogram decoding model to be trained and the to-be-determined style encoding model converge.

[0144] In the embodiments of the present disclosure, by combining the output style features of the style encoding model, the parameters of the U-shaped network composed of the spectrogram decoding model and the style encoding model are updated, so that the U-shaped network can more accurately predict detailed acoustic information and improve the accuracy of speech style transfer.

[0145] Next, the process of training and testing the entire speech style transfer network in the present disclosure will be introduced.

[0146] See Figure 7 , Figure 7 which shows the process of pre-training, training and testing the speech style transfer network in the embodiments of the present disclosure:

[0147] In the pre-training stage, using the open-source data Aishell1-3, a speaker recognition model is trained on a multi-layer TDNN-xvector network, and this speaker recognition model can extract speaker features (Speaker Embedding).

[0148] Then, a multi-speaker synthesis system combining a content encoding (Content Encoder) model and a spectrogram decoding (Mel Decoder) model is trained through the speaker features of the sample audio and the phoneme sequence of the sample audio to obtain a content encoding model.

[0149] Specifically, the sample phoneme sequence of the sample audio can be input into the content encoding model to be trained, and the content features of each sample phoneme output by the content encoding model to be trained can be obtained ( Figure 7 in which the sample phoneme sequence contains 3 sample phonemes). Then, based on the duration of each sample phoneme, the content features of each phoneme are copied and combined to obtain the content features of the sample phoneme sequence. The content features and the speaker features are input into the spectrogram decoding model to be trained, and the target spectrogram of the sample phoneme sequence output by the spectrogram decoding model to be trained is obtained. Based on the target spectrogram and the spectrogram features of the sample audio, a loss function is calculated to update the parameters of the content encoding model until the model converges, and a trained content encoding model is obtained.

[0150] In the pre-training stage, a phoneme duration prediction model can also be trained, that is, it is trained through the phoneme annotation and duration annotation of the audio in Aishell3.

[0151] In the training stage, each sample audio in the training data has a phoneme annotation. Combining with the specific sample audio, the mean (Mean) and variance (Std) of the duration of each sample phoneme of the sample audio can be calculated. And through the sample audio, its spectrogram features can be calculated. The sample phoneme sequence of the sample audio is input into the trained content encoding model, and the content feature vectors of each sample phoneme can be obtained.

[0152] Then, input the above sample phoneme sequence into the trained duration prediction model to obtain the predicted basic durations of each sample phoneme, and adjust them by combining the mean and variance of the durations of each sample phoneme calculated, so as to obtain the final target duration information. Based on this target duration, copy and combine the content feature vectors output by the above trained content encoding model, and the real content feature vector of the sample audio can be obtained.

[0153] Input the real spectrogram features of the sample audio into the style encoding model (composed of multiple {ResCNN1D layers + IN layers}), calculate the style features of each intermediate layer and the content feature vector of the output layer, calculate the error (loss2) with the real content feature vector of the above sample audio, and update the network backward.

[0154] Input the above real content feature vector into the spectrogram decoding model (composed of multiple {ResCNN1D layers + AdaIN layers}), and combine the style features (including the mean and variance of the real spectrogram) output by each intermediate layer of the style encoding model to generate the target spectrogram with the style of the sample audio, calculate the error (loss1) with the real spectrogram, and update the entire U-shaped network backward. Iterate multiple rounds until convergence to complete the training.

[0155] In the test stage, for the target audio to be migrated, extract the spectrogram features and calculate the mean and variance of the durations of the internal phonemes.

[0156] For the phoneme sequence to be synthesized, obtain the predicted basic duration information through the duration prediction model, and then adjust it according to the mean and variance of the duration of the target audio to obtain the target duration under the guidance of the target audio speech rate. In addition, the phoneme sequence to be synthesized passes through the content encoding model to obtain the content feature vector, and is copied and combined in combination with the target duration.

[0157] The spectrogram features extracted from the target audio calculate the style feature information (including the mean, variance, etc. of the spectrogram) through the style encoding model. Input the above obtained content feature vector into the spectrogram decoding model, and combine the above style feature information to jointly synthesize the target spectrogram with the style of the target audio, and finally convert it into audio, so as to obtain the synthesized audio with the style of the target audio.

[0158] It can be seen that compared with the lack of speaker dynamic or random fine-grained features caused by the prior art of voice style transfer through content feature extraction (phoneme), speaker characteristic extraction (speaker), audio spectral feature prediction (mel-spectrogram), and finally converting the spectrogram into audio through an existing vocoder, the voice style transfer method provided by the embodiment of the present disclosure designs a duration prediction model combined with regular means to achieve matching of synthesized audio speech rate with target audio, laying the foundation for style transfer, and also designs a style encoding model, which, through joint training with a spectral decoding model, well decouples content information from the speaker's speaking style and reduces mutual influence, and finally, through a U-shaped network reconstructed by the spectral, predicts acoustic detailed information with maximum accuracy, such as pitch, harmony, spectral envelope, loudness, etc., as well as a sentence-level speaking style with strong randomness, such as pauses and stress, to achieve voice style transfer based on a sentence.

[0159] According to an embodiment of the present disclosure, the present disclosure also provides a speech style transfer device, such as Figure 8 As shown, the device may include:

[0160] The audio and phoneme sequence acquisition module 810 is used to acquire the target audio to be migrated and the phoneme sequence to be synthesized;

[0161] A target audio feature acquisition module 820 is used to extract the spectral features and the phoneme duration features of the target audio to be migrated, so as to obtain the spectral features and the phoneme duration features of the target audio;

[0162] The feature extraction module 830 of the phoneme sequence to be synthesized is used to extract content features and predict phoneme duration of the phoneme sequence to be synthesized, so as to obtain content features of the phoneme sequence to be synthesized and predicted basic duration of each phoneme to be synthesized;

[0163] A phoneme duration adjustment module 840, configured to adjust the predicted basic duration of each to-be-synthesized phoneme based on the phoneme duration feature of the target audio, to obtain a target duration of each to-be-synthesized phoneme;

[0164] The target spectrum acquisition module 850 is used to acquire a target spectrum having a target audio style corresponding to the phoneme sequence to be synthesized based on the spectrum features of the target audio, the content features of the phoneme sequence to be synthesized, and the target duration of each phoneme to be synthesized;

[0165] The synthesized audio acquisition module 860 is used to convert the target sound spectrum into audio to obtain a synthesized audio with a target audio style corresponding to the to-be-synthesized phoneme sequence.

[0166] In the voice style transfer device provided by the present disclosure, spectral features and phoneme duration features of the target audio to be transferred are extracted to obtain its spectral features and phoneme duration features. Content features of each phoneme are extracted from the phoneme sequence to be synthesized, and the predicted basic duration of each phoneme is predicted to obtain the content features and predicted basic durations of each phoneme. Then, based on the phoneme duration features of the target audio, the predicted basic durations of the phonemes in the phoneme sequence to be synthesized are adjusted to obtain the target durations of each phoneme to be synthesized. Based on the spectral features of the target audio, the content features, and the target durations of each phoneme to be synthesized, the target spectrum corresponding to the phoneme sequence to be synthesized with the style of the target audio is obtained, and audio conversion is performed on it to obtain the synthesized audio corresponding to the phoneme sequence to be synthesized with the style of the target audio. Applying the embodiments of the present disclosure, by combining the speech rate that has a greater impact on the style of the target audio to be transferred for voice style transfer, the audio transfer effect is better, and the audio transfer accuracy is improved.

[0167] In one embodiment of the present disclosure, the target audio feature acquisition module 820 extracts phoneme duration features from the target audio to be transferred, including: calculating the duration of each phoneme included in the target audio to be transferred to obtain the mean and variance of the phoneme durations of the target audio;

[0168] The phoneme duration adjustment module 840 is configured to adjust the predicted basic duration of each phoneme to be synthesized according to the mean and variance of the phoneme durations of the target audio to obtain the target duration of each phoneme to be synthesized that conforms to the speech rate of the target audio.

[0169] In one embodiment of the present disclosure, the above device may further include a style feature extraction module (not shown in the figure) for extracting the style features of the target audio based on the spectral features of the target audio;

[0170] The target spectrum acquisition module 850 is configured to copy and combine the content features corresponding to each phoneme in the phoneme sequence to be synthesized based on the target duration of each phoneme in the phoneme sequence to be synthesized to obtain the target content features of the phoneme sequence to be synthesized;

[0171] Based on the target content features of the phoneme sequence to be synthesized and the style features of the target audio, the phoneme sequence to be synthesized is decoded to obtain the target spectrum corresponding to the phoneme sequence to be synthesized with the style of the target audio.

[0172] In one embodiment of the present disclosure, the target audio feature acquisition module 820 may be configured to input the target audio to be transferred into a preset spectral feature extraction model to obtain the spectral features of the target audio;

[0173] The to-be-synthesized phoneme sequence feature extraction module 830 is configured to input the to-be-synthesized phoneme sequence into a preset content encoding model to obtain the content features of the to-be-synthesized phoneme sequence; and input the to-be-synthesized phoneme sequence into a preset duration prediction model to obtain the predicted basic duration of each to-be-synthesized phoneme.

[0174] The style feature extraction module is configured to input the spectrogram features of the target audio into a preset style encoding model to obtain the style features of the target audio.

[0175] The target spectrogram acquisition module 850 is configured to input the target content features of the to-be-synthesized phoneme sequence and the style features of the target audio into a preset spectrogram decoding model to obtain the target spectrogram with the style of the target audio corresponding to the to-be-synthesized phoneme sequence.

[0176] In an embodiment of the present disclosure, the preset style encoding model is a first U-shaped network model.

[0177] The style feature extraction module is configured to input the spectrogram features of the target audio into the first U-shaped network model for content feature extraction, and use the features output by the middle layer of the first U-shaped network model as the style features of the target audio.

[0178] In other embodiments of the present disclosure, the preset spectrogram decoding model is a second U-shaped network model.

[0179] The target spectrogram acquisition module 850 is configured to input the target content features of the to-be-synthesized phoneme sequence and the style features of the target audio into the second U-shaped network model to obtain the target spectrogram with the style of the target audio corresponding to the to-be-synthesized phoneme sequence output by the second U-shaped network model.

[0180] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved are all in compliance with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0181] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0182] Figure 9FIG. 0 is a schematic block diagram of an exemplary electronic device 900 that may be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, for example, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0183] As Figure 9 shown, the device 900 includes a computing unit 901 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0184] A plurality of components in the device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0185] The computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 executes the various methods and processes described above, such as the voice style transfer method. For example, in some embodiments, the voice style transfer method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the voice style transfer method described above can be executed. Alternatively, in other embodiments, the computing unit 901 can be configured to execute the voice style transfer method in any other suitable manner (e.g., by means of firmware).

[0186] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0187] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program code is executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0188] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0189] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0190] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0191] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The relationship of the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, can also be a server of a distributed system, or a server incorporating a blockchain.

[0192] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.

[0193] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A method for voice style transfer, comprising: Obtain a target audio to be migrated and a phoneme sequence to be synthesized; wherein, the target audio is an audio of any content of a speaker; Extract the spectrogram feature and phoneme duration feature of the target audio to be migrated, and obtain the spectrogram feature and phoneme duration feature of the target audio; Extract the content feature and predict the phoneme duration of the phoneme sequence to be synthesized, and obtain the content feature of the phoneme sequence to be synthesized and the predicted basic duration of each phoneme to be synthesized; Based on the phoneme duration feature of the target audio, adjust the predicted basic duration of each phoneme to be synthesized to obtain the target duration of each phoneme to be synthesized; Based on the spectrogram feature of the target audio, the content feature of the phoneme sequence to be synthesized, and the target duration of each phoneme to be synthesized, obtain a target spectrogram corresponding to the phoneme sequence to be synthesized and having the style of the target audio; Convert the target spectrogram into an audio to obtain a synthesized audio corresponding to the phoneme sequence to be synthesized and having the style of the target audio.

2. The method according to claim 1, wherein, The step of extracting the phoneme duration feature of the target audio to be migrated includes: Calculate the duration of each phoneme included in the target audio to be migrated to obtain the mean and variance of the phoneme duration of the target audio; The step of adjusting the predicted basic duration of each phoneme to be synthesized based on the phoneme duration feature of the target audio to obtain the target duration of each phoneme to be synthesized includes: Adjust the predicted basic duration of each phoneme to be synthesized according to the mean and variance of the phoneme duration of the target audio to obtain the target duration of each phoneme to be synthesized that conforms to the speech rate of the target audio.

3. Before the step of obtaining the target spectrogram with the target audio style corresponding to the to-be-synthesized phoneme sequence based on the spectrogram features of the target audio, the content features of the to-be-synthesized phoneme sequence, and the target duration of each to-be-synthesized phoneme, the method according to claim 1 further includes: Extract the style feature of the target audio based on the spectrogram feature of the target audio; The step of obtaining a target spectrogram corresponding to the phoneme sequence to be synthesized and having the style of the target audio based on the spectrogram feature of the target audio, the content feature of the phoneme sequence to be synthesized, and the target duration of each phoneme to be synthesized includes: Based on the target duration of each phoneme to be synthesized in the phoneme sequence to be synthesized, copy and combine the content features corresponding to each phoneme in the phoneme sequence to be synthesized to obtain the target content feature of the phoneme sequence to be synthesized; Based on the target content feature of the phoneme sequence to be synthesized and the style feature of the target audio, decode the phoneme sequence to be synthesized to obtain a target spectrogram corresponding to the phoneme sequence to be synthesized and having the style of the target audio.

4. The method according to claim 3, wherein, The step of extracting the spectrogram feature of the target audio to be migrated includes: inputting the target audio to be migrated into a preset spectrogram feature extraction model to obtain the spectrogram feature of the target audio; and / or The step of extracting the content feature and predicting the phoneme duration of the phoneme sequence to be synthesized includes: inputting the phoneme sequence to be synthesized into a preset content encoding model to obtain the content feature of the phoneme sequence to be synthesized; and inputting the phoneme sequence to be synthesized into a preset duration prediction model to obtain the predicted basic duration of each phoneme to be synthesized; and / or The step of extracting the style feature of the target audio based on the spectrogram feature of the target audio includes: inputting the spectrogram feature of the target audio into a preset style encoding model to obtain the style feature of the target audio; and / or The step of decoding the to-be-synthesized phoneme sequence based on the target content feature of the to-be-synthesized phoneme sequence and the style feature of the target audio to obtain the target spectrogram with the style of the target audio corresponding to the to-be-synthesized phoneme sequence includes: inputting the target content feature of the to-be-synthesized phoneme sequence and the style feature of the target audio into a preset spectrogram decoding model to obtain the target spectrogram with the style of the target audio corresponding to the to-be-synthesized phoneme sequence output by the second U-shaped network model.

5. The method according to claim 4, wherein, The preset style encoding model is a first U-shaped network model; The step of inputting the spectrogram feature of the target audio into a preset style encoding model to obtain the style feature of the target audio includes: Inputting the spectrogram feature of the target audio into the first U-shaped network model to perform content feature extraction, and taking the feature output by the middle layer of the first U-shaped network model as the style feature of the target audio.

6. The method according to claim 4, wherein, The preset spectrogram decoding model is a second U-shaped network model; The step of inputting the target content feature of the to-be-synthesized phoneme sequence and the style feature of the target audio into a preset spectrogram decoding model to obtain the target spectrogram with the style of the target audio corresponding to the to-be-synthesized phoneme sequence includes: Inputting the target content feature of the to-be-synthesized phoneme sequence and the style feature of the target audio into the second U-shaped network model to obtain the target spectrogram with the style of the target audio corresponding to the to-be-synthesized phoneme sequence output by the second U-shaped network model.

7. The method according to claim 4, wherein, The content encoding model is trained by the following steps: Inputting a first sample audio into a pre-trained speaker recognition model to obtain the sample speaker feature corresponding to the sample audio; Inputting the first phoneme sequence of the first sample audio into the content encoding model to be trained to obtain the first sample content feature of each phoneme in the first sample phoneme sequence; Based on the duration of each phoneme in the first sample phoneme sequence, copying and combining the first sample content feature of each phoneme to obtain the first sample target content feature of each phoneme in the first sample phoneme sequence; Inputting the speaker feature and each first sample target content feature into the spectrogram decoding model to be trained to obtain the first sample spectrogram feature; Updating the parameters of the content encoding model to be trained based on the error between the first sample spectrogram feature and the true spectrogram feature of the first sample audio until the content encoding model to be trained converges.

8. The method according to claim 4, wherein, The style encoding model and the spectrogram decoding model are trained by the following steps: Calculating the mean and variance of the phoneme duration and the true spectrogram feature of the second sample audio; Using the trained content encoding model and the trained duration prediction model to perform content feature extraction and phoneme duration prediction on the second sample phoneme sequence of the second sample audio to obtain the second sample content feature and the sample predicted basic duration of each sample phoneme; Based on the mean and variance of the phoneme durations of the second sample audio, adjust the predicted basic duration of each phoneme in the second sample phoneme sequence to obtain the sample target durations of the phonemes in the second sample phoneme sequence; Based on the respective sample target durations, copy and combine the second sample content features corresponding to each phoneme in the second sample phoneme sequence to obtain the second sample target content features of the second sample phoneme sequence, serving as the true content features of the second sample audio; Input the true spectrogram features of the second sample audio into the style encoding model to be trained to obtain the sample style features and sample audio content features of the second sample audio; Based on the error between the sample audio content features and the true content features, update the parameters of the style encoding model to be trained until the style encoding model to be trained converges, obtaining the to-be-determined style encoding model; Input the true content features and the sample style features output by the to-be-determined style encoding model into the spectrogram decoding model to be trained to obtain the sample spectrogram features; Based on the error between the sample spectrogram features and the true spectrogram features, update the parameters of the spectrogram decoding model to be trained and the to-be-determined style encoding model until the spectrogram decoding model to be trained and the to-be-determined style encoding model converge.

9. A voice style transfer device, comprising: An audio and phoneme sequence acquisition module, configured to acquire a target audio to be migrated and a phoneme sequence to be synthesized; wherein, the target audio is an audio of any content of a speaker; A target audio feature acquisition module, configured to perform spectrogram feature extraction and phoneme duration feature extraction on the target audio to be migrated to obtain the spectrogram features and phoneme duration features of the target audio; A phoneme sequence to be synthesized feature extraction module, configured to perform content feature extraction and phoneme duration prediction on the phoneme sequence to be synthesized to obtain the content features of the phoneme sequence to be synthesized and the predicted basic duration of each phoneme to be synthesized; A phoneme duration adjustment module, configured to adjust the predicted basic duration of each phoneme to be synthesized based on the phoneme duration features of the target audio to obtain the target duration of each phoneme to be synthesized; A target spectrogram acquisition module, configured to obtain a target spectrogram with the style of the target audio corresponding to the phoneme sequence to be synthesized based on the spectrogram features of the target audio, the content features of the phoneme sequence to be synthesized, and the target duration of each phoneme to be synthesized; A synthesized audio acquisition module, configured to convert the target spectrogram into an audio to obtain a synthesized audio with the style of the target audio corresponding to the phoneme sequence to be synthesized.

10. The device according to claim 9, wherein, The target audio feature acquisition module performs phoneme duration feature extraction on the target audio to be migrated, including: calculating the durations of the respective phonemes included in the target audio to be migrated to obtain the mean and variance of the phoneme durations of the target audio; The phoneme duration adjustment module is configured to adjust the predicted basic duration of each phoneme to be synthesized according to the mean and variance of the phoneme durations of the target audio to obtain the target duration of each phoneme to be synthesized that conforms to the speech rate of the target audio.

11. The device according to claim 9, further comprising: A style feature extraction module, configured to extract the style features of the target audio based on the spectrogram features of the target audio; The target spectrogram acquisition module is configured to copy and combine the content features corresponding to each phoneme in the to-be-synthesized phoneme sequence based on the target duration of each to-be-synthesized phoneme in the to-be-synthesized phoneme sequence, so as to obtain the target content features of the to-be-synthesized phoneme sequence; Based on the target content features of the to-be-synthesized phoneme sequence and the style features of the target audio, decode the to-be-synthesized phoneme sequence to obtain a target spectrogram corresponding to the to-be-synthesized phoneme sequence and having the style of the target audio.

12. The device according to claim 11, wherein, The target audio feature acquisition module is configured to input the target audio to be migrated into a preset spectrogram feature extraction model to obtain the spectrogram features of the target audio; The to-be-synthesized phoneme sequence feature extraction module is configured to input the to-be-synthesized phoneme sequence into a preset content encoding model to obtain the content features of the to-be-synthesized phoneme sequence; and input the to-be-synthesized phoneme sequence into a preset duration prediction model to obtain the predicted basic duration of each to-be-synthesized phoneme; The style feature extraction module is configured to input the spectrogram features of the target audio into a preset style encoding model to obtain the style features of the target audio; The target spectrogram acquisition module is configured to input the target content features of the to-be-synthesized phoneme sequence and the style features of the target audio into a preset spectrogram decoding model to obtain a target spectrogram corresponding to the to-be-synthesized phoneme sequence and having the style of the target audio.

13. The device according to claim 12, wherein, The preset style encoding model is a first U-shaped network model; The style feature extraction module is configured to input the spectrogram features of the target audio into the first U-shaped network model for content feature extraction, and use the features output by the middle layer of the first U-shaped network model as the style features of the target audio.

14. The device according to claim 12, wherein, The preset spectrogram decoding model is a second U-shaped network model; The target spectrogram acquisition module is configured to input the target content features of the to-be-synthesized phoneme sequence and the style features of the target audio into the second U-shaped network model to obtain a target spectrogram corresponding to the to-be-synthesized phoneme sequence and having the style of the target audio output by the second U-shaped network model.

15. An electronic device, comprising: At least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-8.

16. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-8.

17. A computer program product, comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Speech synthesis method and system for new tone generation

    CN112802448A