Speech synthesis method and device, electronic equipment and readable storage medium
By using a decoupled speech synthesis model, efficient speech synthesis from text to specific timbres is achieved, solving the need for personalized speech synthesis in scenarios such as video creation, and possessing timbre control capabilities and cross-style speech synthesis capabilities.
Patent Information
- Application Number
- CN202111107876.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-22
- Publication Date
- 2026-01-13
- Estimated Expiration
- 2041-09-22
AI Technical Summary
Existing technologies struggle to achieve efficient text-to-speech synthesis with specific timbres, especially in scenarios such as video creation, where users' demand for personalized speech synthesis remains unmet.
A decoupled speech synthesis model is adopted, including a first feature extraction sub-model and a second feature extraction sub-model, which are used to establish the mapping from text to bottleneck features and from bottleneck features to Mel-spectral features, respectively. The timbre and other features are independently controlled through a pre-trained model.
It achieves stable synthesis of timbre, reduces reliance on high-quality target audio, meets the needs of personalized speech synthesis, and can synthesize speech across styles.
Smart Images

Figure CN115910021B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to a speech synthesis method, apparatus, electronic device, and readable storage medium. Background Technology
[0002] TTS (text-to-speech) is a technology that converts text into audio. It is one of the more popular speech synthesis technologies today and has been widely used in various industries. For example, video creation, intelligent customer service, intelligent reading aloud, and intelligent dubbing.
[0003] With the continuous development of artificial intelligence technology, speech synthesis technology has penetrated into various production and daily life scenarios, and users have put forward higher requirements for speech synthesis, such as personalized speech synthesis. Taking video creation as an example, when creating videos, users want to convert a piece of text into audio with a specific timbre and add that audio as background music to enhance the personalization of the video. However, how to achieve text-to-audio conversion with a specific timbre is a problem that urgently needs to be solved. Summary of the Invention
[0004] In order to solve the above-mentioned technical problems, or at least partially solve the above-mentioned technical problems, this disclosure provides a speech synthesis method, apparatus, electronic device and readable storage medium.
[0005] In a first aspect, this disclosure provides a speech synthesis method, including:
[0006] Get the text to be processed;
[0007] The text to be processed is input into a target speech synthesis model, and the Mel spectrum sequence corresponding to the text to be processed is obtained from the output of the target speech synthesis model; wherein, the target speech synthesis model includes: a first feature extraction sub-model and a second feature extraction sub-model, the first feature extraction sub-model is used to output a first acoustic feature based on the input text to be processed, the first acoustic feature including a bottleneck feature corresponding to the text to be processed; the second feature extraction sub-model is used to output the Mel spectrum feature corresponding to the text to be processed based on the input first acoustic feature;
[0008] Based on the Mel spectrum features corresponding to the text to be processed, the target audio corresponding to the text to be processed is obtained, and the target audio has a target timbre.
[0009] In some possible implementations, the first feature extraction sub-model is trained based on the labeled text corresponding to the first sample audio and the second acoustic features corresponding to the first sample audio, wherein the second acoustic features include the first labeled bottleneck features corresponding to the first sample audio.
[0010] In some possible implementations, the second feature extraction sub-model is obtained by training based on the third acoustic feature and the first labeled Mel spectral feature corresponding to the second sample audio, and the fourth acoustic feature and the second labeled Mel spectral feature corresponding to the third sample audio.
[0011] The third acoustic feature includes the second labeled bottleneck feature corresponding to the second sample audio; the fourth acoustic feature includes the third labeled bottleneck feature corresponding to the third sample audio; and the third sample audio is a sample audio with the target timbre.
[0012] In some possible implementations, the first annotation bottleneck feature corresponding to the first sample audio, the second annotation bottleneck feature corresponding to the second sample audio, and the third annotation bottleneck feature corresponding to the third sample audio are obtained by the encoder of the end-to-end speech recognition model extracting bottleneck features from the input first sample audio, second sample audio, and third sample audio, respectively.
[0013] In some possible implementations, the second acoustic feature further includes: a first labeled fundamental frequency feature corresponding to the first sample audio; the third acoustic feature further includes: a second labeled fundamental frequency feature corresponding to the second sample audio; the fourth acoustic feature further includes: a third labeled fundamental frequency feature corresponding to the third sample audio.
[0014] Accordingly, the first acoustic feature output by the first feature extraction sub-model also includes: the fundamental frequency feature corresponding to the text to be processed.
[0015] In some possible implementations, the first labeled fundamental frequency feature corresponding to the first sample audio, the second labeled fundamental frequency feature corresponding to the second sample audio, and the third labeled fundamental frequency feature corresponding to the third sample audio are obtained by performing digital signal processing on the first sample audio, the second sample audio, and the third sample audio respectively.
[0016] In some possible implementations, the language of the first sample audio is the same as that of the second sample audio; or the language of the first sample audio is different from that of the third sample audio.
[0017] Secondly, this disclosure provides a speech synthesis apparatus, comprising:
[0018] The acquisition module is used to acquire the text to be processed;
[0019] A processing module is configured to input the text to be processed into a target speech synthesis model and obtain the Mel spectrum sequence corresponding to the text to be processed output by the target speech synthesis model; wherein, the target speech synthesis model includes: a first feature extraction sub-model and a second feature extraction sub-model, the first feature extraction sub-model being configured to output a first acoustic feature based on the input text to be processed, the first acoustic feature including a bottleneck feature corresponding to the text to be processed; the second feature extraction sub-model being configured to output the Mel spectrum feature corresponding to the text to be processed based on the input first acoustic feature;
[0020] The processing module is further configured to obtain the target audio corresponding to the text to be processed based on the Mel spectrum features corresponding to the text to be processed, wherein the target audio has a target timbre.
[0021] Thirdly, this disclosure provides an electronic device, including: a memory, a processor, and a computer program;
[0022] The memory is configured to store the computer program;
[0023] The processor is configured to execute the computer program to implement the speech synthesis method as described in any of the first aspects.
[0024] Fourthly, this disclosure provides a readable storage medium, including: computer program instructions;
[0025] When the computer program instructions are executed by at least one processor of an electronic device, they implement the speech synthesis method as described in any of the first aspects.
[0026] Fifthly, this disclosure provides a program product comprising: computer program instructions; the computer program instructions are stored in a readable storage medium, an electronic device obtains the computer program instructions from the readable storage medium, and at least one processor of the electronic device executes the computer program instructions to implement the speech synthesis method as described in any of the first aspects.
[0027] This disclosure provides a speech synthesis method, apparatus, electronic device, and readable storage medium. The proposed solution achieves audio conversion from text to a target timbre through a pre-trained speech synthesis model. The speech synthesis model includes a first feature extraction sub-model and a second feature extraction sub-model. The first feature extraction sub-model outputs a first acoustic feature, including bottleneck features, based on the input text to be processed. The second feature extraction sub-model outputs Mel-spectral features corresponding to the text to be processed based on the input first acoustic feature. The target audio, possessing a target timbre, is obtained based on the Mel-spectral features corresponding to the text to be processed. This solution decouples the speech synthesis model into two models through the first acoustic feature containing bottleneck features, achieving relatively independent control of timbre and other features over speech synthesis, thus meeting users' needs for personalized speech synthesis. Attached Figure Description
[0028] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0029] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0030] Figures 1a to 1c This is a schematic diagram of the structure of the speech synthesis model provided in this disclosure;
[0031] Figure 2 A flowchart of a speech synthesis method provided in an embodiment of this disclosure;
[0032] Figure 3 This is a schematic diagram of the structure of a speech synthesis device provided in an embodiment of the present disclosure;
[0033] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure. Detailed Implementation
[0034] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.
[0035] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.
[0036] This disclosure provides a speech synthesis method, apparatus, electronic device, readable storage medium, and program product. The method achieves audio conversion from text to target timbre through a pre-trained speech synthesis model. The speech synthesis model can achieve relatively independent control of timbre and other features on speech synthesis, meeting users' needs for personalized speech synthesis.
[0037] The speech synthesis method provided in this disclosure can be executed by an electronic device. This electronic device can be a tablet computer, mobile phone (such as a foldable phone, a large-screen phone, etc.), wearable device, in-vehicle device, augmented reality (AR) / virtual reality (VR) device, laptop computer, ultra-mobile personal computer (UMPC), netbook, personal digital assistant (PDA), smart TV, smart screen, high-definition TV, 4K TV, smart speaker, smart projector, and other Internet of Things (IoT) devices. This disclosure does not impose any restrictions on the specific type of electronic device.
[0038] It should be noted that the electronic device used to train and acquire the speech synthesis model and the electronic device used to execute speech synthesis services using the speech synthesis model can be different electronic devices or the same electronic device; this disclosure does not limit this. For example, the speech synthesis model can be trained and acquired by a server device, and then the server device can distribute the trained speech synthesis model to the terminal device / server device, which can then execute speech synthesis services based on the speech synthesis model. Alternatively, the speech synthesis model can be trained and acquired by a server device, and then the trained speech synthesis model can be deployed on the server device, which can then call the speech synthesis model to process speech synthesis services. This disclosure does not impose any restrictions, and in practical applications, it can be flexibly configured.
[0039] The speech synthesis model in this scheme will be introduced below.
[0040] The speech synthesis model in this scheme decouples itself into two sub-models by introducing acoustic features, including bottleneck features: a first feature extraction sub-model and a second feature extraction sub-model. The first feature extraction sub-model establishes a deep mapping between the text sequence and the acoustic features containing the bottleneck features; the second feature extraction sub-model establishes a deep mapping between the acoustic features containing the bottleneck features and the Mel-spectral features. Based on this, at least the following can be achieved:
[0041] 1. The two decoupled feature extraction sub-models can be trained using different sample audio.
[0042] Specifically, the model used to establish a deep mapping between text sequences and acoustic features containing bottleneck features, i.e., the first feature extraction sub-model, needs to be trained using high-quality first sample audio and the labeled text corresponding to the first sample audio as sample data.
[0043] The model used to establish a deep mapping between acoustic features containing bottleneck features and Mel spectral features, i.e., the second feature extraction sub-model, can be trained using second sample audio without corresponding text annotation. Since there is no need to annotate the text corresponding to the second sample audio, this can greatly reduce the cost of obtaining the second sample audio.
[0044] 2. By decoupling the speech synthesis model, the control of speech synthesis is achieved by making timbre and other features relatively independent.
[0045] Specifically, the acoustic features containing bottleneck features are features unrelated to timbre. Among them, the acoustic features output by the first feature extraction sub-model mainly include information in dimensions such as prosody and content.
[0046] The Mel spectrum features output by the second feature extraction sub-model include not only information on the prosody and content dimensions mentioned above, but also information on the timbre dimension.
[0047] 3. Reduced the requirements for third-sample audio with the target timbre.
[0048] This speech synthesis model can be trained with a small number of third-sample audio samples of the target timbre, and the final speech synthesis model can synthesize audio with the target timbre. Even if the quality of the third-sample audio is not high, such as non-standard pronunciation or disfluency, the speech synthesis model can still stably synthesize audio with the target timbre.
[0049] Because the second feature extraction sub-model, which controls the timbre, is trained using the second sample audio, it already possesses a high level of speech synthesis control capability for timbre. Therefore, even if the second feature extraction sub-model learns a small amount of third sample audio, it can still master the target timbre well.
[0050] The following detailed examples illustrate how to train and obtain a speech synthesis model. In these examples, an electronic device is used as an example, and the accompanying drawings are provided for detailed explanation.
[0051] in, Figure 1a A framework diagram for training and acquiring a speech synthesis model is shown; Figure 1b and Figure 1cThe structural diagrams of the first feature extraction sub-model and the second feature extraction sub-model included in the speech synthesis model are shown as examples.
[0052] Reference Figure 1a As shown, the speech synthesis model 100 includes a first feature extraction sub-model 101 and a second feature extraction sub-model 102. The training process for the speech synthesis model 100 includes training the first feature extraction sub-model 101 and training the second feature extraction sub-model 102.
[0053] The following sections describe the training process for the first feature extraction sub-model 101 and the training process for the second feature extraction sub-model 102.
[0054] I. Training the first feature extraction sub-model 101
[0055] The first feature extraction sub-model 101 is trained based on the labeled text corresponding to the first sample audio and the labeled acoustic features (hereinafter referred to as the labeled acoustic features corresponding to the first sample audio as the second acoustic features). By learning the relationship between the labeled text corresponding to the first sample audio and the second acoustic features, the first feature extraction sub-model 101 acquires the ability to establish a deep mapping between the text and the acoustic features containing the bottleneck features.
[0056] Specifically, the first feature extraction sub-model 101 is used to analyze the labeled text corresponding to the input first sample audio, model the intermediate feature sequence, perform feature transformation and dimensionality reduction on the intermediate feature sequence, and output the fifth acoustic feature corresponding to the labeled text.
[0057] Then, based on the second acoustic feature corresponding to the first sample audio, the fifth acoustic feature corresponding to the first sample audio, and the pre-constructed loss function, the loss function information for this round of training is calculated, and the coefficient values of the parameters included in the first feature extraction sub-model 101 are adjusted according to the loss function information for this round of training.
[0058] Through iterative training of multiple first sample audios, the labeled text corresponding to the first sample audios, and the second acoustic features (including the first labeled bottleneck features) corresponding to the first sample audios, a first feature extraction model 101 that satisfies the corresponding convergence conditions is finally obtained.
[0059] During training, the second acoustic feature corresponding to the first sample audio can be understood as the learning target of the first feature extraction sub-model 101.
[0060] The first sample audio can include high-quality audio files (or clean audio), and the corresponding annotation text can be a character sequence or phoneme sequence; this disclosure does not limit this. The first sample audio can be obtained through recording and multiple cleaning processes based on actual needs, or it can be obtained by filtering and cleaning TTS data stored in an audio database. This disclosure does not restrict the method of obtaining the first sample audio. Similarly, the annotation text corresponding to the first sample audio can also be obtained through repeated annotation and correction to ensure the accuracy of the annotation text.
[0061] Furthermore, the fifth acoustic feature corresponding to the labeled text can be understood as the predicted acoustic feature corresponding to the labeled text output by the first feature extraction sub-model 101. The fifth acoustic feature corresponding to the labeled text can also be understood as the fifth acoustic feature corresponding to the first sample audio.
[0062] In some embodiments, the second acoustic feature includes: a first labeled bottleneck feature corresponding to the first sample audio.
[0063] Among them, the bottleneck is a non-linear feature transformation technique and an effective dimensionality reduction technique. In the speech synthesis scenario for specific timbres mentioned in this solution, the bottleneck features can include information in dimensions such as prosody and content.
[0064] One possible implementation is that the first labeled bottleneck feature corresponding to the first sample audio can be obtained by the encoder of the end-to-end speech recognition (ASR) model.
[0065] In the following text, the end-to-end ASR model will be referred to as the ASR model.
[0066] For example, refer to Figure 1a As shown, the first sample audio can be input into the ASR model 104 to obtain the first labeled bottleneck feature corresponding to the first sample audio output by the encoder of the ASR model 104. The encoder of the ASR model 104 is equivalent to the extractor of the bottleneck feature in advance. In this scheme, the encoder of the ASR model 104 can be used to prepare sample data.
[0067] It should be noted that ASR model 104 may also include other modules, such as Figure 1a As shown, the ASR model 104 also includes a decoder and an attention network. No processing is required on the outputs of modules other than the encoder in the ASR model 104, and this disclosure does not limit the functionality or implementation of modules or networks other than the encoder in the ASR model.
[0068] The example of obtaining the labeled bottleneck features corresponding to the first sample audio through the encoder of ASR model 104 is merely an illustration and not a limitation on the implementation method of obtaining the first labeled bottleneck features corresponding to the first sample audio. In practical applications, other methods can also be used, and this disclosure does not impose any restrictions on this. For example, sample audio and corresponding labeled bottleneck features are stored in some databases, and these can also be obtained from the database.
[0069] In other embodiments, the second acoustic feature corresponding to the first sample audio includes: a first labeled bottleneck feature corresponding to the first sample audio and a first labeled fundamental frequency feature corresponding to the first sample audio.
[0070] The first labeled bottleneck feature can be described in detail in the previous example; for the sake of brevity, it will not be repeated here.
[0071] Pitch represents the subjective perception of the highness or lowness of a sound by the human ear. The pitch primarily depends on the fundamental frequency of the sound; a higher fundamental frequency results in a higher pitch, and a lower fundamental frequency results in a lower pitch. In speech synthesis, pitch is also a crucial factor affecting the synthesis effect. To enable the final speech synthesis model to control pitch, this scheme introduces both bottleneck features and fundamental frequency features, allowing the final first feature extraction sub-model 101 to output corresponding bottleneck and fundamental frequency features based on the input text.
[0072] Specifically, the first feature extraction sub-model 101 is used to analyze the labeled text corresponding to the input first sample audio, model the intermediate feature sequence, perform feature transformation and dimensionality reduction on the intermediate feature sequence, and output the fifth acoustic feature corresponding to the labeled text.
[0073] The fifth acoustic feature corresponding to the labeled text can be understood as the predicted acoustic feature corresponding to the labeled text output by the first feature extraction sub-model 101. The fifth acoustic feature corresponding to the labeled text can also be understood as the fifth acoustic feature corresponding to the first sample audio.
[0074] It should be noted that the second acoustic feature corresponding to the first sample audio includes the first labeled bottleneck feature and the first labeled fundamental frequency feature. Therefore, during the training process, the fifth acoustic feature corresponding to the first sample audio output by the first feature extraction sub-model 101 also includes the predicted bottleneck feature and the predicted fundamental frequency feature corresponding to the first sample audio.
[0075] Then, based on the second acoustic feature corresponding to the first sample audio, the fifth acoustic feature corresponding to the first sample audio, and the pre-constructed loss function, the loss function information for this round of training is calculated, and the coefficient values of the parameters included in the first feature extraction sub-model 101 are adjusted according to the loss function information.
[0076] Through iterative training of massive amounts of first sample audio, the labeled text corresponding to the first sample audio, and the second acoustic features (including the first labeled bottleneck feature and the first labeled fundamental frequency feature) corresponding to the first sample audio, a first feature extraction model 101 that satisfies the corresponding convergence condition is finally obtained.
[0077] During training, the second acoustic feature corresponding to the first sample audio can be understood as the learning target of the first feature extraction sub-model 101.
[0078] One possible implementation is that the first labeled fundamental frequency feature corresponding to the first sample audio can be obtained by analyzing the first sample audio using digital signal processing (DSP) methods. For example, such as... Figure 1a As shown, the first sample audio can be digitally processed by the digital signal processor 105 to obtain the first labeled fundamental frequency feature corresponding to the first sample audio. The specific implementation of the digital signal processor 105 is not limited, as long as it can extract the first labeled fundamental frequency feature corresponding to the input first sample audio.
[0079] Furthermore, the first labeled fundamental frequency feature corresponding to the first sample audio is not limited to being obtained through digital signal processing methods, and this disclosure does not limit the implementation method for obtaining the first labeled fundamental frequency feature. For example, sample audio and the labeled fundamental frequency feature corresponding to the sample audio are stored in some databases, and it can also be obtained from the database.
[0080] It should be noted that the convergence condition for the first feature extraction sub-model can be, but is not limited to, evaluation metrics such as the number of iterations or the loss threshold. This disclosure does not impose restrictions on the convergence condition for training the first feature extraction sub-model. Furthermore, training can be performed based on the labeled bottleneck features corresponding to the first sample audio, or based on the first labeled bottleneck features and the first labeled fundamental frequency features corresponding to the first sample audio, and the convergence conditions can differ.
[0081] Furthermore, training is performed based on the first annotation bottleneck features corresponding to the first sample audio, or based on the first annotation bottleneck features and the first annotation fundamental frequency features corresponding to the first sample audio. The loss function of the pre-constructed first feature extraction sub-model can be the same or different. This disclosure does not limit the implementation method of the loss function of the pre-constructed first feature extraction sub-model.
[0082] The network structure of the first feature extraction sub-model is shown below as an example.
[0083] Figure 1b An exemplary implementation of the first feature extraction sub-model 101 is shown. (Refer to...) Figure 1bAs shown, the first feature extraction sub-model 101 may include: a text encoder network 1011, an attention network 1012, and a decoder network 1013.
[0084] The text encoding network 1011 is used to receive text as input, analyze the context and temporal relationship of the input text, and model the intermediate feature sequence, which contains contextual information and temporal relationship.
[0085] Decoding network 1013 can adopt an autoregressive network structure, using the output of the previous time step as the input of the next time step.
[0086] Attention network 1012 is primarily used to output attention coefficients. These attention coefficients are then weighted and averaged with the intermediate feature sequence output by text encoding network 1011 to obtain a weighted average result. This weighted average result serves as another conditional input to decoding network 1013 at each time step. Decoding network 1013 performs feature transformation on the input (i.e., the weighted average result and the output of the previous time step) to output the predicted acoustic features corresponding to the text.
[0087] Combining the two aforementioned implementation methods, the predicted acoustic features corresponding to the text may include: the predicted bottleneck features corresponding to the text; or, the predicted acoustic features corresponding to the text may include: the bottleneck features corresponding to the text and the fundamental frequency features corresponding to the text.
[0088] Furthermore, the initial values of the coefficients of the parameters included in the first feature extraction sub-model 101 may be randomly generated, preset, or determined in other ways, and this disclosure does not limit them.
[0089] The first feature extraction sub-model 101 is iteratively trained using the labeled text corresponding to multiple first sample audios and the first labeled acoustic features corresponding to the first sample audios. The coefficient values of the parameters included in the first feature extraction sub-model 101 are continuously optimized until the convergence condition of the first feature extraction sub-model 101 is met, at which point the training of the first feature extraction sub-model 101 is stopped.
[0090] It should be understood that the first sample audio described above corresponds one-to-one with the corresponding labeled text, and they are paired sample data.
[0091] II. Training the second feature extraction sub-model 102
[0092] Training the second feature extraction sub-model 102 includes two stages. The first stage is to train the second feature extraction sub-model based on the second sample audio to obtain an intermediate model. The second stage is to fine-tune the intermediate model based on the third sample audio to obtain the final second feature extraction sub-model.
[0093] In this disclosure, the timbre of the second sample audio is not limited; in addition, the third sample audio is a sample audio with the target timbre.
[0094] The training process of the second feature extraction sub-model 102 is described in detail below:
[0095] Phase 1:
[0096] In the first stage of training, the second feature extraction sub-model 102 is used to iteratively train based on the second sample audio to obtain an intermediate model.
[0097] The second feature extraction sub-model 102 learns the mapping relationship between the third acoustic feature corresponding to the second sample audio and the first labeled Mel spectral feature of the second sample audio to obtain an intermediate model with a certain speech synthesis control capability for timbre. The Mel spectral feature includes timbre information.
[0098] This disclosure does not limit parameters such as the timbre, duration, storage format, or number of second sample audio files. The second sample audio files may include audio with a specific target timbre, or audio with a non-target timbre.
[0099] In the first stage of training, the second feature extraction sub-model 102 is used to analyze the second acoustic features corresponding to the input second sample audio and output the predicted Mel spectrum features corresponding to the second sample audio. Then, based on the first labeled Mel spectrum features corresponding to the second sample audio and the predicted Mel spectrum features corresponding to the second sample audio, the coefficient values of the parameters included in the second feature extraction sub-model 102 are adjusted. Through continuous iterative training of the second feature extraction sub-model 102 with a large number of second sample audio samples, an intermediate model is obtained.
[0100] In the first stage of training, the first labeled Mel spectrum features can be understood as the learning objective of the second feature extraction sub-model 102 in the first stage.
[0101] Since the input to the second feature extraction sub-model 102 is the third acoustic feature corresponding to the second sample audio, the second sample audio does not need to be labeled with corresponding text, thus greatly reducing the time and manpower costs associated with acquiring the second sample audio. Furthermore, a large amount of audio can be obtained at a low cost as second sample audio for iterative training of the second feature extraction sub-model 102. By training the second feature extraction sub-model 102 with a large amount of second sample audio, the intermediate model acquires a high level of control over speech synthesis based on timbre.
[0102] Phase Two:
[0103] In the second stage, the intermediate model is trained based on the third sample, so that the intermediate model learns the target timbre and obtains the speech synthesis control capability for the target timbre.
[0104] It should be noted that since the intermediate model already has a high degree of control over speech synthesis based on timbre, the requirements for the third sample audio have been reduced. For example, the requirements for the duration and quality of the third sample audio have been reduced. Even if the duration of the third sample audio is short or the pronunciation is unclear, the final second feature extraction sub-model 102 obtained through training can still achieve a high degree of control over speech synthesis based on the target timbre.
[0105] Furthermore, the third sample audio has a target timbre. The third sample audio can be audio recorded by the user or audio with the desired timbre uploaded by the user. This disclosure does not limit this.
[0106] Specifically, the fourth acoustic feature corresponding to the third sample audio is input into the intermediate model to obtain the predicted Mel spectrum feature corresponding to the third sample audio output by the intermediate model; then, based on the second labeled Mel spectrum feature corresponding to the third sample audio and the predicted Mel spectrum feature corresponding to the third sample audio, the loss function information corresponding to this round of training is calculated; according to the loss function information, the coefficient values of the parameters included in the intermediate model are adjusted to obtain the final second feature extraction sub-model 102.
[0107] During the second phase of training, the second labeled Mel spectrum features corresponding to the third sample audio can be understood as the learning target of the intermediate model.
[0108] Based on the aforementioned introduction of the first feature extraction sub-model 101, during the training process, if the fifth acoustic feature output by the first feature extraction sub-model 101 based on the labeled text of the input first sample audio includes the predicted bottleneck feature, that is, the first feature extraction sub-model 101 can achieve the mapping from text to bottleneck feature, then the third acoustic feature corresponding to the second sample audio input to the second feature extraction sub-model 102 includes the second labeled bottleneck feature corresponding to the second sample audio, and the fourth acoustic feature corresponding to the third sample audio input to the intermediate model includes the third labeled bottleneck feature corresponding to the third sample audio.
[0109] The second and third annotation bottleneck features can be obtained by extracting bottleneck features from the second and third sample audio samples respectively using the encoder of the ASR model. The implementation method is similar to that of obtaining the first annotation bottleneck features, and for the sake of simplicity, it will not be described in detail here.
[0110] During training, if the first feature extraction sub-model 101 outputs a fifth acoustic feature that includes a predicted bottleneck feature and a predicted fundamental frequency feature based on the labeled text of the input first sample audio, that is, the first feature extraction sub-model 101 can achieve the mapping from text to bottleneck features and fundamental frequency features, then the third acoustic feature corresponding to the second sample audio input to the second feature extraction sub-model 102 includes the second labeled bottleneck feature and the second labeled fundamental frequency feature corresponding to the second sample audio, and the fourth acoustic feature corresponding to the third sample audio input to the intermediate model includes the third labeled bottleneck feature and the third labeled fundamental frequency feature corresponding to the third sample audio.
[0111] The second and third annotation bottleneck features can be obtained by extracting bottleneck features from the second and third sample audio samples respectively using the encoder of the ASR model, similar to the method of obtaining the first annotation bottleneck features. The second and third annotation fundamental frequency features can be obtained by analyzing the second and third sample audio samples respectively using digital signal processing technology, similar to the method of obtaining the first annotation fundamental frequency features. For the sake of simplicity, they will not be elaborated here.
[0112] In summary, during the training process, the input of the second feature extraction sub-model 102 and the output of the first feature extraction sub-model 101 remain consistent.
[0113] Furthermore, when training the second feature extraction sub-model 102, the initial values of the coefficients corresponding to each parameter included in the second feature extraction sub-model 102 can be preset or randomly initialized, and this disclosure does not limit this.
[0114] Furthermore, the pre-constructed loss functions used in the first and second training phases can be the same or different, and this disclosure does not impose any restrictions on this.
[0115] in, Figure 1c An exemplary implementation of the second feature extraction sub-model 102 is shown. (Refer to...) Figure 1c As shown, the second feature extraction sub-model 102 can be implemented using a self-attention network structure.
[0116] Figure 1c In this model, the second feature extraction sub-model 102 includes a convolutional network 1021 and one or more residual networks 1022. Each residual network 1022 includes a self-attention network 1022a and a linear network 1022b.
[0117] The convolutional network 1021 is mainly used to perform convolution processing on the acoustic features corresponding to the input sample audio to model local feature information. The convolutional network 1021 may include one or more convolutional layers; this disclosure does not limit the number of convolutional layers included in the convolutional network 1021. Furthermore, the convolutional network 1021 inputs the local feature information to the connected residual network 1022.
[0118] The aforementioned one or more residual networks 1022, after passing through one or more residual networks 1022, are converted into Mel spectrum features.
[0119] It should be understood that intermediate models and Figure 1c The second feature extraction sub-model 102 shown has the same structure, but the weight coefficients of the included parameters are not exactly the same.
[0120] By training the first feature extraction sub-model 101 and the second feature extraction sub-model 102 respectively, a first feature extraction model and a second feature extraction model that meet the requirements of speech synthesis are finally obtained; then the first feature extraction model and the second feature extraction model are spliced together to obtain a speech synthesis model that can synthesize the target timbre.
[0121] In some cases, the speech synthesis model 100 may also include a vocoder 103. The vocoder 103 is used to convert the Mel-spectral features output by the second feature extraction sub-model 102 into audio. Of course, the vocoder can also be a standalone module, not bound to the speech synthesis model. Furthermore, this scheme does not restrict the specific type of vocoder.
[0122] exist Figures 1a to 1c Based on the illustrated embodiment, this solution can also be trained using sample audio from different languages, so that the final speech synthesis model has the ability to control speech synthesis across styles.
[0123] Optionally, the first sample audio and the second sample audio are in the same language; the first sample audio and the third sample audio are in different languages; and the second sample audio and the third sample audio are in different languages.
[0124] Assuming that the first and second sample audio samples are both Chinese audio data, and the third sample audio sample with the target timbre is English audio data, then the final speech synthesis model can synthesize Chinese audio with an English speaking style, and the Chinese audio has the target timbre.
[0125] Assuming that the first and second sample audio samples are both English audio data, and the third sample audio sample with the target timbre is Chinese audio data, then the final speech synthesis model can synthesize Chinese audio with a Chinese speaking style, and the Chinese audio has the target timbre.
[0126] In practical applications, the first and second sample audio files can be of any language, while the third sample audio file can be in a different language than the first sample audio file, and also a different language than the second sample audio file, thus enabling cross-style speech synthesis. Furthermore, stable cross-style speech synthesis can be achieved even with only a small number of third sample audio files.
[0127] In the above Figures 1a to 1c Based on the illustrated embodiment, the target speech synthesis model obtained through training has the ability to stably synthesize audio with the target timbre. Based on this, the target speech synthesis model can be used to process corresponding speech synthesis services.
[0128] Figure 2 A flowchart illustrating a speech synthesis method provided in an embodiment of this disclosure. (Refer to...) Figure 2 As shown, the speech synthesis method provided in this embodiment includes:
[0129] S201. Obtain the text to be processed.
[0130] The text to be processed may include either a character sequence or a phoneme sequence. The text to be processed, including either a character sequence or a phoneme sequence, is used to synthesize the target audio timbre (hereinafter referred to as: target audio).
[0131] This embodiment does not impose restrictions on parameters such as the length, language, and storage format of the text to be processed. The language of the text to be processed is consistent with the language of the first and second sample audio samples during the training phase; for example, if the first and second sample audio samples are in Chinese, then the text to be processed will also be in Chinese.
[0132] S202. Input the text to be processed into the target speech synthesis model and obtain the Mel spectrum sequence corresponding to the text to be processed output by the target speech synthesis model.
[0133] In combination with the above Figures 1a to 1c In the illustrated embodiment, the target speech synthesis model includes a trained first feature extraction sub-model and a second feature extraction sub-model. The first and second feature extraction sub-models can be found in the detailed description of the foregoing embodiments, and will not be repeated here.
[0134] In some embodiments, the text to be processed is input into the target speech synthesis model. The first feature extraction model extracts features from the text to be processed and outputs the bottleneck features corresponding to the text to be processed. The second feature extraction model receives the bottleneck features corresponding to the text to be processed as input and outputs the Mel spectrum features corresponding to the text to be processed.
[0135] In other embodiments, the text to be processed is input into the target speech synthesis model. The first feature extraction model extracts features from the text to be processed and outputs the bottleneck features and fundamental frequency features corresponding to the text to be processed. The second feature extraction model receives the bottleneck features and fundamental frequency features corresponding to the text to be processed as input and outputs the Mel spectrum features corresponding to the text to be processed.
[0136] S203. Based on the Mel spectrum features corresponding to the text to be processed, obtain the target audio corresponding to the text to be processed, wherein the target audio has a target timbre.
[0137] Based on any of the above methods, and further based on a vocoder, Mel spectral features can be converted into audio with the target timbre.
[0138] In some cases, the vocoder can be part of the target speech synthesis model, in which case the target speech synthesis model can output audio with the target timbre; in other cases, the vocoder is an independent module outside the target speech synthesis model. The target speech synthesis model outputs the Mel spectrum features corresponding to the text to be processed, and the vocoder receives the Mel spectrum features corresponding to the text to be processed as input and converts them into audio with the target timbre.
[0139] The speech synthesis method provided in this embodiment synthesizes speech based on the specific timbre of the target speech synthesis model by inputting the text sequence to be processed into the target speech synthesis model. In the speech synthesis scenario, users can upload third-party sample audio to specify the timbre of the synthesized audio, and can even upload third-party sample audio in different languages to specify the desired speaking style, thus meeting the user's personalized needs.
[0140] Figure 3 This is a schematic diagram of a speech synthesis device provided according to an embodiment of the present disclosure. (Refer to...) Figure 3 As shown, the speech synthesis device 300 provided in this embodiment includes:
[0141] The acquisition module 301 is used to acquire the text to be processed.
[0142] As shown above, the text to be processed can include a sequence of characters or a sequence of phonemes.
[0143] Processing module 302 is used to input the text to be processed into a target speech synthesis model and obtain a first Mel-spectral sequence corresponding to the text to be processed output by the target speech synthesis model; wherein, the target speech synthesis model includes: a first feature extraction sub-model and a second feature extraction sub-model, the first feature extraction sub-model is used to output a first acoustic feature based on the input text sequence, the first acoustic feature including a bottleneck feature corresponding to the text to be processed; the second feature extraction sub-model is used to output Mel-spectral features corresponding to the text to be processed based on the input first acoustic feature;
[0144] The processing module 302 is further configured to obtain the target audio corresponding to the text to be processed based on the Mel spectrum features corresponding to the text to be processed, wherein the target audio has a target timbre.
[0145] In some possible implementations, the first feature extraction sub-model is trained based on the labeled text corresponding to the first sample audio and the second acoustic features corresponding to the first sample audio, wherein the second acoustic features include the first labeled bottleneck features corresponding to the first sample audio.
[0146] In some possible implementations, the second feature extraction sub-model is obtained by training based on the third acoustic feature and the first labeled Mel spectral feature corresponding to the second sample audio, and the fourth acoustic feature and the second labeled Mel spectral feature corresponding to the third sample audio.
[0147] The third acoustic feature includes the second labeled bottleneck feature corresponding to the second sample audio; the fourth acoustic feature includes the third labeled bottleneck feature corresponding to the third sample audio; and the third sample audio is a sample audio with the target timbre.
[0148] In some possible implementations, the first annotation bottleneck feature corresponding to the first sample audio, the second annotation bottleneck feature corresponding to the second sample audio, and the third annotation bottleneck feature corresponding to the third sample audio are obtained by the encoder of the end-to-end speech recognition model extracting bottleneck features from the input first sample audio, second sample audio, and third sample audio, respectively.
[0149] In some possible implementations, the second acoustic feature further includes: a first labeled fundamental frequency feature corresponding to the first sample audio; the third acoustic feature further includes: a second labeled fundamental frequency feature corresponding to the second sample audio; the fourth acoustic feature further includes: a third labeled fundamental frequency feature corresponding to the third sample audio.
[0150] Accordingly, the first acoustic feature output by the first feature extraction sub-model also includes: the fundamental frequency feature corresponding to the text to be processed.
[0151] In some possible implementations, the first labeled fundamental frequency feature corresponding to the first sample audio, the second labeled fundamental frequency feature corresponding to the second sample audio, and the third labeled fundamental frequency feature corresponding to the third sample audio are obtained by performing digital signal processing on the first sample audio, the second sample audio, and the third sample audio respectively.
[0152] In some possible implementations, the language of the first sample audio is the same as that of the second sample audio; or the language of the first sample audio is different from that of the third sample audio.
[0153] The speech synthesis device provided in this embodiment can be used to execute the technical solutions of any of the above method embodiments. Its implementation principle and technical effect are similar, and can be referred to the detailed description of the foregoing method embodiments. For the sake of brevity, it will not be repeated here.
[0154] Figure 4 This is a schematic diagram of the structure of an electronic device provided according to an embodiment of the present disclosure. (Refer to...) Figure 4 As shown, the electronic device 400 provided in this embodiment includes a memory 401 and a processor 402.
[0155] The memory 401 can be a separate physical unit, connected to the processor 402 via a bus 403. Alternatively, the memory 401 and processor 402 can be integrated together, implemented in hardware, etc.
[0156] The memory 401 is used to store program instructions, and the processor 402 calls the program instructions to execute the operations of any of the above method embodiments.
[0157] Optionally, when some or all of the methods in the above embodiments are implemented by software, the electronic device 400 may also include only the processor 402. The memory 401 for storing programs is located outside the electronic device 400, and the processor 402 is connected to the memory via circuits / wires to read and execute the programs stored in the memory.
[0158] Processor 402 can be a central processing unit (CPU), a network processor (NP), or a combination of CPU and NP.
[0159] The processor 402 may further include a hardware chip. This hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0160] Memory 401 may include volatile memory, such as random-access memory (RAM); memory may also include non-volatile memory, such as flash memory, hard disk drive (HDD) or solid-state drive (SSD); memory may also include combinations of the above types of memory.
[0161] This disclosure also provides a computer-readable storage medium (also referred to as a readable storage medium) that includes computer program instructions, which, when executed by at least one processor of an electronic device, perform the technical solutions of any of the above method embodiments.
[0162] This disclosure also provides a program product including computer program instructions stored in a readable storage medium. At least one processor of the electronic device can read the computer program instructions from the readable storage medium, and the at least one processor executes the computer program instructions to cause the electronic device to perform the technical solution of any of the above method embodiments.
[0163] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0164] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A speech synthesis method, characterized in that, include: Get the text to be processed; The text to be processed is input into a target speech synthesis model, and the Mel spectrum sequence corresponding to the text to be processed is obtained from the output of the target speech synthesis model; wherein, the target speech synthesis model includes: a first feature extraction sub-model and a second feature extraction sub-model, the first feature extraction sub-model is used to output a first acoustic feature based on the input text to be processed, the first acoustic feature including a bottleneck feature corresponding to the text to be processed; the second feature extraction sub-model is used to output the Mel spectrum feature corresponding to the text to be processed based on the input first acoustic feature; Based on the Mel spectrum features corresponding to the text to be processed, the target audio corresponding to the text to be processed is obtained, and the target audio has a target timbre.
2. The method according to claim 1, characterized in that, The first feature extraction sub-model is trained based on the labeled text corresponding to the first sample audio and the second acoustic features corresponding to the first sample audio. The second acoustic features include the first labeled bottleneck features corresponding to the first sample audio.
3. The method according to claim 2, characterized in that, The second feature extraction sub-model is obtained by training based on the third acoustic feature and the first labeled Mel spectral feature corresponding to the second sample audio, and the fourth acoustic feature and the second labeled Mel spectral feature corresponding to the third sample audio. The third acoustic feature includes the second labeled bottleneck feature corresponding to the second sample audio; the fourth acoustic feature includes the third labeled bottleneck feature corresponding to the third sample audio; and the third sample audio is a sample audio with the target timbre.
4. The method according to claim 3, characterized in that, The first annotation bottleneck feature corresponding to the first sample audio, the second annotation bottleneck feature corresponding to the second sample audio, and the third annotation bottleneck feature corresponding to the third sample audio are obtained by the encoder of the end-to-end speech recognition model extracting bottleneck features from the input first sample audio, second sample audio, and third sample audio respectively.
5. The method according to claim 3, characterized in that, The second acoustic feature further includes: a first labeled fundamental frequency feature corresponding to the first sample audio; the third acoustic feature further includes: a second labeled fundamental frequency feature corresponding to the second sample audio; the fourth acoustic feature further includes: a third labeled fundamental frequency feature corresponding to the third sample audio. Accordingly, the first acoustic feature output by the first feature extraction sub-model also includes: the fundamental frequency feature corresponding to the text to be processed.
6. The method according to claim 5, characterized in that, The first labeled fundamental frequency feature corresponding to the first sample audio, the second labeled fundamental frequency feature corresponding to the second sample audio, and the third labeled fundamental frequency feature corresponding to the third sample audio are obtained by performing digital signal processing on the first sample audio, the second sample audio, and the third sample audio respectively.
7. The method according to claim 3, characterized in that, The language of the first sample audio is the same as that of the second sample audio; the language of the first sample audio is different from that of the third sample audio.
8. A speech synthesis device, characterized in that, include: The acquisition module is used to acquire the text to be processed; A processing module is configured to input the text to be processed into a target speech synthesis model and obtain a first Mel-spectral sequence corresponding to the text to be processed output by the target speech synthesis model; wherein, the target speech synthesis model includes: a first feature extraction sub-model and a second feature extraction sub-model, the first feature extraction sub-model being configured to output a first acoustic feature based on the input text sequence, the first acoustic feature including a bottleneck feature corresponding to the text to be processed; the second feature extraction sub-model being configured to output a Mel-spectral feature corresponding to the text to be processed based on the input first acoustic feature; The processing module is further configured to obtain the target audio corresponding to the text to be processed based on the Mel spectrum features corresponding to the text to be processed, wherein the target audio has a target timbre.
9. An electronic device, characterized in that, include: Memory, processor, and computer program instructions; The memory is configured to store the computer program instructions; The processor is configured to execute the computer program instructions to implement the speech synthesis method as described in any one of claims 1 to 7.
10. A readable storage medium, characterized in that, include: Computer program instructions; The computer program instructions are executed by at least one processor of the electronic device to implement the speech synthesis method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Method and device for generating audio, equipment and medium
CN111402842A
Voice conversion system, method and application
CN112017644A