Voice signal generation method, device and electronic equipment
By using trained prosody coding predictors, duration predictors, and spectrum predictors, combined with adversarial generative networks and timbre encoders, speech signals with diverse timbres are generated, solving the problems of single timbre and bland prosody in existing technologies and improving the expressiveness of speech synthesis.
Patent Information
- Application Number
- CN202411339594.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-23
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2044-09-23
AI Technical Summary
Existing speech synthesis technology has difficulty in accurately capturing and reproducing the unique voice characteristics and complex emotional changes of different characters when faced with highly role-based and emotional tasks such as novel reading, resulting in a monotonous timbre and bland rhythmic performance, and is unable to meet the diverse needs in multi-character dialogue scenarios.
The pre-trained prosody coding predictor, duration predictor and spectrum predictor are used, combined with the adversarial generative network and timbre encoder to generate speech signals with diverse timbre by obtaining the text features and timbre embedding information of the target text.
It improves the rhythmic performance of speech synthesis, realizes the generation of speech with diverse timbres, and meets the needs of multi-role dialogue scenarios in fields such as literary interpretation and audiobooks.
Smart Images

Figure CN119169993B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a method, device and electronic device for generating a speech signal. Background Art
[0002] In current speech synthesis technology, the typical process involves three main stages: First, the front-end module analyzes the plain text input and converts it into a set of structured text features; then, the acoustic model uses these features to generate corresponding acoustic parameters; finally, the vocoder converts these acoustic parameters into an audible speech waveform. This process makes text-to-speech conversion possible and plays an important role in application scenarios such as news broadcasting and navigation guidance.
[0003] However, although this technology has demonstrated excellent performance in many aspects, its limitations gradually become apparent when faced with speech synthesis tasks that require a high degree of characterization and emotion, such as novel reading. Specifically, the novel voices generated by current speech synthesis technology often exhibit problems such as monotonous timbre and bland rhythmic performance. It is difficult to accurately capture and reproduce the unique voice characteristics and complex emotional changes of different characters in the novel, and thus cannot meet users' diverse needs for speech synthesis in multi-character dialogue scenarios. This situation not only limits the further expansion of speech synthesis technology in fields such as literary interpretation and audiobooks, but also stimulates an urgent desire within and outside the industry to develop more intelligent, flexible, and expressive speech synthesis technology. Summary of the Invention
[0004] The technical problem to be solved by this application is to provide a method, device and electronic device for generating a speech signal, which can improve the rhythmic performance of speech synthesis and realize the synthesis of speech with diverse timbres. The specific solution is as follows:
[0005] A method for generating a speech signal, comprising:
[0006] Obtaining a target text to be processed; the target text includes N sentence texts and a narration dialogue tag for each sentence text, each narration dialogue tag is used to indicate the type of the sentence text to which it belongs, and the type of the sentence text is one of a narration type and a dialogue type;
[0007] Based on a pre-trained prosodic coding predictor, the target text, the text features of the target text, and the timbre embedding information of each sentence text, the prosodic information of the central sentence text is obtained; the timbre embedding information of each sentence text is generated based on the narration dialogue label and the reference speech of each sentence text; the central sentence text is the Mth sentence text in the target text, wherein 1 <M<N;
[0008] Obtaining duration information of the central sentence text based on a pre-trained duration predictor, the target text, text features of the target text, and timbre embedding information of each sentence text;
[0009] Based on a pre-trained spectrum predictor, the text features of the central sentence text, the timbre embedding information of the central sentence text, the duration information of the central sentence text and the prosody information of the central sentence text, the speech signal of the central sentence text is obtained.
[0010] In the above method, optionally, the training process of the prosody coding predictor includes:
[0011] Obtaining a generative adversarial network and a first training dataset; the generative adversarial network includes a generator and a discriminator; the generator includes an initial prosody code predictor and a first timbre encoder corresponding to the initial prosody code; the first training dataset includes a plurality of first training data, each of the first training data includes N first training sentence texts, as well as text features, pre-extracted prosody codes, narration dialogue labels, and voice information of the N first training sentence texts;
[0012] Selecting first target training data currently used for training from each first training data in the first training data set;
[0013] Obtaining timbre embedding information of the first target training data according to the first timbre encoder, the narration dialogue tag of the first target training data, and the voice information;
[0014] Obtaining prosodic codes of the N first training sentence texts in the first target training data according to the initial prosodic code predictor, the N first training sentence texts in the first target training data, and text features;
[0015] Obtaining loss function values of the generator and the discriminator according to the generator, the discriminator, the first target training data, and the prosodic encoding of N first training sentence texts in the first target training data;
[0016] Using the loss function value of the generator to update the model parameters of the first timbre encoder and the model parameters of the initial prosody coding predictor in the generator; and using the loss function value of the discriminator to update the model parameters of the discriminator;
[0017] If the updated generator and the discriminator do not meet the preset first training completion condition, returning to the step of selecting the first target training data currently used for training from each training data in the training data set;
[0018] When the updated generator and the discriminator meet a preset first training completion condition, the initial prosody coding predictor in the updated generator is determined as the trained prosody coding predictor.
[0019] In the above method, optionally, the training process of the spectrum predictor includes:
[0020] Obtaining an initial spectrum predictor to be trained, a second timbre encoder, a prosody code extractor, and a second training data set; the second training data set includes a plurality of second training data, each second training data including a second training sentence text, and prosodic features, text features, voice information, narration dialogue labels, and a target signal of the second training sentence text; the prosodic features include duration, fundamental frequency, and energy;
[0021] Selecting second target training data currently used for training from each second training data in the second training data set;
[0022] Inputting the prosodic features in the second target training data into the prosodic code extractor to obtain the prosodic code of the second training sentence text in the second target training data;
[0023] Obtaining timbre embedding information of a second training sentence text in the second target training data according to the second timbre encoder, the narration dialogue tag of the second target training data, and the voice information;
[0024] Obtaining a speech signal of a second training sentence text in the second target training data according to the initial spectrum predictor, the text features in the second target training data, the duration in the prosodic features, the prosodic coding of the second training sentence text, and the timbre embedding information;
[0025] Calculating a first loss function value using a preset first loss function, a speech signal of a second training sentence text in the second target training data, and a target signal;
[0026] Updating model parameters of the initial spectrum predictor, the second timbre encoder, and the prosody code extractor using the first loss function value;
[0027] If the updated initial spectrum predictor does not meet the preset second training completion condition, returning to the step of selecting the second target training data currently used for training from each second training data in the second training data set;
[0028] In a case where the updated initial spectrum predictor meets a preset second training completion condition, the initial spectrum predictor meeting the second training completion condition is determined as a trained spectrum predictor.
[0029] In the above method, optionally, the training process of the duration predictor includes:
[0030] Obtaining an initial duration predictor to be trained, a third timbre encoder, and a third training data set; the third training data set includes a plurality of third training data; each of the third training data includes N third training sentence texts, and text features, narration dialogue labels, voice information, and target duration information of the N third training sentence texts;
[0031] Selecting third target training data currently used for training from each third training data in the third training data set;
[0032] Obtaining timbre embedding information of N third training sentence texts in the third target training data according to the third timbre encoder, the narration dialogue tag of the third target training data, and the voice information;
[0033] Obtaining duration information of the N third training sentence texts in the third target training data according to the initial duration predictor, the N third training sentence texts in the third target training data, and text features;
[0034] Calculate a second loss function value using a preset second loss function, duration information of the N third training sentence texts in the third target training data, and target duration information;
[0035] Updating the model parameters of the initial duration predictor and the third timbre encoder using the second loss function value;
[0036] If the updated initial duration predictor does not meet the preset third training completion condition, returning to the step of selecting third target training data currently used for training from each third training data in the third training data set;
[0037] In the case that the updated initial duration predictor meets the preset third training completion condition, the initial duration predictor meeting the third training completion condition is used as the trained duration predictor.
[0038] Optionally, the method of obtaining the speech signal of the central sentence text based on a pre-trained spectrum predictor, text features of the central sentence text, timbre embedding information of the central sentence text, duration information of the central sentence text, and prosody information of the central sentence text includes:
[0039] Obtaining a timbre encoder corresponding to the spectrum predictor;
[0040] Obtaining timbre embedding information of the central sentence text based on the timbre encoder corresponding to the spectrum predictor, the narration dialogue tag of each sentence text, and the reference voice information of each sentence text;
[0041] Multiplying the timbre embedding information of the central sentence text by a preset control coefficient to obtain target timbre embedding information;
[0042] Based on a pre-trained spectrum predictor, the text features of the central sentence text, the target timbre embedding information of the central sentence text, the duration information of the central sentence text and the prosody information of the central sentence text, the speech signal of the central sentence text is obtained.
[0043] A speech signal generating device, comprising:
[0044] an acquisition unit, configured to acquire a target text to be processed; the target text includes N sentence texts and a narration dialogue tag for each sentence text, wherein each narration dialogue tag is used to indicate a type of the sentence text to which it belongs, and the type of the sentence text is one of a narration type and a dialogue type;
[0045] The first execution unit is configured to obtain the prosody information of the central sentence text based on a pre-trained prosody coding predictor, the target text, the text features of the target text, and the timbre embedding information of each sentence text; the timbre embedding information of each sentence text is generated based on the narration dialogue tag of each sentence text; the central sentence text is the Mth sentence text in the target text, wherein 1 <M<N;
[0046] A second execution unit is configured to obtain duration information of the central sentence text based on a pre-trained duration predictor, the target text, text features of the target text, and timbre embedding information of each sentence text;
[0047] The third execution unit is used to obtain the speech signal of the central sentence text based on a pre-trained spectrum predictor, the text features of the central sentence text, the timbre embedding information of the central sentence text, the duration information of the central sentence text and the rhythm information of the central sentence text.
[0048] In the above device, optionally, the first execution unit includes:
[0049] A first acquisition subunit is configured to acquire a generative adversarial network and a first training dataset; the generative adversarial network includes a generator and a discriminator; the generator includes an initial prosody code predictor and a first timbre encoder corresponding to the initial prosody code; the first training dataset includes a plurality of first training data, each of the first training data includes N first training sentence texts, as well as text features, pre-extracted prosody codes, narration dialogue tags, and voice information of the N first training sentence texts;
[0050] A first selection subunit is configured to select first target training data currently used for training from each first training data in the first training data set;
[0051] a first execution subunit, configured to obtain timbre embedding information of the first target training data according to the first timbre encoder, the narration dialogue tag of the first target training data, and the voice information;
[0052] A second execution subunit is configured to obtain prosodic codes of the N first training sentence texts in the first target training data according to the initial prosodic code predictor, the N first training sentence texts in the first target training data, and text features;
[0053] A first calculation subunit is configured to obtain loss function values of the generator and the discriminator based on the generator, the discriminator, the first target training data, and the prosodic encoding of the N first training sentence texts in the first target training data;
[0054] A first updating subunit is configured to update the model parameters of the first timbre encoder and the model parameters of the initial prosody coding predictor in the generator using the loss function value of the generator; and to update the model parameters of the discriminator using the loss function value of the discriminator;
[0055] a third execution subunit, configured to, if the updated generator and the discriminator do not meet the preset first training completion condition, return to trigger the first selection subunit to execute the step of selecting first target training data currently used for training from each training data in the training data set;
[0056] The first determining subunit is configured to determine the initial prosody coding predictor in the updated generator as the trained prosody coding predictor when the updated generator and the discriminator meet a preset first training completion condition.
[0057] In the above device, optionally, the third execution unit includes:
[0058] a second acquisition subunit, configured to acquire an initial spectrum predictor to be trained, a second timbre encoder, a prosody code extractor, and a second training data set; the second training data set includes a plurality of second training data, each second training data including a second training sentence text, and prosodic features, text features, voice information, narration dialogue tags, and a target signal of the second training sentence text; the prosodic features include duration, fundamental frequency, and energy;
[0059] A second selection subunit is used to select second target training data currently used for training from each second training data in the second training data set;
[0060] A first input subunit, configured to input the prosodic features in the second target training data into the prosodic code extractor to obtain a prosodic code of a second training sentence text in the second target training data;
[0061] a fourth execution subunit, configured to obtain timbre embedding information of a second training sentence text in the second target training data according to the second timbre encoder, the narration dialogue tag of the second target training data, and the voice information;
[0062] a fifth execution subunit, configured to obtain a speech signal of a second training sentence text in the second target training data based on the initial spectrum predictor, the text features in the second target training data, the duration in the prosodic features, the prosodic coding of the second training sentence text, and the timbre embedding information;
[0063] A second calculation subunit is configured to calculate a first loss function value by using a preset first loss function, a speech signal of a second training sentence text in the second target training data, and a target signal;
[0064] a second updating subunit, configured to update model parameters of the initial spectrum predictor, the second timbre encoder, and the prosody code extractor using the first loss function value;
[0065] a sixth execution subunit, configured to, if the updated initial spectrum predictor does not meet the preset second training completion condition, return to trigger the second acquisition subunit to execute the step of selecting second target training data currently used for training from each second training data in the second training data set;
[0066] The second determining subunit is configured to, when the updated initial spectrum predictor satisfies a preset second training completion condition, determine the initial spectrum predictor that satisfies the second training completion condition as a trained spectrum predictor.
[0067] In the above device, optionally, the second execution unit includes:
[0068] a third acquisition subunit, configured to acquire an initial duration predictor to be trained, a third timbre encoder, and a third training data set; the third training data set comprising a plurality of third training data; each of the third training data comprising N third training sentence texts, and text features, narration dialogue labels, voice information, and target duration information of the N third training sentence texts;
[0069] A third selection subunit is used to select third target training data currently used for training from each third training data in the third training data set;
[0070] a seventh execution subunit, configured to obtain timbre embedding information of N third training sentence texts in the third target training data according to the third timbre encoder, the narration dialogue tag of the third target training data, and the voice information;
[0071] an eighth execution subunit, configured to obtain duration information of the N third training sentence texts in the third target training data according to the initial duration predictor, the N third training sentence texts in the third target training data, and text features;
[0072] A third calculation subunit is configured to calculate a second loss function value by using a preset second loss function, duration information of the N third training sentence texts in the third target training data, and target duration information;
[0073] a third updating subunit, configured to update the model parameters of the initial duration predictor and the third timbre encoder using the second loss function value;
[0074] a ninth execution subunit, configured to, if the updated initial duration predictor does not satisfy the preset third training completion condition, return to trigger the third selection subunit to execute a process of selecting third target training data currently used for training from each third training data in the third training data set;
[0075] The tenth execution subunit is configured to, when the updated initial duration predictor satisfies a preset third training completion condition, use the initial duration predictor that satisfies the third training completion condition as the trained duration predictor.
[0076] An electronic device includes a memory and one or more instructions, wherein the one or more instructions are stored in the memory and configured to cause one or more processors to execute the above-mentioned method for generating a speech signal.
[0077] Based on the speech signal generation method, device and electronic device provided by the embodiments of the present application, the method includes: obtaining a target text to be processed; the target text includes N sentence texts and the voice-over dialogue tags of each sentence text, and each voice-over dialogue tag is used to indicate the type of the sentence text to which it belongs, and the type of the sentence text is one of the voice-over type and the dialogue type; based on a pre-trained prosody encoding predictor, the target text, the text features of the target text, and the timbre embedding information of each sentence text, obtaining the prosody information of the central sentence text; the timbre embedding information of each sentence text is generated based on the voice-over dialogue tag of each sentence text and a reference speech; the central sentence text is the Mth sentence text in the target text, where 1 < M < N; based on a pre-trained duration predictor, the target text, the text features of the target text, and the timbre embedding information of each sentence text, obtaining the duration information of the central sentence text; based on a pre-trained spectral predictor, the text features of the central sentence text, the timbre embedding information of the central sentence text, the duration information of the central sentence text, and the prosody information of the central sentence text, obtaining the speech signal of the central sentence text. Applying the method provided by the embodiments of the present application can improve the prosody performance of speech synthesis and can realize the synthesis of voices with diverse timbres. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on the provided drawings without creative efforts.
[0079] Figure 1 It is a flowchart of a speech signal generation method provided by the present application;
[0080] Figure 2 It is a schematic structural diagram of a prosody encoding predictor provided by the present application;
[0081] Figure 3 It is a flowchart of a process for obtaining a Mel spectrum provided by the present application; <0000A schematic diagram of a training process of a prosody coding predictor provided in this application;
[0085] Figure 7 A flowchart of a training process of a spectrum predictor provided in this application;
[0086] Figure 8 A schematic diagram of the training phase of a spectrum predictor provided in this application;
[0087] Figure 9 A schematic diagram of a speech signal synthesis framework provided in this application;
[0088] Figure 10 A schematic diagram of the structure of a speech signal generating device provided in this application;
[0089] Figure 11 This is a schematic diagram of the structure of an electronic device provided in this application. DETAILED DESCRIPTION
[0090] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0091] In this application, the terms "comprises," "comprising," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not preclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.
[0092] The embodiment of the present application provides a method for generating a speech signal, which can be applied to electronic devices, such as computers, smart phones, tablet devices, smart wearable devices, server clusters or cloud computing centers. The method flow chart of the method is as follows: Figure 1 As shown, specifically including:
[0093] S101: Obtain a target text to be processed; the target text includes N sentence texts and a narration dialogue tag for each sentence text, each narration dialogue tag is used to indicate the type of the sentence text to which it belongs, and the type of the sentence text is one of a narration type and a dialogue type.
[0094] Optionally, N = 2K + 1, where K is a positive integer.
[0095] S102: Based on a pre-trained prosody encoding predictor, the target text, the text features of the target text, and the timbre embedding information of each sentence text, obtain the prosody information of the central sentence text; the timbre embedding information of each sentence text is generated based on the narrator dialogue label and the reference speech of each sentence text; the central sentence text is the Mth sentence text in the target text, where 1 < M < N. In some embodiments, M = K + 1.
[0096] In this embodiment, as Figure 2 shown, the prosody encoding predictor may include a first BERT model, a first encoder, and a first decoder; the target text may be input into the first BERT model to obtain the first semantic features of the target text; then, the first semantic features of the target text are upsampled to obtain the upsampled first semantic features; the upsampled first semantic features are concatenated with the text features of the target text to obtain the first target text features; then, the first target text features are input into the first encoder to obtain the output of the first encoder; obtain the timbre encoder corresponding to the prosody encoding predictor; according to the timbre encoder corresponding to the prosody encoding predictor, the narrator dialogue label of each sentence text, and the reference speech information of each sentence text, obtain the first timbre embedding information of each sentence text; concatenate the first timbre embedding information of each sentence text with the output of the first encoder, and input the concatenated data into the first decoder to obtain the prosody information of the target text; extract the prosody information of the central sentence text from the prosody information of the target text, where the reference speech information of each sentence text is one of multiple selectable preset reference speech information of the sentence text, that is, the reference speech information of the sentence text can be automatically selected or selected by the user.
[0097] S103: Based on a pre-trained duration predictor, the target text, the text features of the target text, and the timbre embedding information of each sentence text, obtain the duration information of the central sentence text.
[0098] In this embodiment, the duration predictor has a similar structure to the prosody coding predictor, and the duration predictor may include a second BERT model, a second encoder, and a second decoder; the target text may be input into the second BERT model to obtain the second semantic features of the target text; the second semantic features of the target text may then be upsampled to obtain the upsampled second semantic features; the upsampled second semantic features may be concatenated with the text features of the target text to obtain the second target text features; the second target text features may then be input into the second encoder to obtain the output of the second encoder; a timbre encoder corresponding to the duration predictor may be obtained; second timbre embedding information of each sentence text may be obtained based on the timbre encoder corresponding to the duration predictor, the narration dialogue tag of each sentence text, and the reference voice information of each sentence text; the second timbre embedding information of each sentence text may be concatenated with the output of the second encoder, and the concatenated data may be input into the second decoder to obtain the duration information of the target text; the duration information of the central sentence text may be extracted from the duration information of the target text.
[0099] S104: Obtaining a speech signal of the central sentence text based on a pre-trained spectrum predictor, text features of the central sentence text, timbre embedding information of the central sentence text, duration information of the central sentence text, and prosody information of the central sentence text.
[0100] In this embodiment, the spectrum predictor includes a third encoder, a length regularizer and a third decoder. The text features of the central sentence text can be input into the third encoder to obtain the output of the third encoder; the timbre encoder corresponding to the spectrum predictor is obtained; the timbre embedding information of the central sentence text is obtained based on the timbre encoder corresponding to the spectrum predictor, the narration dialogue tag of the central sentence text, and the reference voice information of the central sentence text; the timbre embedding information, text features, rhythm information and duration information of the central sentence text are input into the length regularizer to obtain the output of the length regularizer, and the third decoder is used to obtain the speech signal of the central sentence text based on the output of the length regularizer, which can be a mel spectrum.
[0101] Optionally, the timbre encoders corresponding to the prosody coding predictor, duration predictor, and spectrum predictor may be the same or different, and may be used to generate timbre embedding information of the sentence text based on the narration dialogue tags and voice information of the sentence text.
[0102] In some embodiments, the prosody code predictor, duration predictor, and spectrum predictor each correspond to a different timbre coder, see Figure 3The spectrum predictor generates mel-spectrograms sentence by sentence, while the prosody code predictor and duration predictor generate the prosody codes and durations of 2K+1 sentences at one time. When computing resources are limited, prosody codes and durations can be generated for 2K+1 sentences first, and then speech can be synthesized for each sentence separately. Specifically, for each sentence text to be synthesized (i.e., the central sentence text), the prosody code predictor can be used to convert the text information (target text) of the 2K+1 sentences centered on the central sentence text into corresponding prosody codes. Then, the prosody code corresponding to the central sentence can be extracted based on the sentence boundaries. Similarly, the duration predictor predicts the duration of the 2K+1 sentences and extracts the duration of the central sentence. Finally, the spectrum predictor converts the text features of the central sentence into mel-spectrograms in combination with the prosody code and duration. While implementing the above process, this embodiment uses reference waveforms and narration dialogue tags as inputs to the timbre encoder to help synthesize conversational speech with various timbres.
[0103] See also Figure 4 , which is a flowchart of the process of outputting timbre embedding information from the timbre encoder. When the narration dialogue tag of a sentence text indicates that the sentence text is of the dialogue type, the timbre embedding information includes an intermediate timbre embedding and a narration dialogue embedding; when the narration dialogue tag of a sentence text indicates that the sentence text is of the narration type, the timbre embedding information includes an all-zero embedding and a narration dialogue embedding. Specifically, during the data preprocessing and annotation stage, sentences are divided into narration sentences and dialogue sentences and labeled with narration dialogue tags. In the timbre encoder, these narration dialogue tags are first converted into trainable narration dialogue embeddings using a lookup table, which form part of the output. This embedding informs the model whether the currently synthesized sentence is a narration sentence or a dialogue sentence, making it easier for the model to capture the phonetic differences between narration and dialogue sentences in audiobooks. To ensure the stability of dialogue sentence training and basic control of the timbre of the dialogue portion of the speech, this embodiment introduces a timbre encoder based on the ECAPA-TDNN model. This ECAPA-TDNN model is pre-trained for speaker recognition tasks on a large-scale multi-speaker dataset.
[0104] Optionally, during the training phase, the timbre encoder first determines the type of each sentence using the narration dialogue label. For narration sentences, it is assumed that the speaker uses a fixed timbre, so an all-zero embedding plus the narration dialogue embedding is used as the output. For dialogue sentences, the real speech is input as the reference speech into the ECAPA-TDNN to extract the sentence-level timbre representation, which is then converted into an intermediate timbre embedding through five fully connected layers and finally added to the narration dialogue embedding to obtain the output timbre embedding. The ECAPA-TDNN modules in these timbre encoders will be fine-tuned to learn more detailed timbre control. The timbre control in this embodiment is completely driven by the real speech waveform during the training phase, so there is no need to label the characters.
[0105] Optionally, before synthesizing speech, it is necessary to manually select reference waveforms of different timbres from the training set and construct a character timbre library. In the timbre encoder, the narration dialogue embedding is first obtained based on the narration dialogue label of the sentence. Then, an all-zero embedding is used for the narration sentence; for the dialogue sentence, the corresponding speech is automatically or manually selected from the character timbre library according to the specific character as the reference speech and input into the timbre encoder to obtain the corresponding timbre embedding. In addition, additional control can be applied to the intermediate timbre embedding within the timbre encoder. If the timbre change of the dialogue part is not required, the all-zero embedding can be used as the intermediate timbre embedding to synthesize the dialogue sentence. In addition, the degree of timbre imitation in the synthesized dialogue can be controlled by multiplying the intermediate timbre embedding by a coefficient.
[0106] The method provided in the embodiments of the present application can improve the rhythmic performance of speech synthesis and can realize the synthesis of speech with diverse timbres.
[0107] In an embodiment provided by the present application, based on the above implementation process, optionally, the training process of the prosody coding predictor is as follows: Figure 5 As shown, including:
[0108] S501: Obtain a generative adversarial network and a first training data set; the generative adversarial network includes a generator and a discriminator; the generator includes an initial prosody coding predictor and a first timbre encoder corresponding to the initial prosody coding; the first training data set includes a plurality of first training data, each of the first training data includes N first training sentence texts, and text features, narration dialogue tags, pre-extracted prosody coding and voice information of the N first training sentence texts.
[0109] In this embodiment, see Figure 6The discriminator of the adversarial generative network consists of four decoding blocks DBlock. Each DBlock contains a prosodic side block and a conditional side block, both of which contain residual connections and two layers of one-dimensional convolution. Taking into account the significant correlation between prosodic coding and the corresponding text, this embodiment adds an independent trainable conditional network to enhance the discriminator's ability to distinguish between true and false samples. The structure of this conditional network is consistent with that of the prosodic coding predictor. It includes an encoder, a decoder, and a pre-trained BERT network, and is also equipped with a dedicated timbre encoder. The output of the conditional network is directly connected to the conditional side input of the first Dblock. The conditional network participates in generative adversarial training as part of the discriminator. The training of the adversarial generative network in this embodiment uses the training criterion of least squares GAN, and uses feature matching loss and reconstruction loss.
[0110] In this embodiment, the prosody code may be pre-extracted by a trained prosody extractor.
[0111] S502: Select first target training data currently used for training from each first training data in the first training data set.
[0112] In this embodiment, the first target training data currently used for training may be selected in a preset order, or may be selected randomly.
[0113] S503: Obtaining timbre embedding information of the first target training data according to the first timbre encoder, the narration dialogue tag of the first target training data, and the voice information.
[0114] S504: Obtain prosodic codes of the N first training sentence texts in the first target training data according to the initial prosodic code predictor, the N first training sentence texts in the first target training data, and text features.
[0115] S505: Obtain loss function values of the generator and the discriminator according to the discriminator, the first target training data, and the prosodic encoding of N first training sentence texts in the first target training data.
[0116] S506: Using the loss function value of the generator to update the model parameters of the first timbre encoder and the model parameters of the initial prosody coding predictor in the generator; and using the loss function value of the discriminator to update the loss function value of the discriminator.
[0117] S507: Determine whether the updated generator meets the preset first training completion condition; if not, return to execute S502; if so, execute S508.
[0118] S508: Determine the updated initial prosody coding predictor in the generator as the trained prosody coding predictor.
[0119] In this embodiment, the first timbre encoder that meets the first training completion condition may be used as the timbre encoder corresponding to the prosody coding predictor.
[0120] In some embodiments provided in this application, based on the above implementation process, optionally, the training process of the spectrum predictor is as follows: Figure 7 Shown, including:
[0121] S701: Obtain an initial spectrum predictor, a second timbre encoder, a prosody code extractor, and a second training data set to be trained; the second training data set includes a plurality of second training data, each second training data includes a second training sentence text, and prosodic features, text features, voice information, narration dialogue tags, and a target signal of the second training sentence text; the prosodic features include duration, fundamental frequency, and energy.
[0122] S702: Select second target training data currently used for training from each second training data in the second training data set.
[0123] S703: Input the prosodic features in the second target training data into the prosodic code extractor to obtain the prosodic code of the second training sentence text in the second target training data.
[0124] S704: Obtain timbre embedding information of a second training sentence text in the second target training data according to the second timbre encoder, the narration dialogue tag of the second target training data, and the voice information.
[0125] S705: Obtain the speech signal of the second training sentence text in the second target training data based on the initial spectrum predictor, the text features in the second target training data, the duration in the prosodic features, the prosodic coding of the second training sentence text, and the timbre embedding information.
[0126] Optionally, the speech signal is a Mel spectrum.
[0127] S706: Calculate a first loss function value using a preset first loss function, a speech signal of a second training sentence text in the second target training data, and a target signal.
[0128] S707: Using the first loss function value, update the model parameters of the initial spectrum predictor, the second timbre encoder, and the prosody code extractor.
[0129] S708: Determine whether the updated initial spectrum predictor, second timbre encoder, and prosody code extractor meet a preset second training completion condition; if not, return to S702; if so, execute S709.
[0130] S709: Determine the initial spectrum predictor, second timbre encoder, and prosody code extractor that meet the second training completion condition as the trained spectrum predictor, second timbre encoder, and prosody code extractor.
[0131] In this embodiment, see Figure 8 , is a schematic diagram of the training stage of the spectrum predictor, which can be trained jointly with the prosody code extractor. In order to save computing resources, the second training data can be used for training. The second training data is the training data of a single sentence, not the paragraph-level data. The spectrum predictor consists of an encoder and a decoder, both of which contain a 4-layer Conformer structure. The phoneme-level text features are input into the encoder and converted into phoneme-level hidden layer representations. The length regularizer then expands the phoneme-level hidden layer representations to the frame level according to the duration, and then the decoder converts them into Mel-spectrogram output. The Mel-spectrogram is reconstructed using the first loss function L1, which also guides the training of the prosody code extractor and the timbre encoder.
[0132] In this embodiment, prosodic features extracted from the waveform, such as duration, fundamental frequency, and energy, are used as input to the prosodic code extractor, allowing the extractor to focus on capturing the prosodic changes in speech. Fenfen and energy are at the frame level, while duration operates at the phoneme level.
[0133] In this embodiment, the prosody code extractor consists of six 1D convolutional layers, two bidirectional LSTM layers, and two fully connected layers. The fundamental frequency and energy are first concatenated and fed into the prosody code extractor. After the first bidirectional LSTM layer, the hidden layer representation is averaged from the frame level to the phoneme level based on duration, then concatenated with the corresponding phoneme duration value before being fed into the subsequent network. The output of the prosody extractor is defined as the prosody code and concatenated with the output of the encoder in the spectrum predictor as conditioning information.
[0134] In this embodiment, the prosody extractor and the spectrum predictor may be trained first, and then the prosody code may be extracted based on the trained prosody extractor to train the prosody code predictor.
[0135] In this embodiment, the trained second timbre encoder may be used as the timbre encoder corresponding to the spectrum predictor.
[0136] In an embodiment provided in the present application, based on the above implementation process, optionally, the training process of the duration predictor includes:
[0137] Obtaining an initial duration predictor to be trained, a third timbre encoder, and a third training data set; the third training data set includes a plurality of third training data; each of the third training data includes N third training sentence texts, and text features, narration dialogue labels, voice information, and target duration information of the N third training sentence texts;
[0138] Selecting third target training data currently used for training from each third training data in the third training data set;
[0139] Obtaining timbre embedding information of N third training sentence texts in the third target training data according to the third timbre encoder, the narration dialogue tag of the third target training data, and the voice information;
[0140] Obtaining duration information of the N third training sentence texts in the third target training data according to the initial duration predictor, the N third training sentence texts in the third target training data, and text features;
[0141] Calculate a second loss function value using a preset second loss function, duration information of the N third training sentence texts in the third target training data, and target duration information;
[0142] Updating the model parameters of the initial duration predictor and the third timbre encoder using the second loss function value;
[0143] If the updated initial duration predictor does not meet the preset third training completion condition, returning to the step of selecting third target training data currently used for training from each third training data in the third training data set;
[0144] In the case that the updated initial duration predictor meets the preset third training completion condition, the initial duration predictor meeting the third training completion condition is used as the trained duration predictor.
[0145] In this embodiment, the trained third timbre encoder may be used as the timbre encoder corresponding to the duration predictor.
[0146] In some embodiments provided herein, based on the above implementation process, optionally, obtaining the speech signal of the central sentence text based on a pre-trained spectrum predictor, text features of the central sentence text, timbre embedding information of the central sentence text, duration information of the central sentence text, and prosody information of the central sentence text includes:
[0147] Obtaining a timbre encoder corresponding to the spectrum predictor;
[0148] Obtaining timbre embedding information of the central sentence text based on the timbre encoder corresponding to the spectrum predictor, the narration dialogue tag of each sentence text, and the reference voice information of each sentence text;
[0149] Multiplying the timbre embedding information of the central sentence text by a preset control coefficient to obtain target timbre embedding information;
[0150] Based on a pre-trained spectrum predictor, the text features of the central sentence text, the target timbre embedding information of the central sentence text, the duration information of the central sentence text and the prosody information of the central sentence text, the speech signal of the central sentence text is obtained.
[0151] See also Figure 9 , a schematic diagram of a speech signal synthesis framework provided in an embodiment of the present application; the speech signal synthesis framework is primarily composed of a duration predictor, a spectrum predictor, and a prosody coder. The duration predictor, spectrum predictor, and prosody coder can all employ arbitrary neural network structures. Furthermore, each predictor is equipped with a dedicated timbre encoder.
[0152] In this embodiment, an autoencoder-based method can be used to extract and predict prosodic codes to model the prosody of novel speech. During the training phase, an additional prosodic code extractor is first required to encode the actual acoustic features into prosodic codes. The prosodic codes are then input into the spectrum predictor as conditional information. The prosodic code extractor and the spectrum predictor are trained using a unified acoustic feature reconstruction loss function. After training, the prosodic code extractor is used to extract the prosodic codes corresponding to all sentences in the training set and use these as the actual prosodic codes.
[0153] Optionally, the prosody code predictor can predict the actual prosody code from text features. The training principle uses a generative adversarial network architecture. The generator is a prosody code predictor that predicts the prosody code from the text. The discriminator network is used to distinguish between the predicted prosody code and the actual prosody code. Because prosody codes are highly correlated with textual information, a separately trainable conditional network module is added to the discriminator to extract conditional information from text features to help the discriminator distinguish between predicted and actual samples. After training, the prosody code predictor can be used in the generation phase to predict the prosody code and provide it to the spectrum predictor.
[0154] In this embodiment, a timbre encoder can be used to provide timbre information to the conditional networks within the prosody code predictor, spectrum predictor, and discriminator within the main framework, thereby controlling the timbre of the novel's dialogue. The timbre encoder primarily consists of a speaker recognition network and is pre-trained on speaker recognition tasks, thus enabling it to distinguish different types of timbre from speech. During the training phase, real speech is input as reference speech into each timbre encoder, where timbre information is extracted and incorporated as a condition in each subsequent module. The parameters of the speaker recognition network within the timbre encoder are fine-tuned through joint training with subsequent modules to capture more detailed timbre information from the novel. During the generation phase, representative dialogue audio corresponding to several timbres can be pre-selected from the training set to construct a character timbre library. For the dialogue sentences to be synthesized, audio corresponding to the desired timbre can be automatically or manually selected from the timbre library as reference speech and input into each timbre encoder to synthesize dialogue speech of a specific timbre.
[0155] and Figure 1 Corresponding to the method described above, the embodiment of the present application further provides a speech signal generating device for Figure 1 The specific implementation of the method is shown in the following diagram: Figure 10 As shown, specifically including:
[0156] An acquisition unit 1001 is configured to acquire a target text to be processed; the target text includes N sentence texts and a narration dialogue tag for each sentence text, wherein each narration dialogue tag is used to indicate the type of the sentence text to which it belongs, and the type of the sentence text is one of a narration type and a dialogue type;
[0157] The first execution unit 1002 is configured to obtain the prosody information of the central sentence text based on a pre-trained prosody coding predictor, the target text, the text features of the target text, and the timbre embedding information of each sentence text; the timbre embedding information of each sentence text is generated based on the narration dialogue label of each sentence text and the reference speech; the central sentence text is the Mth sentence text in the target text, wherein 1 <M<N;
[0158] The second execution unit 1003 is configured to obtain the duration information of the central sentence text based on the pre-trained duration predictor, the target text, the text features of the target text, and the timbre embedding information of each sentence text;
[0159] The third execution unit 1004 is used to obtain the speech signal of the central sentence text based on a pre-trained spectrum predictor, the text features of the central sentence text, the timbre embedding information of the central sentence text, the duration information of the central sentence text and the rhythm information of the central sentence text.
[0160] In an embodiment provided in the present application, based on the above solution, optionally, the first execution unit 1002 includes:
[0161] A first acquisition subunit is configured to acquire a generative adversarial network and a first training dataset; the generative adversarial network includes a generator and a discriminator; the generator includes an initial prosody code predictor and a first timbre encoder corresponding to the initial prosody code; the first training dataset includes a plurality of first training data, each of the first training data includes N first training sentence texts, as well as text features, pre-extracted prosody codes, narration dialogue tags, and voice information of the N first training sentence texts;
[0162] A first selection subunit is configured to select first target training data currently used for training from each first training data in the first training data set;
[0163] a first execution subunit, configured to obtain timbre embedding information of the first target training data according to the first timbre encoder, the narration dialogue tag of the first target training data, and the voice information;
[0164] A second execution subunit is configured to obtain prosodic codes of the N first training sentence texts in the first target training data according to the initial prosodic code predictor, the N first training sentence texts in the first target training data, and text features;
[0165] A first calculation subunit is configured to obtain a loss function value of the generator based on the discriminator, the first target training data, and the prosodic encoding of N first training sentence texts in the first target training data;
[0166] A first updating subunit, configured to update the model parameters of the first timbre encoder and the model parameters of the initial prosody coding predictor in the generator using the loss function value of the generator;
[0167] a third execution subunit, configured to, if the updated generator does not satisfy the preset first training completion condition, return to executing the step of selecting first target training data currently used for training from each training data in the training data set;
[0168] The first determining subunit is configured to determine the initial prosody coding predictor in the updated generator as the trained prosody coding predictor if the updated generator does not meet a preset first training completion condition.
[0169] In an embodiment provided in the present application, based on the above solution, optionally, the third execution unit 1004 includes:
[0170] a second acquisition subunit, configured to acquire an initial spectrum predictor to be trained, a second timbre encoder, a prosody code extractor, and a second training data set; the second training data set includes a plurality of second training data, each second training data including a second training sentence text, and prosodic features, text features, voice information, narration dialogue tags, and a target signal of the second training sentence text; the prosodic features include duration, fundamental frequency, and energy;
[0171] A second selection subunit is used to select second target training data currently used for training from each second training data in the second training data set;
[0172] A first input subunit, configured to input the prosodic features in the second target training data into the prosodic code extractor to obtain a prosodic code of a second training sentence text in the second target training data;
[0173] a fourth execution subunit, configured to obtain timbre embedding information of a second training sentence text in the second target training data according to the second timbre encoder, the narration dialogue tag of the second target training data, and the voice information;
[0174] a fifth execution subunit, configured to obtain a speech signal of a second training sentence text in the second target training data based on the initial spectrum predictor, the text features in the second target training data, the duration in the prosodic features, the prosodic coding of the second training sentence text, and the timbre embedding information;
[0175] A second calculation subunit is configured to calculate a first loss function value by using a preset first loss function, a speech signal of a second training sentence text in the second target training data, and a target signal;
[0176] a second updating subunit, configured to update model parameters of the initial spectrum predictor, the second timbre encoder, and the prosody code extractor using the first loss function value;
[0177] a sixth execution subunit, configured to, if the updated initial spectrum predictor does not meet the preset second training completion condition, return to the step of selecting the second target training data currently used for training from each second training data in the second training data set;
[0178] The second determining subunit is configured to, when the updated initial spectrum predictor satisfies a preset second training completion condition, determine the initial spectrum predictor that satisfies the second training completion condition as a trained spectrum predictor.
[0179] In an embodiment provided in the present application, based on the above solution, optionally, the second execution unit 1003 includes:
[0180] a third acquisition subunit, configured to acquire an initial duration predictor to be trained, a third timbre encoder, and a third training data set; the third training data set comprising a plurality of third training data; each of the third training data comprising N third training sentence texts, and text features, narration dialogue labels, voice information, and target duration information of the N third training sentence texts;
[0181] A third selection subunit is used to select third target training data currently used for training from each third training data in the third training data set;
[0182] a seventh execution subunit, configured to obtain timbre embedding information of N third training sentence texts in the third target training data according to the third timbre encoder, the narration dialogue tag of the third target training data, and the voice information;
[0183] an eighth execution subunit, configured to obtain duration information of the N third training sentence texts in the third target training data according to the initial duration predictor, the N third training sentence texts in the third target training data, and text features;
[0184] A third calculation subunit is configured to calculate a second loss function value by using a preset second loss function, duration information of the N third training sentence texts in the third target training data, and target duration information;
[0185] a third updating subunit, configured to update the model parameters of the initial duration predictor and the third timbre encoder using the second loss function value;
[0186] a ninth execution subunit, configured to, if the updated initial duration predictor does not satisfy the preset third training completion condition, return to trigger the third selection subunit to execute a process of selecting third target training data currently used for training from each third training data in the third training data set;
[0187] The tenth execution subunit is configured to, when the updated initial duration predictor satisfies a preset third training completion condition, use the initial duration predictor that satisfies the third training completion condition as the trained duration predictor.
[0188] The specific principles and execution processes of each unit and module in the voice signal generating device disclosed in the above embodiment of the present application are the same as those of the voice signal generating method disclosed in the above embodiment of the present application. Please refer to the corresponding parts of the voice signal generating method provided in the above embodiment of the present application, and no further details will be given here.
[0189] An embodiment of the present application further provides a storage medium, which includes stored instructions, wherein when the instructions are executed, the device where the storage medium is located is controlled to execute the above-mentioned voice signal generation method.
[0190] The present application also provides an electronic device, the structure of which is shown in FIG. Figure 11 As shown, it specifically includes a memory 1101 and one or more instructions 1102, wherein the one or more instructions 1102 are stored in the memory 1101 and are configured to be executed by one or more processors 1103 to perform the following operations:
[0191] Obtaining a target text to be processed; the target text includes N sentence texts and a narration dialogue tag for each sentence text, each narration dialogue tag is used to indicate the type of the sentence text to which it belongs, and the type of the sentence text is one of a narration type and a dialogue type;
[0192] Based on a pre-trained prosodic coding predictor, the target text, the text features of the target text, and the timbre embedding information of each sentence text, the prosodic information of the central sentence text is obtained; the timbre embedding information of each sentence text is generated based on the narration dialogue label and the reference speech of each sentence text; the central sentence text is the Mth sentence text in the target text, wherein 1 <M<N;
[0193] Obtaining duration information of the central sentence text based on a pre-trained duration predictor, the target text, text features of the target text, and timbre embedding information of each sentence text;
[0194] Based on a pre-trained spectrum predictor, the text features of the central sentence text, the timbre embedding information of the central sentence text, the duration information of the central sentence text and the prosody information of the central sentence text, the speech signal of the central sentence text is obtained.
[0195] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similarities between the various embodiments can be referred to in conjunction with each other. For device embodiments, since they are generally similar to method embodiments, their description is relatively simple, and for relevant details, reference can be made to the description of the method embodiments.
[0196] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are merely used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations.
[0197] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing this application, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0198] Through the description of the above embodiments, it can be seen that those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present application or certain parts of the embodiments.
[0199] The above is a detailed introduction to a speech signal generation method provided by the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for general technical personnel in this field, based on the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A method for generating a speech signal, characterized in that: include: Get the target text to be processed; The target text includes N sentence texts and a narration dialogue tag for each sentence text, wherein each narration dialogue tag is used to indicate the type of the sentence text to which it belongs, and the type of the sentence text is one of a narration type and a dialogue type; Based on a pre-trained prosodic coding predictor, the target text, the text features of the target text, and the timbre embedding information of each sentence text, the prosodic information of the central sentence text is obtained; the timbre embedding information of each sentence text is generated based on the narration dialogue label and the reference speech of each sentence text; the central sentence text is the Mth sentence text in the target text, wherein 1 <M<N; Obtaining duration information of the central sentence text based on a pre-trained duration predictor, the target text, text features of the target text, and timbre embedding information of each sentence text; Based on a pre-trained spectrum predictor, the text features of the central sentence text, the timbre embedding information of the central sentence text, the duration information of the central sentence text and the prosody information of the central sentence text, the speech signal of the central sentence text is obtained.
2. The method according to claim 1, characterized in that The training process of the prosody coding predictor includes: Obtaining a generative adversarial network and a first training dataset; the generative adversarial network includes a generator and a discriminator; the generator includes an initial prosody code predictor and a first timbre encoder corresponding to the initial prosody code; the first training dataset includes a plurality of first training data, each of the first training data includes N first training sentence texts, as well as text features, pre-extracted prosody codes, narration dialogue labels, and voice information of the N first training sentence texts; Selecting first target training data currently used for training from each first training data in the first training data set; Obtaining timbre embedding information of the first target training data according to the first timbre encoder, the narration dialogue tag of the first target training data, and the voice information; Obtaining prosodic codes of the N first training sentence texts in the first target training data according to the initial prosodic code predictor, the N first training sentence texts in the first target training data, and text features; Obtaining loss function values of the generator and the discriminator according to the generator, the discriminator, the first target training data, and the prosodic encoding of N first training sentence texts in the first target training data; Using the loss function value of the generator to update the model parameters of the first timbre encoder and the model parameters of the initial prosody coding predictor in the generator; and using the loss function value of the discriminator to update the model parameters of the discriminator; If the updated generator and the discriminator do not meet the preset first training completion condition, returning to the step of selecting the first target training data currently used for training from each training data in the training data set; When the updated generator and the discriminator meet a preset first training completion condition, the initial prosody coding predictor in the updated generator is determined as the trained prosody coding predictor.
3. The method according to claim 1, characterized in that The training process of the spectrum predictor includes: Obtaining an initial spectrum predictor to be trained, a second timbre encoder, a prosody code extractor, and a second training data set; the second training data set includes a plurality of second training data, each second training data including a second training sentence text, and prosodic features, text features, voice information, narration dialogue labels, and a target signal of the second training sentence text; the prosodic features include duration, fundamental frequency, and energy; Selecting second target training data currently used for training from each second training data in the second training data set; Inputting the prosodic features in the second target training data into the prosodic code extractor to obtain the prosodic code of the second training sentence text in the second target training data; Obtaining timbre embedding information of a second training sentence text in the second target training data according to the second timbre encoder, the narration dialogue tag of the second target training data, and the voice information; Obtaining a speech signal of a second training sentence text in the second target training data according to the initial spectrum predictor, the text features in the second target training data, the duration in the prosodic features, the prosodic coding of the second training sentence text, and the timbre embedding information; Calculating a first loss function value using a preset first loss function, a speech signal of a second training sentence text in the second target training data, and a target signal; Updating model parameters of the initial spectrum predictor, the second timbre encoder, and the prosody code extractor using the first loss function value; If the updated initial spectrum predictor does not meet the preset second training completion condition, returning to the step of selecting the second target training data currently used for training from each second training data in the second training data set; In a case where the updated initial spectrum predictor meets a preset second training completion condition, the initial spectrum predictor meeting the second training completion condition is determined as a trained spectrum predictor.
4. The method according to claim 1, wherein The training process of the duration predictor includes: Obtaining an initial duration predictor to be trained, a third timbre encoder, and a third training data set; the third training data set includes a plurality of third training data; each of the third training data includes N third training sentence texts, and text features, narration dialogue labels, voice information, and target duration information of the N third training sentence texts; Selecting third target training data currently used for training from each third training data in the third training data set; Obtaining timbre embedding information of N third training sentence texts in the third target training data according to the third timbre encoder, the narration dialogue tag of the third target training data, and the voice information; Obtaining duration information of the N third training sentence texts in the third target training data according to the initial duration predictor, the N third training sentence texts in the third target training data, and text features; Calculate a second loss function value using a preset second loss function, duration information of the N third training sentence texts in the third target training data, and target duration information; Updating the model parameters of the initial duration predictor and the third timbre encoder using the second loss function value; If the updated initial duration predictor does not meet the preset third training completion condition, returning to the step of selecting third target training data currently used for training from each third training data in the third training data set; In the case that the updated initial duration predictor meets the preset third training completion condition, the initial duration predictor meeting the third training completion condition is used as the trained duration predictor.
5. The method according to claim 1, wherein The method of obtaining a speech signal of the central sentence text based on a pre-trained spectrum predictor, text features of the central sentence text, timbre embedding information of the central sentence text, duration information of the central sentence text, and prosody information of the central sentence text includes: Obtaining a timbre encoder corresponding to the spectrum predictor; Obtaining timbre embedding information of the central sentence text based on the timbre encoder corresponding to the spectrum predictor, the narration dialogue tag of each sentence text, and the reference voice information of each sentence text; Multiplying the timbre embedding information of the central sentence text by a preset control coefficient to obtain target timbre embedding information; Based on a pre-trained spectrum predictor, the text features of the central sentence text, the target timbre embedding information of the central sentence text, the duration information of the central sentence text and the prosody information of the central sentence text, the speech signal of the central sentence text is obtained.
6. A speech signal generating device, characterized in that: include: An acquisition unit, used for acquiring target text to be processed; The target text includes N sentence texts and a narration dialogue tag for each sentence text, wherein each narration dialogue tag is used to indicate the type of the sentence text to which it belongs, and the type of the sentence text is one of a narration type and a dialogue type; The first execution unit is configured to obtain the prosody information of the central sentence text based on a pre-trained prosody coding predictor, the target text, the text features of the target text, and the timbre embedding information of each sentence text; the timbre embedding information of each sentence text is generated based on the narration dialogue label of each sentence text and the reference speech; the central sentence text is the Mth sentence text in the target text, wherein 1 <M<N; A second execution unit is configured to obtain duration information of the central sentence text based on a pre-trained duration predictor, the target text, text features of the target text, and timbre embedding information of each sentence text; The third execution unit is used to obtain the speech signal of the central sentence text based on a pre-trained spectrum predictor, the text features of the central sentence text, the timbre embedding information of the central sentence text, the duration information of the central sentence text and the rhythm information of the central sentence text.
7. The device according to claim 6, characterized in that The first execution unit includes: A first acquisition subunit is configured to acquire a generative adversarial network and a first training dataset; the generative adversarial network includes a generator and a discriminator; the generator includes an initial prosody code predictor and a first timbre encoder corresponding to the initial prosody code; the first training dataset includes a plurality of first training data, each of the first training data includes N first training sentence texts, as well as text features, pre-extracted prosody codes, narration dialogue tags, and voice information of the N first training sentence texts; A first selection subunit is configured to select first target training data currently used for training from each first training data in the first training data set; a first execution subunit, configured to obtain timbre embedding information of the first target training data according to the first timbre encoder, the narration dialogue tag of the first target training data, and the voice information; A second execution subunit is configured to obtain prosodic codes of the N first training sentence texts in the first target training data according to the initial prosodic code predictor, the N first training sentence texts in the first target training data, and text features; A first calculation subunit is configured to obtain loss function values of the generator and the discriminator based on the generator, the discriminator, the first target training data, and the prosodic encoding of the N first training sentence texts in the first target training data; A first updating subunit is configured to update the model parameters of the first timbre encoder and the model parameters of the initial prosody coding predictor in the generator using the loss function value of the generator; and to update the model parameters of the discriminator using the loss function value of the discriminator; a third execution subunit, configured to, if the updated generator and discriminator do not meet the preset first training completion condition, return to trigger the first selection subunit to execute the step of selecting first target training data currently used for training from each training data in the training data set; The first determining subunit is configured to determine the initial prosody coding predictor in the updated generator as the trained prosody coding predictor when the updated generator and the discriminator meet a preset first training completion condition.
8. The device according to claim 6, characterized in that The third execution unit includes: a second acquisition subunit, configured to acquire an initial spectrum predictor to be trained, a second timbre encoder, a prosody code extractor, and a second training data set; the second training data set includes a plurality of second training data, each second training data including a second training sentence text, and prosodic features, text features, voice information, narration dialogue tags, and a target signal of the second training sentence text; the prosodic features include duration, fundamental frequency, and energy; A second selection subunit is used to select second target training data currently used for training from each second training data in the second training data set; A first input subunit, configured to input the prosodic features in the second target training data into the prosodic code extractor to obtain a prosodic code of a second training sentence text in the second target training data; a fourth execution subunit, configured to obtain timbre embedding information of a second training sentence text in the second target training data according to the second timbre encoder, the narration dialogue tag of the second target training data, and the voice information; a fifth execution subunit, configured to obtain a speech signal of a second training sentence text in the second target training data based on the initial spectrum predictor, the text features in the second target training data, the duration in the prosodic features, the prosodic coding of the second training sentence text, and the timbre embedding information; A second calculation subunit is configured to calculate a first loss function value by using a preset first loss function, a speech signal of a second training sentence text in the second target training data, and a target signal; a second updating subunit, configured to update model parameters of the initial spectrum predictor, the second timbre encoder, and the prosody code extractor using the first loss function value; a sixth execution subunit, configured to, if the updated initial spectrum predictor does not meet the preset second training completion condition, return to trigger the second selection subunit to execute the step of selecting second target training data currently used for training from each second training data in the second training data set; The second determining subunit is configured to, when the updated initial spectrum predictor satisfies a preset second training completion condition, determine the initial spectrum predictor that satisfies the second training completion condition as a trained spectrum predictor.
9. The device according to claim 6, characterized in that The second execution unit includes: a third acquisition subunit, configured to acquire an initial duration predictor to be trained, a third timbre encoder, and a third training data set; the third training data set comprising a plurality of third training data; each of the third training data comprising N third training sentence texts, and text features, narration dialogue labels, voice information, and target duration information of the N third training sentence texts; A third selection subunit is used to select third target training data currently used for training from each third training data in the third training data set; a seventh execution subunit, configured to obtain timbre embedding information of N third training sentence texts in the third target training data according to the third timbre encoder, the narration dialogue tag of the third target training data, and the voice information; an eighth execution subunit, configured to obtain duration information of the N third training sentence texts in the third target training data according to the initial duration predictor, the N third training sentence texts in the third target training data, and text features; A third calculation subunit is configured to calculate a second loss function value by using a preset second loss function, duration information of the N third training sentence texts in the third target training data, and target duration information; a third updating subunit, configured to update the model parameters of the initial duration predictor and the third timbre encoder using the second loss function value; a ninth execution subunit, configured to, if the updated initial duration predictor does not satisfy the preset third training completion condition, return to trigger the third selection subunit to execute a process of selecting third target training data currently used for training from each third training data in the third training data set; The tenth execution subunit is configured to, when the updated initial duration predictor satisfies a preset third training completion condition, use the initial duration predictor that satisfies the third training completion condition as the trained duration predictor.
10. An electronic device, characterized in that: The system comprises a memory and one or more instructions, wherein the one or more instructions are stored in the memory and configured to execute the speech signal generation method according to any one of claims 1 to 5 by one or more processors.
Citation Information
Patent Citations
Speech synthesis method and device, equipment and storage medium
CN113990286A
Speech synthesis method and related device, electronic equipment and storage medium
CN114283781A