Speech synthesis methods, devices, equipment and storage media

By introducing user auditory feedback signals into the speech synthesis model and optimizing the model parameters, the problem of existing speech libraries not matching user auditory perception is solved, and a speech synthesis effect that better matches user auditory perception is achieved.

CN116612742BActive Publication Date: 2026-03-31IFLYTEK CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-28
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

During the training process, existing speech synthesis systems may encounter pronunciation flaws and issues that do not conform to the user's auditory perception in the original speech in the audio library. This results in the synthesized speech not meeting the user's auditory perception goals in real-world scenarios, thus reducing the effectiveness of speech synthesis.

Method used

By pre-training the speech synthesis model and adding user auditory feedback signals as reward signals, the parameters of the basic speech synthesis model are adjusted to guide the model to optimize in a direction that is more in line with human hearing. The phoneme sequence is obtained by text analysis and input into the trained model to generate synthesized speech that is more in line with user hearing.

Benefits of technology

The speech synthesis effect has been improved, and the speech generated by the speech synthesis system is more in line with the user's auditory goals, thus improving the quality of speech synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116612742B_ABST
    Figure CN116612742B_ABST
Patent Text Reader

Abstract

This application discloses a speech synthesis method, apparatus, device, and storage medium. The method involves analyzing the original text to be synthesized to obtain a phoneme sequence; inputting the phoneme sequence into a configured speech synthesis model to obtain synthesized speech output by the model. The speech synthesis model is a final speech synthesis model after parameter adjustment of the basic speech synthesis model, using the scoring results of multiple candidate speech samples corresponding to the input test text synthesized by the basic speech synthesis model as reward signals. The scoring results of each candidate speech sample conform to the user's auditory perception goals. This application adds user auditory feedback signals (i.e., the scoring results as reward signals) to the training process of the speech synthesis model, guiding the speech synthesis model to optimize model parameters in a direction that better conforms to the user's auditory perception, making the synthesized speech more in line with the user's auditory perception goals and improving the speech synthesis effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech synthesis technology, and more specifically, to a speech synthesis method, apparatus, device, and storage medium. Background Technology

[0002] Humans communicate with each other in various ways in daily life, but the most direct, easy-to-understand, and natural mode of communication is voice. The rapid development of computer and internet technology has greatly changed people's lifestyles, and the relationship between humans and computers is inseparable. Today, speech synthesis is widely used in interactive fields such as smart homes and intelligent robots.

[0003] However, current speech synthesis systems typically train by reproducing the original speech from a sound library. For example, they train the system's model parameters by minimizing the error between the synthesized and original speech as the loss function. In this process, it's questionable whether the original speech in the sound library best matches human hearing. The original speech in the sound library generally only guarantees correctness—that is, the synthesized speech matches the text content—but it may still have pronunciation flaws, such as feedback, and may not meet the user's auditory expectations in terms of timbre and rhythm. Therefore, in real-world speech synthesis scenarios, the synthesized speech produced by existing systems trained to reproduce the original speech from the sound library may not meet the user's auditory expectations, thus reducing the effectiveness of speech synthesis. Summary of the Invention

[0004] In view of the above problems, this application is made to provide a speech synthesis method, apparatus, device, and storage medium, so as to achieve the goal of synthesizing speech that better meets the user's auditory experience and improving the speech synthesis effect. The specific solution is as follows:

[0005] Firstly, a speech synthesis method is provided, including:

[0006] Obtain the original text of the speech to be synthesized;

[0007] The original text is analyzed to obtain the phoneme sequence corresponding to the original text;

[0008] The phoneme sequence corresponding to the original text is input into the configured speech synthesis model to obtain the synthesized speech output by the model;

[0009] The speech synthesis model is a final speech synthesis model obtained by adjusting the parameters of the basic speech synthesis model, using the scoring results of multiple candidate speech synthesized by the basic speech synthesis model corresponding to the input test text as reward signals. The scoring results of each candidate speech meet the user's listening perception goals.

[0010] Secondly, a speech synthesis device is provided, comprising:

[0011] The raw text acquisition unit is used to acquire the raw text of the speech to be synthesized;

[0012] The text analysis unit is used to perform text analysis on the original text to obtain the phoneme sequence corresponding to the original text.

[0013] The speech synthesis model processing unit is used to input the phoneme sequence corresponding to the original text into the configured speech synthesis model to obtain the synthesized speech output by the model; the speech synthesis model is the final speech synthesis model after adjusting the parameters of the basic speech synthesis model by using the scoring results of multiple candidate speech synthesized by the basic speech synthesis model corresponding to the input test text as reward signals, wherein the scoring results of each candidate speech meet the user's listening perception target.

[0014] Thirdly, a speech synthesis device is provided, including: a memory and a processor;

[0015] The memory is used to store programs;

[0016] The processor is used to execute the program to implement the various steps of the speech synthesis method as described above.

[0017] Fourthly, a storage medium is provided on which a computer program is stored, which, when executed by a processor, implements the various steps of the speech synthesis method as described above.

[0018] Using the above technical solution, this application pre-trains a speech synthesis model, which is obtained by updating and adjusting the parameters of a basic speech synthesis model. The basic speech synthesis model can be obtained using traditional training methods. During the parameter update and adjustment of the basic speech synthesis model, the scoring results of multiple candidate speech samples synthesized by the basic speech synthesis model corresponding to the input test text can be obtained. These scoring results meet the user's auditory perception goals. Based on this, the scoring results are used as a reward signal to update and adjust the parameters of the basic speech synthesis model. As can be seen, this application adds user auditory perception feedback signals (i.e., the scoring results as reward signals) to the training process of the speech synthesis model, guiding the speech synthesis model to optimize and adjust its parameters in a direction that better matches human auditory perception. Based on this, for the original text to be synthesized, a phoneme sequence is first obtained through text analysis, and then the phoneme sequence is input into the trained speech synthesis model to obtain the synthesized speech output by the model. This synthesized speech better meets the user's auditory perception goals, thereby greatly improving the speech synthesis effect. Attached Figure Description

[0019] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:

[0020] Figure 1 This is a schematic flowchart of a speech synthesis method provided in an embodiment of this application;

[0021] Figure 2 An example is provided, illustrating the process of speech synthesis using a basic speech synthesis model with a given structure.

[0022] Figure 3 This example illustrates the process of speech synthesis using a base speech synthesis model with another structure.

[0023] Figure 4 An example is shown in the diagram illustrating the training process of a speech synthesis model;

[0024] Figure 5 A schematic diagram illustrating the training process of another speech synthesis model is provided.

[0025] Figure 6 This is a schematic diagram of a speech synthesis device provided in an embodiment of this application;

[0026] Figure 7 This is a schematic diagram of the structure of the speech synthesis device provided in the embodiments of this application. Detailed Implementation

[0027] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0028] This application provides a speech synthesis scheme applicable to various scenarios requiring speech synthesis, such as speech synthesis in mobile phones, in-vehicle voice assistants, smart homes, and smart robots. Furthermore, to ensure the effectiveness of the speech synthesis in this application, a novel training method for the speech synthesis model is provided, enabling the trained model to output speech that better meets the user's auditory goals, thereby improving the speech synthesis effect.

[0029] The proposed solution can be implemented based on a terminal with data processing capabilities, such as a mobile phone, computer, server, or cloud platform.

[0030] Next, combined Figure 1 The speech synthesis method of this application may include the following steps:

[0031] Step S100: Obtain the original text of the speech to be synthesized.

[0032] Specifically, in a speech synthesis scenario, the original text that needs to be synthesized is obtained.

[0033] Step S110: Perform text analysis on the original text to obtain the phoneme sequence corresponding to the original text.

[0034] Specifically, the purpose of text analysis is to convert the input raw text into a sequence of symbols that can be used for speech synthesis, that is, the phoneme sequence in this step.

[0035] The text analysis process may include the following sub-steps:

[0036] S1. Text Normalization: Converts special expressions such as numbers, abbreviations, and currency in the original text into standard text form. S2. Word Segmentation: Divides the standard text into basic units such as words and punctuation marks. S3. Prosody Prediction: Based on the prosodic features of the input original text predictor, such as L1, L2, L3, L4, and / or L5 prosodic features. In addition, it can include prediction of tone sandhi and polyphonic characters. S4. Phonetic Conversion: Converts each basic unit after word segmentation into a phoneme sequence, that is, uses phonemes to represent pronunciation, resulting in a phoneme sequence.

[0037] Step S120: Input the phoneme sequence corresponding to the original text into the configured speech synthesis model to obtain the synthesized speech output by the model. The speech synthesis model uses the score results of each candidate speech that meets the user's listening perception target as a reward signal to adjust the parameters of the basic speech synthesis model.

[0038] Specifically, in this embodiment of the application, a speech synthesis model can be pre-trained. The speech synthesis model can be a final speech synthesis model after adjusting the parameters of the basic speech synthesis model by using the scoring results of multiple candidate speech synthesized by the basic speech synthesis model corresponding to the input test text as a reward signal. The scoring results of each candidate speech meet the user's listening perception target.

[0039] The basic speech synthesis model can be trained using traditional methods, such as training it with the goal of making the synthesized speech of the test text approximate the original speech in the speech library corresponding to the test text.

[0040] The speech synthesis method provided in this application pre-trains a speech synthesis model, which is obtained by updating and adjusting the parameters of a basic speech synthesis model. The basic speech synthesis model can be obtained using traditional training methods. During the parameter update and adjustment of the basic speech synthesis model, the scoring results of multiple candidate speech samples synthesized by the basic speech synthesis model corresponding to the input test text can be obtained. These scoring results meet the user's auditory perception goals. Based on this, the scoring results are used as a reward signal to update and adjust the parameters of the basic speech synthesis model. As can be seen, this application adds a user auditory perception feedback signal (i.e., the scoring results as a reward signal) to the training process of the speech synthesis model, guiding the speech synthesis model to optimize and adjust its parameters in a direction that better matches human auditory perception. Furthermore, for the original text to be synthesized, a phoneme sequence is first obtained through text analysis, and then the phoneme sequence is input into the trained speech synthesis model to obtain the synthesized speech output by the model. This synthesized speech better meets the user's auditory perception goals, thereby greatly improving the speech synthesis effect.

[0041] In one optional embodiment, during the training phase of the aforementioned speech synthesis model, the scoring results of each candidate speech, serving as the reward signal, can be obtained in various ways. For example, users can manually score the candidate speech based on their listening preferences. In addition, this embodiment also provides a method for automatically scoring candidate speech using a machine learning model. Specifically:

[0042] In order to score candidate speech and ensure that the scoring results meet the user's listening perception goals, a reward model can be pre-trained in this embodiment to score each candidate speech.

[0043] The reward model can employ neural network structures such as CNN, RNN, LSTM, and Transformer. The training process for the reward model can involve using text-speech pairs containing the test text and a corresponding synthesized speech as training samples, and using the user ratings of the synthesized speech contained in the text-speech pairs as sample labels.

[0044] In other words, in this embodiment, test text and its corresponding synthesized speech can be obtained, test text and each synthesized speech can be combined to form a text-speech pair, and the user can score the synthesized speech in the text-speech pair. When scoring, the user can do so according to their own listening preferences. The score represents the relative quality of each synthesized speech as measured by the user from the listening perspective.

[0045] The input to the reward model during the training phase is the test text x and the corresponding synthesized speech y. i The output is a score of the synthesized speech in the current input text-speech pair:

[0046] s i =f θ (x,y i )

[0047] Among them, y i Let θ represent the i-th synthesized speech in the synthesized speech corresponding to the test text x. θ represents the model parameters of the reward model.

[0048] The loss function for training the reward model can be that, for the same test text, the higher the score of the synthesized speech and the test text combined, the higher the score of the text-speech pair. Specifically, it can be achieved by minimizing the following loss function:

[0049]

[0050] In the case of input test text x, y w It is more than y l The higher-scoring synthesized speech, where K represents the total number of synthesized speech samples corresponding to the test text x.

[0051] The reward model provided in this embodiment can be used to automatically score multiple candidate speech synthesized by the basic speech synthesis model corresponding to the input test text, and ensure that the score meets the user's listening perception target. The score result of the candidate speech is used as a reward signal to update the parameters of the basic speech synthesis model.

[0052] In some embodiments of this application, two optional structures for the basic speech synthesis model are provided.

[0053] Reference Figure 2 and Figure 3 The paper illustrates the speech synthesis process of two different basic speech synthesis models.

[0054] The first model structure, such as Figure 2 As shown, a basic speech synthesis model can include a cascaded acoustic model and a vocoder.

[0055] The acoustic model is used to derive acoustic features based on the input phoneme sequence, such as fundamental frequency (F0) and spectral features. These acoustic features are related to the characteristics of the vocal tract and vocal cords during human vocalization. The acoustic model can employ structures such as RNNs, LSTMs, or Transformers, and is responsible for learning the mapping relationship between phonemes and acoustic features.

[0056] A vocoder is used to convert acoustic features into audio signals. Specifically, the task of a vocoder can be to convert acoustic features into a continuous audio signal of speech samples. Common waveform generation algorithms include, but are not limited to, the Griffin-Lim algorithm, WaveNet, WaveRNN, and WaveGlow. These algorithms generate speech signals that approximate the original speech in the audio database through generator-discriminator architectures (e.g., GAN network structures) or autoregressive models.

[0057] As can be seen from the above description of the functions of the acoustic model and vocoder, the acoustic model and vocoder are modeled separately during the training process.

[0058] During the training phase, acoustic models can be trained by minimizing the error between acoustic features (such as speech spectrum features) and the acoustic features of the original speech as a loss function, such as MSE loss.

[0059] During the training phase, a generative-adversarial approach can be adopted for vocoders. That is, the goal of the generator is to make the generated audio signal as close as possible to the original speech signal, while the goal of the discriminator is to distinguish the generated audio signal from the original audio signal. When the discriminator cannot distinguish the speech generated by the generator from the original speech, it is considered that the speech generated by the generator is highly realistic and can be used as a vocoder after training.

[0060] Both the training phase of the acoustic model and the training phase of the vocoder aim to infinitely recover the original speech in the voice library.

[0061] The second model structure, such as Figure 2 As shown, the basic speech synthesis model can be modeled end-to-end. That is, the phoneme sequence obtained from text analysis can be directly input into the end-to-end basic speech synthesis model to obtain the synthesized speech output by the model.

[0062] The basic speech synthesis model for end-to-end modeling can adopt the generative-adversarial criterion during the training phase. The specific process is similar to that described above and will not be repeated here.

[0063] This embodiment provides the structures of two basic speech synthesis models. Both basic speech synthesis models are trained during the training phase with the goal of making the synthesized speech approximate the original speech in the sound library that corresponds to the test text.

[0064] Combination Figure 4 and Figure 5 This embodiment describes the training process of an optional speech synthesis model.

[0065] The speech synthesis model described in this embodiment is also applicable to the two different basic speech synthesis models with the aforementioned embodiments. The training process will be described in detail below, including the following steps:

[0066] S1. Obtain the basic speech synthesis model

[0067] Specifically, the basic speech synthesis model is trained with the goal of making the synthesized speech of the test text approximate the original speech in the sound library corresponding to the test text. The specific structure and training process can be referred to the relevant introduction above.

[0068] S2, Data Collection

[0069] Based on the basic speech synthesis model, multiple candidate speech samples are generated for each test text in the first test set through sampling and decoding, resulting in a candidate speech set for each test text.

[0070] The first test set may include several test texts. For each test text, multiple candidate speech can be generated by sampling and decoding using a basic speech synthesis model, thus obtaining a candidate speech set corresponding to each test text.

[0071] Sampling decoding methods include, but are not limited to, vocoders or end-to-end basic speech synthesis models. When generating audio sample points, a probability-based decoding method can be used. That is, when generating each audio sample point, the probability of its value occurring is considered, and sampling is performed based on this probability to determine the value of the audio sample point (for example, for sampling 16-bit encoded sample points, there are 256 possible values). This method allows the basic speech synthesis model to generate multiple candidate speech values ​​for a single test text. For example, for the test text x = "It's nice to meet you today," the basic speech synthesis model generates candidate speech values ​​y1, y2, y3, y4, and y5 based on the sampling decoding strategy.

[0072] S3, User Evaluation Feedback

[0073] Specifically, the user's rating of each candidate speech in the candidate speech set corresponding to each test text is obtained, resulting in a candidate speech set carrying the rating results for each test text.

[0074] In practical applications, human evaluators can be asked to score each candidate speech in the candidate speech set. The evaluators will give a comprehensive score to the candidate speech based on factors such as pronunciation, sound quality, rhythm and cadence.

[0075] S4, Training the reward model

[0076] Specifically, the reward model is trained using the set of candidate speech containing the scoring results corresponding to each test text as training data.

[0077] Specifically, the test text and each candidate speech in the corresponding candidate speech set can be combined into a text-speech pair as training samples, and the scores of the candidate speech in the text-speech pair can be used as sample labels to train the reward model. The specific training process of the reward model can be referred to the description of the relevant embodiments above, and will not be repeated here.

[0078] S5. Update the basic speech synthesis model.

[0079] Specifically, the trained reward model is used to score the multiple candidate speech synthesized by the basic speech synthesis model for each test text in the second test set, and the score results of each candidate speech are used as reward signals to adjust the parameters of the basic speech synthesis model to obtain the adjusted speech synthesis model.

[0080] like Figure 4 As shown, when the basic speech synthesis model is a structure where the acoustic model and the vocoder are modeled separately, in this step, the parameters of the vocoder in the basic speech synthesis model can be adjusted based on the reward signal to obtain the adjusted vocoder. The adjusted vocoder and the acoustic model together form the adjusted speech synthesis model.

[0081] like Figure 5 As shown, when the basic speech synthesis model is an end-to-end basic speech synthesis model, the parameters of the end-to-end basic speech synthesis model can be adjusted based on the reward signal in this step to obtain the adjusted speech synthesis model.

[0082] In a preferred implementation, the second test set in this step is preferably different from the first test set in step S2 above, so as to avoid the problem of overfitting when using the reward model on the same test text.

[0083] By employing steps S1-S5 of the above embodiments, the training process of the speech synthesis model can be completed. The speech synthesis model training method provided in this application breaks away from the traditional goal of speech synthesis models that aim to infinitely recover the original speech of the sound bank. Instead, it guides the speech synthesis model to optimize in a direction that is more in line with human hearing, thereby greatly improving the upper limit of the performance of the speech synthesis model.

[0084] Alternatively, to improve the training effect of the speech synthesis model, a multi-round iterative optimization approach can be used for model training.

[0085] Specifically, after completing one round of training in step S5 to obtain the adjusted speech synthesis model, the adjusted speech synthesis model can be updated to the basic speech synthesis model, and the training process of steps S2-S5 can be iteratively executed until the set iteration termination condition is met, and the final adjusted speech synthesis model is obtained.

[0086] The iteration termination condition can be set by the number of iterations reaching a set number of iteration rounds, or by the change in the model loss function during the reinforcement learning process being less than a set threshold.

[0087] It's important to explain that in each iteration of training, the base speech synthesis model used to generate candidate speech in step S2 is the same as the one adjusted in the previous round. This means the base speech synthesis model is different in each round, leading to different generated candidate speech. After human evaluation, the training reward model is updated. The evaluation results from the reward model are then used as reward signals to reinforce the base speech synthesis model. In each iteration, the performance of the base speech synthesis model gradually improves, ultimately resulting in a speech synthesis model whose synthesized sound meets the user's auditory expectations.

[0088] In some embodiments of this application, two different implementations of updating the basic speech synthesis model in step S5 above are described.

[0089] The first type

[0090] Reinforcement learning can be used to adjust the parameters of the basic speech synthesis model.

[0091] Specifically, the scoring results of each candidate speech are used as reward signals, and reinforcement learning is used to adjust the parameters of the basic speech synthesis model to obtain the adjusted speech synthesis model. The goal of reinforcement learning may include fine-tuning and updating the parameters of the basic speech synthesis model using the reward signals to maximize the expected reward, that is, to make the synthesized speech output by the basic speech synthesis model more in line with the user's listening perception goals.

[0092] In the reinforcement learning process described above, reinforcement learning algorithms such as Proximal Policy Optimization (PPO) or other reinforcement learning algorithms can be used. Considering that the reinforcement learning process fine-tunes the model parameters of the basic speech synthesis model to avoid excessively large model parameter updates reducing training stability, this embodiment can further incorporate a constraint policy into the reinforcement learning objective to limit the magnitude of model parameter updates. Specifically, the following loss function can be used:

[0093]

[0094] in, These are the model parameters of the basic speech synthesis model, E[] represents the expected value to be calculated, and f θ (x, y) represents the score of the reward model for the test text x and the output synthesized speech y, and KL[] represents the calculation of the KL divergence between the two distributions. This represents the parameter distribution of the base speech synthesis model after reinforcement learning updates. This represents the parameter distribution of the basic speech synthesis model before reinforcement learning.

[0095] for Figure 4 The training process of the example speech synthesis model, as described above. and These represent the parameter distributions of the vocoder before and after the reinforcement learning update, respectively; for Figure 5 The training process of the example speech synthesis model, as described above. and These represent the parameter distributions of the end-to-end basic speech synthesis model before and after reinforcement learning updates, respectively.

[0096] As can be seen from the above loss function, the goal of reinforcement learning in this embodiment is to maximize the expected reward while ensuring that the parameter distribution of the updated speech synthesis model does not deviate too much from the parameter distribution of the original speech synthesis model.

[0097] The second type

[0098] Candidate speech with higher reward scores can be used to update the basic speech synthesis model a second time.

[0099] Specifically, the top N candidate speech samples with the highest scores can be selected from the candidate speech samples corresponding to the test text. The test text and the top N candidate speech samples are then used to form a second-updated training set. Furthermore, using the second-updated training set, the parameters of the basic speech synthesis model are adjusted using gradient descent to obtain the adjusted speech synthesis model.

[0100] N can be 1, 2, or other values. When N is greater than 1, update weights can be set for each candidate speech, and the candidate speech with the higher score can have a larger update weight. For example, when N is 2, the candidate speech with the highest score can be assigned an update weight a1, and the candidate speech with the second highest score can be assigned an update weight a2, where a1 > a2.

[0101] By using candidate speech samples with higher scores to fine-tune and update the parameters of the basic speech synthesis model, the adjusted speech synthesis model can synthesize speech that is closer to the user's auditory perception goal.

[0102] The speech synthesis apparatus provided in the embodiments of this application is described below. The speech synthesis apparatus described below can be referred to in correspondence with the speech synthesis method described above.

[0103] See Figure 6 , Figure 6 This is a schematic diagram of the structure of a speech synthesis device disclosed in an embodiment of this application.

[0104] like Figure 6 As shown, the device may include:

[0105] The original text acquisition unit 11 is used to acquire the original text of the speech to be synthesized;

[0106] Text analysis unit 12 is used to perform text analysis on the original text to obtain the phoneme sequence corresponding to the original text;

[0107] The speech synthesis model processing unit 13 is used to input the phoneme sequence corresponding to the original text into the configured speech synthesis model to obtain the synthesized speech output by the model; the speech synthesis model is the final speech synthesis model after adjusting the parameters of the basic speech synthesis model by using the scoring results of multiple candidate speech synthesized by the basic speech synthesis model corresponding to the input test text as reward signals, wherein the scoring results of each candidate speech meet the user's listening perception target.

[0108] Optionally, the apparatus of this application may further include:

[0109] A speech synthesis model training unit is used to train the speech synthesis model.

[0110] The process of scoring each candidate speech in the speech synthesis model training unit during the training of the speech synthesis model includes:

[0111] The reward model is implemented using a configured reward model. The reward model is trained by using a text-speech pair containing the test text and a corresponding synthesized speech as training samples, and using the user rating results of the synthesized speech contained in the text-speech pair as sample labels.

[0112] Optionally, the process of training the speech synthesis model using the aforementioned speech synthesis model training unit may include:

[0113] A basic speech synthesis model is obtained, which is trained with the goal of making the synthesized speech of the test text approximate the original speech in the sound library corresponding to the test text;

[0114] Based on the aforementioned basic speech synthesis model, multiple candidate speech messages are generated for each test text in the first test set through sampling and decoding, resulting in a candidate speech set corresponding to each test text.

[0115] Obtain the user's rating results for each candidate speech in the candidate speech set corresponding to each test text, and obtain the candidate speech set carrying the rating results for each test text;

[0116] The reward model is trained using the set of candidate speech containing the scoring results corresponding to each test text as training data.

[0117] The trained reward model is used to score the multiple candidate speech synthesized by the basic speech synthesis model for each test text in the second test set. The score results of each candidate speech are used as reward signals to adjust the parameters of the basic speech synthesis model, resulting in the adjusted speech synthesis model.

[0118] Optionally, the basic speech synthesis model obtained by the above speech synthesis model training unit may include two structures, namely:

[0119] The first type

[0120] The basic speech synthesis model includes a cascaded acoustic model and a vocoder. The acoustic model is used to obtain acoustic features based on the input phoneme sequence, and the vocoder is used to convert the acoustic features into audio signals. The acoustic model and the vocoder are modeled separately.

[0121] The second type

[0122] The basic speech synthesis model adopts an end-to-end modeling approach.

[0123] Optionally, after obtaining the adjusted speech synthesis model, the above-mentioned speech synthesis model training unit can also be used for:

[0124] The adjusted speech synthesis model is updated to the basic speech synthesis model, and the training process of the speech synthesis model is iteratively executed until the set iteration termination condition is met, thus obtaining the final adjusted speech synthesis model.

[0125] Optionally, the process by which the above-mentioned speech synthesis model training unit uses the scoring results of each candidate speech as a reward signal to adjust the parameters of the basic speech synthesis model to obtain the adjusted speech synthesis model may include:

[0126] Using the score results of each candidate speech as a reward signal, the parameters of the basic speech synthesis model are adjusted using reinforcement learning to obtain an adjusted speech synthesis model. The goal of reinforcement learning includes fine-tuning and updating the parameters of the basic speech synthesis model using the reward signal to maximize the expected reward.

[0127] Alternatively, the process by which the above-mentioned speech synthesis model training unit uses the scoring results of each candidate speech as a reward signal to adjust the parameters of the basic speech synthesis model to obtain the adjusted speech synthesis model may include:

[0128] The top N candidate voices with the highest scores are selected from each candidate voice corresponding to the test text, and the test text and the top N candidate voices are used to form a second update training set;

[0129] Using the second updated training set, the parameters of the basic speech synthesis model are adjusted by gradient descent to obtain the adjusted speech synthesis model.

[0130] The speech synthesis apparatus provided in this application embodiment can be applied to speech synthesis devices, such as terminals, servers, and cloud computing. Optionally, Figure 6 The hardware structure block diagram of the speech synthesis device is shown below, with reference to... Figure 6 The hardware structure of a speech synthesis device may include: at least one processor 1, at least one communication interface 2, at least one memory 3, and at least one communication bus 4;

[0131] In this embodiment of the application, the number of processor 1, communication interface 2, memory 3, and communication bus 4 is at least one, and processor 1, communication interface 2, and memory 3 communicate with each other through communication bus 4;

[0132] Processor 1 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention.

[0133] Memory 3 may include high-speed RAM, and may also include non-volatile memory, such as at least one disk storage device;

[0134] The memory stores a program, which the processor can call. The program is used for:

[0135] Obtain the original text of the speech to be synthesized;

[0136] The original text is analyzed to obtain the phoneme sequence corresponding to the original text;

[0137] The phoneme sequence corresponding to the original text is input into the configured speech synthesis model to obtain the synthesized speech output by the model;

[0138] The speech synthesis model is a final speech synthesis model obtained by adjusting the parameters of the basic speech synthesis model, using the scoring results of multiple candidate speech synthesized by the basic speech synthesis model corresponding to the input test text as reward signals. The scoring results of each candidate speech meet the user's listening perception goals.

[0139] Optionally, the refined and extended functions of the program can be found in the description above.

[0140] This application embodiment also provides a storage medium that can store a program suitable for execution by a processor, the program being used for:

[0141] Obtain the original text of the speech to be synthesized;

[0142] The original text is analyzed to obtain the phoneme sequence corresponding to the original text;

[0143] The phoneme sequence corresponding to the original text is input into the configured speech synthesis model to obtain the synthesized speech output by the model;

[0144] The speech synthesis model is a final speech synthesis model obtained by adjusting the parameters of the basic speech synthesis model, using the scoring results of multiple candidate speech synthesized by the basic speech synthesis model corresponding to the input test text as reward signals. The scoring results of each candidate speech meet the user's listening perception goals.

[0145] Optionally, the refined and extended functions of the program can be found in the description above.

[0146] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0147] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referred to each other.

[0148] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A speech synthesis method characterized by, The method comprises: obtaining original text of to-be-synthesized speech; performing text analysis on the original text to obtain a phoneme sequence corresponding to the original text; inputting the phoneme sequence corresponding to the original text into a configured speech synthesis model to obtain synthesized speech output by the model; the speech synthesis model is a final speech synthesis model obtained by adjusting parameters of a basic speech synthesis model by taking a score result of a plurality of candidate speeches corresponding to input test text synthesized by the basic speech synthesis model as a reward signal, wherein the score result of each candidate speech conforms to a user listening experience target; wherein the basic speech synthesis model is adjusted in parameters by using reinforcement learning, and a loss function of the parameter adjustment process is: taking the difference between the score result of the reward model for the test text and the output synthesized speech and the KL divergence of the basic speech synthesis model parameter distribution after and before reinforcement learning update, calculating the expected value, and taking the opposite result.

2. The method of claim 1, wherein, The process of scoring each candidate speech is implemented by using a configured reward model; the reward model takes a text-speech pair containing a test text and a corresponding synthesized speech as a training sample, and takes a user score result of the synthesized speech contained in the text-speech pair as a sample label to train.

3. The method of claim 1, wherein, The basic speech synthesis model comprises a concatenated acoustic model and a vocoder, the acoustic model is used to obtain acoustic features based on input phoneme sequence, the vocoder is used to convert the acoustic features into audio signals, and the acoustic model and the vocoder are modeled separately; or, The basic speech synthesis model adopts an end-to-end modeling manner.

4. The method of claim 2, wherein, The training process of the speech synthesis model comprises: obtaining a basic speech synthesis model, the basic speech synthesis model is trained to approach synthesized speech of a test text to original speech corresponding to the test text in an audio library as a target; based on the basic speech synthesis model, a plurality of candidate speeches are generated for each test text in a first test set by using a sampling decoding manner, to obtain a candidate speech set corresponding to each test text; obtaining a user score result of each candidate speech in the candidate speech set corresponding to each test text, to obtain a candidate speech set carrying a score result corresponding to each test text; training a reward model by taking the candidate speech set carrying the score result corresponding to each test text as training data; using the trained reward model to score a plurality of candidate speeches synthesized by the basic speech synthesis model for each test text in a second test set, and taking the score result of each candidate speech as a reward signal to adjust parameters of the basic speech synthesis model to obtain an adjusted speech synthesis model.

5. The method of claim 4, wherein, After obtaining the adjusted speech synthesis model, further comprising: updating the adjusted speech synthesis model to the basic speech synthesis model, and iteratively performing the training process of the speech synthesis model until a set iteration end condition is reached to obtain a final adjusted speech synthesis model.

6. The method of claim 4, wherein, the process of taking the score result of each candidate speech as a reward signal to adjust parameters of the basic speech synthesis model to obtain an adjusted speech synthesis model comprises: The scoring results of the candidate speeches are taken as reward signals, and the basic speech synthesis model is adjusted in a reinforcement learning manner to obtain an adjusted speech synthesis model, wherein the target of the reinforcement learning includes fine-tuning and updating the parameters of the basic speech synthesis model by using the reward signals to maximize the expected reward.

7. The method of claim 6, wherein, The loss function is: wherein, is a model parameter of a base speech synthesis model, E[] represents a calculation of an expected value, represents a score of a reward model for a test text x and an output synthesized speech y, KL[] represents a calculation of a KL divergence of two distributions, represents a parameter distribution of the base speech synthesis model updated through reinforcement learning, represents a parameter distribution of the base speech synthesis model before reinforcement learning.

8. The method of claim 4, wherein, The scoring results of the candidate speeches are taken as reward signals, and the basic speech synthesis model is adjusted in a reinforcement learning manner to obtain an adjusted speech synthesis model, wherein the target of the reinforcement learning includes fine-tuning and updating the parameters of the basic speech synthesis model by using the reward signals to maximize the expected reward. The topN candidate speeches with the highest scoring results are selected from the candidate speeches corresponding to the test text, and the test text and the topN candidate speeches are used to form a secondary update training set. The basic speech synthesis model is adjusted in a gradient descent manner by using the secondary update training set to obtain an adjusted speech synthesis model.

9. A speech synthesis apparatus characterized by comprising: It comprises: An original text acquisition unit is configured to acquire an original text of a speech to be synthesized; A text analysis unit is configured to perform text analysis on the original text to obtain a phoneme sequence corresponding to the original text; A speech synthesis model processing unit is configured to input the phoneme sequence corresponding to the original text into a configured speech synthesis model to obtain a synthesized speech output by the model; the speech synthesis model is a final speech synthesis model obtained by adjusting the parameters of a basic speech synthesis model by taking the scoring results of multiple candidate speeches corresponding to the input test text synthesized by the basic speech synthesis model as reward signals, wherein the scoring results of the candidate speeches meet the user's hearing perception target. The parameters of the basic speech synthesis model are adjusted in a reinforcement learning manner, and the loss function of the parameter adjustment process is: the difference between the scoring results of the reward model of the test text and the output synthesized speech and the KL divergence of the parameter distribution of the basic speech synthesis model after and before the reinforcement learning update, the expected value is calculated, and the result of taking the inverse number is taken.

10. A speech synthesis device characterized by comprising: It comprises: A memory and a processor; The memory is configured to store a program; The processor is configured to execute the program to implement each step of the speech synthesis method according to any one of claims 1-8.

11. A storage medium having stored thereon a computer program, characterized in that The computer program is executed by the processor to implement each step of the speech synthesis method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Speech synthesis method and system

    CN106297766A