Method and device for training speech synthesis model, equipment and storage medium

By performing multiple inferences in the speech synthesis model to generate multiple speech token sequences and adjusting parameters, the problems of instability and insufficient naturalness in speech synthesis in the prior art are solved, and higher quality speech synthesis results are achieved.

CN120954380APending Publication Date: 2025-11-14BEIJING ZITIAO NETWORK TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410585425.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-11
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing speech synthesis technology has shortcomings in terms of stability, similarity, and naturalness of synthesized speech, resulting in a poor user experience.

Method used

By performing multiple inferences in the speech synthesis model to generate multiple speech token sequences, the synthesized speech content is reconstructed. The target loss is determined based on reward information, and the language model parameters are adjusted to optimize the speech synthesis effect through reinforcement learning.

Benefits of technology

It improves the stability and naturalness of speech synthesis, enhancing the quality of synthesized speech and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120954380A_ABST
    Figure CN120954380A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a method and device for training a speech synthesis model, equipment and a storage medium. The method comprises the following steps: executing a plurality of reasoning processes by utilizing a language model in a speech synthesis model, and generating a plurality of speech token sequences corresponding to input information; generating a plurality of synthetic voice contents corresponding to the plurality of voice token sequences; determining a target loss for the speech synthesis model at least based on the reward information of the multiple synthesis speech contents; and adjusting model parameters of the language model based on the target loss. Therefore, the embodiment of the invention can improve the speech synthesis effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The exemplary embodiments disclosed herein generally relate to the field of computers, and particularly to a method, apparatus, device, and computer-readable storage medium for training a speech synthesis model. Background Technology

[0002] With the increasing maturity of artificial intelligence and global wide area network (web) applications, speech synthesis technology is being used more and more widely on global wide area networks. Besides the clarity and intelligibility of synthesized speech, people are also placing higher demands on the naturalness, rhythm, and audio quality of synthesized speech. Therefore, improving the quality of synthesized speech by training a model has become a problem that requires continuous exploration and breakthroughs. Summary of the Invention

[0003] In a first aspect of this disclosure, a method for training a speech synthesis model is provided. The method includes: generating multiple speech token sequences corresponding to input information by performing multiple inference processes using a language model within the speech synthesis model; generating multiple synthesized speech contents corresponding to the multiple speech token sequences; determining a target loss for the speech synthesis model based at least on reward information from the multiple synthesized speech contents; and adjusting model parameters of the language model based on the target loss.

[0004] In a second aspect of this disclosure, an apparatus for training a speech synthesis model is provided. The apparatus includes: a token generation module configured to generate multiple speech token sequences corresponding to input information by performing multiple inference processes using a language model in the speech synthesis model; a speech generation module configured to generate multiple synthesized speech contents corresponding to the multiple speech token sequences; a loss determination module configured to determine a target loss for the speech synthesis model based at least on reward information from the multiple synthesized speech contents; and a parameter adjustment module configured to adjust model parameters of the language model based on the target loss.

[0005] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.

[0006] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program that can be executed by a processor to implement the method of the first aspect.

[0007] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0008] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0009] Figure 1 A schematic diagram is shown of an example environment in which embodiments of the present disclosure may be implemented;

[0010] Figure 2 A schematic diagram of an example architecture for training a speech synthesis model according to some embodiments of the present disclosure is shown;

[0011] Figure 3 A flowchart illustrating an example process for training a speech synthesis model according to some embodiments of the present disclosure is shown;

[0012] Figure 4 A schematic structural block diagram of an example apparatus for training a speech synthesis model according to some embodiments of the present disclosure is shown; and

[0013] Figure 5 A block diagram of an electronic device capable of implementing several embodiments of the present disclosure is shown. Detailed Implementation

[0014] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0015] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.

[0016] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0017] As used in this paper, the term "model" refers to a system that learns the relationship between inputs and outputs from training data, enabling it to generate corresponding outputs for a given input after training. Model generation can be based on machine learning techniques. Deep learning is a machine learning algorithm that uses multiple layers of processing units to process inputs and provide corresponding outputs. In this paper, "model" may also be referred to as a "machine learning model," a "machine learning network," or simply a "network," and these terms are used interchangeably. A model can also include different types of processing units or networks.

[0018] The embodiments of this disclosure may involve user data, data acquisition, and / or use. All of these aspects comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained in accordance with relevant laws and regulations through appropriate means. The specific methods of notification and / or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.

[0019] In this specification and the embodiments, any processing of personal information will be carried out only under the premise of legality (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information other than that necessary for basic functions will not affect the user's use of basic functions.

[0020] Currently, speech synthesis can be performed using language models and diffusion models. The language model predicts a sequence of speech tokens, and the diffusion model and vocoder convert these token sequences into audio. While the speech synthesized in this way possesses high expressiveness and richness, some synthesized results still occasionally produce results that users dislike. These include poor synthesis stability (e.g., unclear pronunciation, omissions, and extra words) and low similarity.

[0021] In view of this, embodiments of the present disclosure propose an improved scheme for training a speech synthesis model. According to this scheme, multiple inference processes are performed using the language model within the speech synthesis model to generate multiple speech token sequences corresponding to the input information. Subsequently, multiple synthesized speech contents corresponding to the multiple speech token sequences are generated. Then, a target loss for the speech synthesis model is determined based at least on the reward information of the multiple synthesized speech contents. The model parameters of the language model are adjusted according to the target loss.

[0022] Therefore, embodiments of this disclosure can improve the performance of speech synthesis through reinforcement learning.

[0023] The following section provides a detailed description of various example implementations of this scheme, with reference to the accompanying drawings.

[0024] Example Environment

[0025] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. For example... Figure 1 As shown, example environment 100 may include electronic device 110.

[0026] In this example environment 100, electronic device 110 may run a speech synthesis model 120 for synthesizing speech. The speech synthesis model 120 may include a language model 140 and a reconstruction module 145.

[0027] In some embodiments, electronic device 110 acquires input information 130 (e.g., text) and input information 135 (e.g., prompt voice). Then, electronic device 110 uses language model 140 to perform multiple inferences on the acquired input information 130 and input information 135 to obtain multiple voice token sequences. Subsequently, electronic device 110 reconstructs the multiple voice token sequences back to audio 150 using reconstruction module 145.

[0028] The electronic device 110 can determine the loss for the speech synthesis model 120 based at least on the reward information of the audio 150. In some embodiments, the electronic device 110 trains the speech synthesis model 120 based on the loss, thereby improving the performance of speech synthesis by utilizing the trained speech synthesis model 120.

[0029] In some embodiments, electronic device 110 can be any type of computing device, including terminal devices or server devices. Terminal devices can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR / AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. In some embodiments, electronic device 110 can also support any type of user-facing interface (such as "wearable" circuitry). Server devices may include, for example, computing systems / servers such as mainframes, edge computing nodes, computing devices in cloud environments, etc.

[0030] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.

[0031] The following description will continue with reference to the accompanying drawings, which will provide some exemplary embodiments of this disclosure.

[0032] Example process

[0033] Figure 2 A schematic diagram of an example architecture 200 for training a speech synthesis model according to some embodiments of the present disclosure is shown. Process 200 can be implemented at electronic device 110. Reference is made below. Figure 1 To describe process 200.

[0034] In some embodiments, electronic device 110 generates multiple speech token sequences corresponding to input information by performing multiple inference processes using a language model in a speech synthesis model. In example architecture 200, electronic device 110 acquires text information 211 and prompt information 212. In some examples, the prompt information 211 acquired by electronic device 110 may be audio, which may be audio corresponding to the text information 211. In some examples, the input information may be the text information 211 and prompt information 212 acquired by electronic device 110.

[0035] Electronic device 110 performs N (N>1) reasoning processes on text information 211 and prompt information 212 by utilizing language model 140 in speech synthesis model 120 to generate N (N>1) voice token sequences corresponding to text information 211 and prompt information 212.

[0036] In some embodiments, electronic device 110 generates multiple synthesized speech content corresponding to multiple voice token sequences. In some examples, electronic device 110 can reconstruct the multiple voice token sequences it generates back into multiple synthesized speech content 213 via reconstruction module 145.

[0037] In some embodiments, the electronic device 110 determines a target loss for the speech synthesis model based on reward information from at least a plurality of synthesized speech contents. The electronic device 110 then adjusts the model parameters of the language model based on the target loss.

[0038] In some embodiments, the electronic device 110 may utilize a reward model to process and determine reward information for multiple synthesized speech contents. In some embodiments, the reward model may be trained based on training preference data. In some embodiments, the reward model may also be configured to determine scores for multiple synthesized speech contents with respect to at least one evaluation metric.

[0039] In some examples, electronic device 110 uses a reward model to determine reward information 214 for each multi-word synthesized speech content 213. In some examples, electronic device 110 may train the reward model based on manually labeled preference data.

[0040] In other examples, the electronic device 110 may also configure the reward model to determine the scores of multiple synthesized speech contents with respect to at least one evaluation metric. In some embodiments, the at least one evaluation metric includes at least one of the following: a first evaluation metric, a second evaluation metric, a third evaluation metric, and a fourth evaluation metric.

[0041] In some embodiments, a first evaluation metric indicates the similarity between the corresponding synthesized speech content and reference speech content. The reference speech content corresponds to input information, for example, the reference speech content corresponds to cue information 212 (e.g., cue audio). In some examples, the electronic device 110 may utilize a speaker recognition model (ASV) to calculate the similarity between the synthesized speech content 213 and the reference speech content.

[0042] In some embodiments, the second evaluation metric indicates the pronunciation stability of the corresponding synthesized speech content. In some examples, the electronic device 110 may utilize a speech recognition model (WER) to calculate the pronunciation stability of the synthesized speech content 213.

[0043] In some embodiments, a third evaluation metric indicates the expressive state of the corresponding synthesized speech content. In some examples, the electronic device 110 may use a model to calculate the expressive state of the synthesized speech content 213. Such expressive state may, for example, indicate the speaking state of the speech content 213.

[0044] In some embodiments, the fourth evaluation metric may indicate the quality of the corresponding synthesized speech content, and examples may include, but are not limited to, the naturalness, expressiveness, and sound quality of the synthesized speech content. In some embodiments, the electronic device 110 may utilize an appropriate quality assessment model to determine the quality of the synthesized speech content.

[0045] Then, the electronic device 110 determines a target loss 218 for the speech synthesis model 120, at least based on the reward information 214 determined by it using the reward model. In some embodiments, the electronic device 110 adjusts the model parameters of the language model 140 based on the target loss 218. The determination of the target loss 218 by the electronic device 110 will be described in detail below.

[0046] In some embodiments, during the adjustment of the model parameters of the language model by the electronic device 110, at least one other model parameter of the speech synthesis model is fixed. It is understood that when at least one other model parameter of the speech synthesis model 120 is fixed, the electronic device 110 adjusts the model parameters of the language model 140 according to the target loss function 218.

[0047] The process by which electronic device 110 determines target loss 218 will be described in detail below with reference to example architecture 200.

[0048] In some embodiments, the electronic device 110 processes multiple voice token sequences using a language model to determine first probability information corresponding to the multiple voice token sequences. Based on the first probability information and reward information, the electronic device 110 determines a first loss for the speech synthesis model. Then, the electronic device 110 determines a target loss based on the first loss.

[0049] In some examples, electronic device 110 inputs multiple speech token sequences into language model 140 to obtain first probability information corresponding to the multiple speech token sequences. Then, electronic device 110 uses reward information 214 and the first probability information to calculate a first loss (also known as reinforcement learning loss) 215 for speech synthesis model 120. Finally, electronic device 110 determines a target loss 218 based on the first loss 215.

[0050] In some embodiments, the electronic device 110 determines the difference between reward information corresponding to multiple synthesized speech contents and a threshold. Then, the electronic device 110 determines a first loss based on the difference and first probability information. Understandably, the electronic device 110 processes multiple synthesized speech contents using a reward model to determine reward information corresponding to each of the multiple synthesized speech contents. In some examples, the reward information can be represented as a score, where a higher score indicates a better performance of the synthesized speech content corresponding to that reward information.

[0051] In some examples, electronic device 110 determines the difference between a pre-set threshold and the corresponding scores of multiple synthesized speech contents. Then, electronic device 110 determines a first loss based on the difference between the threshold and the multiple synthesized speech contents.

[0052] In some embodiments, the first loss can be determined using the following formula:

[0053]

[0054] Where b is a constant hyperparameter, β is the set of N different output speech token sequences corresponding to the same input x, and y i It refers to the i-th output corresponding to the same x, γ(y i |x) is y i Corresponding reward information, It is y i The first probability information is calculated using a language model.

[0055] In some embodiments, the electronic device 110 may also utilize a reference model to process multiple voice token sequences to determine second probability information corresponding to the multiple voice token sequences. The reference model has the same initial parameters as the language model. Then, the electronic device 110 determines a second loss based on a comparison of the first and second probability information. The electronic device 110 determines a target loss based on the first and second losses.

[0056] In some examples, electronic device 110 inputs multiple voice token sequences into a reference model 216 with the same initial parameters as language model 140 to obtain second probability information corresponding to the multiple voice token sequences. Electronic device 110 calculates a second loss 217 based on a comparison of the first and second probability information. As an example, the second loss 217 can be determined based on the KL divergence between the first and second probability information. In some examples, the second loss 217 can be used to constrain language model 140 to prevent forgetting.

[0057] Then, the electronic device 110 determines a target loss 218 for the speech synthesis model 120 based on its calculated first loss 215 and second loss 217. In some embodiments, the target loss is a weighted sum of the first loss and the second loss. It is understood that the first loss 215 and the second loss 217 are weighted and summed to obtain the target loss 218.

[0058] Based on this approach, the embodiments of this disclosure can align the model's responses with preference data, making the model-generated responses more meaningful. Furthermore, by scoring the speech synthesis results across various dimensions using other models and using these scores as reward information for the speech synthesis model, the model can be optimized, thereby improving the speech synthesis effect.

[0059] Example process

[0060] Figure 3 A flowchart of an example process 300 for training a speech synthesis model according to some embodiments of the present disclosure is shown. Process 300 can be implemented at electronic device 110. Reference is made below. Figure 1 To describe process 300.

[0061] As shown in the figure, in box 310, electronic device 110 generates multiple speech token sequences corresponding to the input information by performing multiple inference processes using the language model in the speech synthesis model.

[0062] In box 320, electronic device 110 generates multiple synthesized speech contents corresponding to multiple voice token sequences.

[0063] In box 330, electronic device 110 determines the target loss for the speech synthesis model based on reward information from at least a number of synthesized speech contents.

[0064] In box 340, electronic device 110 adjusts the model parameters of the language model based on the target loss.

[0065] In some embodiments, process 300 further includes: processing reward information for determining multiple synthesized speech content using a reward model, wherein the reward model is trained based on training preference data, or the reward model is configured to determine scores for multiple synthesized speech content with respect to at least one evaluation metric.

[0066] In some embodiments, at least one evaluation metric includes at least one of the following: a first evaluation metric indicating the similarity between the corresponding synthesized speech content and reference speech content, the reference speech content corresponding to the input information; a second evaluation metric indicating the pronunciation stability of the corresponding synthesized speech content; a third evaluation metric indicating the expressive state of the corresponding synthesized speech content; and a fourth evaluation metric indicating the quality of the corresponding synthesized speech content.

[0067] In some embodiments, determining a target loss for a speech synthesis model based at least on reward information of multiple synthesized speech content includes: processing multiple speech token sequences using a language model to determine first probability information corresponding to the multiple speech token sequences; determining a first loss for the speech synthesis model based on the first probability information and reward information; and determining a target loss based on the first loss.

[0068] In some embodiments, determining a first loss for the speech synthesis model based on first probability information and reward information includes: determining the difference between reward information and a threshold corresponding to multiple synthesized speech contents; and determining the first loss based on the difference and the first probability information.

[0069] In some embodiments, determining the target loss based on reinforcement learning loss includes: processing multiple speech token sequences using a reference model to determine second probability information corresponding to the multiple speech token sequences, the reference model having the same initial parameters as the language model; determining a second loss based on a comparison of the first probability information and the second probability information; and determining the target loss based on the first loss and the second loss.

[0070] In some embodiments, the target loss is a weighted sum of the first loss and the second loss.

[0071] In some embodiments, during the adjustment of the model parameters of the language model, at least one other model parameter of the speech synthesis model is fixed.

[0072] Example devices and equipment

[0073] Embodiments of this disclosure also provide corresponding apparatus for implementing the above methods or processes. Figure 4 A schematic structural block diagram of an example apparatus 400 for training a speech synthesis model according to certain embodiments of the present disclosure is shown. Apparatus 400 may be implemented as or included in electronic device 110. Various modules / components in apparatus 400 may be implemented by hardware, software, firmware, or any combination thereof.

[0074] like Figure 4 As shown, the apparatus 400 includes a token generation module 410 configured to generate multiple speech token sequences corresponding to input information by performing multiple inference processes using a language model in a speech synthesis model. The apparatus 400 also includes a speech generation module 420 configured to generate multiple synthesized speech contents corresponding to the multiple speech token sequences. The apparatus 400 further includes a loss determination module 430 configured to determine a target loss for the speech synthesis model based at least on reward information from the multiple synthesized speech contents. The apparatus 400 also includes a parameter adjustment module 440 configured to adjust the model parameters of the language model based on the target loss.

[0075] In some embodiments, the apparatus 400 further includes a reward information determination module configured to process and determine reward information for multiple synthesized speech contents using a reward model, wherein the reward model is trained based on training preference data, or the reward model is configured to determine scores for multiple synthesized speech contents with respect to at least one evaluation metric.

[0076] In some embodiments, at least one evaluation metric includes at least one of the following: a first evaluation metric indicating the similarity between the corresponding synthesized speech content and reference speech content, the reference speech content corresponding to the input information; a second evaluation metric indicating the pronunciation stability of the corresponding synthesized speech content; a third evaluation metric indicating the expressive state of the corresponding synthesized speech content; and a fourth evaluation metric indicating the quality of the corresponding synthesized speech content.

[0077] In some embodiments, the loss determination module 430 is further configured to process multiple speech token sequences using a language model to determine first probability information corresponding to the multiple speech token sequences; determine a first loss for the speech synthesis model based on the first probability information and reward information; and determine a target loss based on the first loss.

[0078] In some embodiments, the loss determination module 430 is further configured to determine the difference between reward information and threshold corresponding to multiple synthesized speech contents; and to determine a first loss based on the difference and first probability information.

[0079] In some embodiments, the loss determination module 430 is further configured to process multiple voice token sequences using a reference model to determine second probability information corresponding to the multiple voice token sequences, the reference model having the same initial parameters as the language model; determine a second loss based on a comparison of the first probability information and the second probability information; and determine a target loss based on the first loss and the second loss.

[0080] In some embodiments, the target loss is a weighted sum of the first loss and the second loss.

[0081] In some embodiments, during the adjustment of the model parameters of the language model, at least one other model parameter of the speech synthesis model is fixed.

[0082] Figure 5 A block diagram of an electronic device 500 in which one or more embodiments of the present disclosure may be implemented is shown. It should be understood that... Figure 5 The electronic device 500 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 5 The electronic device 500 shown can be used to achieve Figure 1 Electronic devices 110.

[0083] like Figure 5 As shown, electronic device 500 is in the form of a general-purpose electronic device. Components of electronic device 500 may include, but are not limited to, one or more processors or processing units 510, memory 520, storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. Processing unit 510 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 520. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 500.

[0084] Electronic device 500 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 520 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 530 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 500.

[0085] Electronic device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 5 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 520 may include computer program product 525 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.

[0086] Communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functionality of components of electronic device 500 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, electronic device 500 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0087] Input device 550 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 560 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 500 can also communicate with one or more external devices (not shown) via communication unit 540 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 500, or with any device that enables electronic device 500 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).

[0088] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.

[0089] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0090] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0091] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0092] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0093] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for training a speech synthesis model, comprising: By utilizing the language model in the speech synthesis model to perform multiple inference processes, multiple speech token sequences corresponding to the input information are generated. Generate multiple synthesized speech contents corresponding to the multiple voice token sequences; Based at least on the reward information of the multiple synthesized speech contents, determine the target loss for the speech synthesis model; as well as Based on the target loss, the model parameters of the language model are adjusted.

2. The method according to claim 1, further comprising: The reward information of the plurality of synthesized speech content is determined by using a reward model, wherein the reward model is trained based on training preference data, or the reward model is configured to determine the score of the plurality of synthesized speech content with respect to at least one evaluation metric.

3. The method according to claim 2, wherein the at least one evaluation index includes at least one of the following: The first evaluation metric indicates the similarity between the corresponding synthesized speech content and the reference speech content, which corresponds to the input information; The second evaluation metric indicates the pronunciation stability of the corresponding synthesized speech content; The third evaluation indicator indicates the expression status of the corresponding synthesized speech content; The fourth evaluation metric indicates the quality of the corresponding synthesized speech content.

4. The method of claim 1, wherein determining the target loss for the speech synthesis model based at least on reward information of the plurality of synthesized speech contents includes: The language model is used to process the plurality of voice token sequences to determine first probability information corresponding to the plurality of voice token sequences; Based on the first probability information and the reward information, a first loss is determined for the speech synthesis model; as well as Based on the first loss, the target loss is determined.

5. The method of claim 4, wherein determining the first loss for the speech synthesis model based on the first probability information and the reward information comprises: Determine the difference between the reward information and the threshold corresponding to the multiple synthesized speech contents; as well as Based on the difference and the first probability information, the first loss is determined.

6. The method of claim 4, wherein determining the target loss based on the reinforcement learning loss comprises: The plurality of voice token sequences are processed using a reference model to determine second probability information corresponding to the plurality of voice token sequences, wherein the reference model has the same initial parameters as the language model; Based on the comparison between the first probability information and the second probability information, a second loss is determined; as well as The target loss is determined based on the first loss and the second loss.

7. The method of claim 6, wherein the target loss is a weighted sum of the first loss and the second loss.

8. The method of claim 1, wherein at least one other model parameter of the speech synthesis model is fixed during the adjustment of the model parameters of the language model.

9. An apparatus for training a speech synthesis model, comprising: The token generation module is configured to generate multiple speech token sequences corresponding to the input information by performing multiple inference processes using the language model in the speech synthesis model. The speech generation module is configured to generate multiple synthesized speech contents corresponding to the multiple speech token sequences; The loss determination module is configured to determine the target loss for the speech synthesis model based at least on reward information of the multiple synthesized speech contents; as well as The parameter adjustment module is configured to adjust the model parameters of the language model based on the target loss.

10. An electronic device, comprising: At least one processing unit; as well as At least one memory, coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, which, when executed by the at least one processing unit, cause the electronic device to perform the method according to any one of claims 1 to 9.

11. A computer-readable storage medium having a computer program stored thereon, the computer program being executable by a processor to implement the method according to any one of claims 1 to 9.