Speech synthesis method, device, electronic device and storage medium

By constructing multiple preference data sets and assigning weights through multi-objective direct preference optimization, the speech synthesis model is trained. This solves the problem that single preference alignment in existing technologies cannot meet diverse needs, and achieves efficient and stable speech synthesis effects in different scenarios.

CN119785766BActive Publication Date: 2025-09-23IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411754113.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-02
Publication Date
2025-09-23
Estimated Expiration
2044-12-02

AI Technical Summary

Technical Problem

Existing speech synthesis models fail to effectively consider human subjective evaluation during the training process, resulting in the difficulty of synthesized speech meeting the diverse human preference needs. Existing methods such as RLHF and DPO can only align to a single preference and cannot adapt to the diverse needs in different scenarios.

Method used

Through a multi-objective direct preference optimization approach, multiple preference data sets are constructed and assigned different preference weights. The speech synthesis model is trained to meet diverse human preference needs. The model training is performed by aligning multiple preferences, avoiding the limitations of existing methods.

Benefits of technology

It achieves the goal of aligning human preferences from multiple dimensions while maintaining training efficiency and stability, meeting diverse needs in different scenarios, synthesizing higher-quality speech, and adapting to users' diverse preferences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119785766B_ABST
    Figure CN119785766B_ABST
Patent Text Reader

Abstract

The present invention provides a speech synthesis method, device, electronic device and storage medium, wherein the method comprises: based on the user's speech synthesis preference, selecting a target speech synthesis model from multiple speech synthesis models, applying the target speech synthesis model to perform speech synthesis based on the text to be synthesized, and obtaining synthesized speech that meets the speech synthesis preference; each speech synthesis model is trained based on a preference data set and preference weight configuration corresponding to multiple preferences, and different speech synthesis models are trained using different preference weight configurations, thereby overcoming the defect that the human preference alignment method for speech synthesis models in traditional solutions can only align a single human preference and cannot meet the diverse human preference needs. The speech synthesis model is trained by a multi-objective direct preference optimization method, which not only makes the training process simpler and more efficient, but also can perform human preference alignment from multiple dimensions, and meet the diverse human preference needs by assigning weights to each preference.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech synthesis technology, and in particular to a speech synthesis method, device, electronic device and storage medium. Background Art

[0002] In recent years, speech synthesis models have become mainstream technology in the field, thanks to their powerful speech generation capabilities and zero-shot adaptability to unseen speakers. However, despite their numerous advantages, practical applications still face a significant challenge: ensuring that the quality of synthesized speech aligns with human subjective evaluations. Traditional speech synthesis model training doesn't directly incorporate metrics such as the Mean Opinion Score (MOS), which is often used for subjective evaluation. This leads to a certain misalignment between the training objectives and human evaluation criteria. This misalignment makes it difficult for the trained models to produce speech that fully meets human expectations.

[0003] To address this, researchers have proposed methods based on human feedback reinforcement learning and direct preference optimization. The former pre-trains a reward model and applies a reinforcement learning algorithm to adjust the speech synthesis model to generate speech with higher reward scores. However, this method relies on online sampling, which is time-consuming and results in low model training efficiency. The latter optimizes the classification function to make the model output closer to the distribution of real data, thereby aligning the speech synthesis model with human preferences. Although both methods can adjust the speech synthesis model to align with human preferences, the human preferences they align to are specific. However, in reality, human preferences for synthesized speech are diverse, and the emphasis of preferences varies in different scenarios. Current solutions cannot adapt to the diverse human preference needs in different scenarios. Summary of the Invention

[0004] The present invention provides a speech synthesis method, device, electronic device and storage medium, which are used to solve the defect that the human preference alignment method for speech synthesis models in the prior art can only align a single human preference and cannot meet the diverse human preference needs. The present invention can align the human preferences of the speech synthesis model from multiple dimensions and meet the diverse human preference needs by setting different preference weight configurations.

[0005] The present invention provides a speech synthesis method, comprising:

[0006] Determine the text to be synthesized and the user's speech synthesis preferences;

[0007] Based on the speech synthesis preference, a target speech synthesis model is selected from a plurality of speech synthesis models, and based on the text to be synthesized, the target speech synthesis model is applied to perform speech synthesis to obtain synthesized speech that meets the speech synthesis preference;

[0008] Among them, each speech synthesis model is trained based on preference data sets and preference weight configurations corresponding to multiple preferences. The preference weight configurations used in training different speech synthesis models are different. Each preference data set contains multiple sets of preference data corresponding to preferences, and each set of preference data includes sample text corresponding to the preference, sample preference speech and sample non-preferential speech.

[0009] According to a speech synthesis method provided by the present invention, a preference data set corresponding to each preference is determined based on the following steps:

[0010] Determining a sample text corresponding to a preference and a prompt voice corresponding to the sample text;

[0011] Determining, based on an initial speech synthesis model, a plurality of initial synthesized speech sounds corresponding to the sample text and the prompt speech sound; wherein the initial speech synthesis model is trained based on a pre-training data set, and the pre-training data set includes pre-trained text and corresponding pre-trained prompt speech sounds;

[0012] Determining a speech score of each initial synthesized speech under a corresponding preference, and determining, from each initial synthesized speech, a sample preferred speech and a sample non-preferred speech corresponding to the corresponding preference based on the speech score corresponding to each initial synthesized speech;

[0013] Based on the sample text and prompt voice corresponding to the preference, and the corresponding sample preferred voice and the sample non-preferred voice, a preference data set corresponding to the preference is constructed.

[0014] According to a speech synthesis method provided by the present invention, determining a speech score of each initial synthesized speech under a corresponding preference, and determining, from each initial synthesized speech based on the speech score corresponding to each initial synthesized speech, a sample preferred speech and a sample non-preferred speech corresponding to the corresponding preference, comprising:

[0015] determining a coarse-grained speech score of each of the initial synthesized speech under the plurality of preferences;

[0016] Based on the coarse-grained speech scores of the initial synthesized speech corresponding to the plurality of preferences, screening the initial synthesized speech to obtain a plurality of candidate synthesized speech;

[0017] Determining a fine-grained speech score for each candidate synthesized speech under corresponding preferences;

[0018] Based on the fine-grained speech scores corresponding to the candidate synthesized speech, sample preferred speech and sample non-preferred speech corresponding to the candidate synthesized speech are determined from the candidate synthesized speech.

[0019] According to a speech synthesis method provided by the present invention, the multiple preferences include a naturalness preference, an intelligibility preference, and a similarity preference; and the method of determining, from the candidate synthesized speech, sample preferred speech and sample non-preferred speech corresponding to the corresponding preference based on the fine-grained speech scores corresponding to the candidate synthesized speech, includes:

[0020] From the candidate synthesized speech corresponding to the naturalness preference, selecting the candidate synthesized speech with the highest and lowest fine-grained speech scores under the naturalness preference as the sample preferred speech and the sample non-preferred speech corresponding to the naturalness preference, respectively;

[0021] From the candidate synthesized speech corresponding to the intelligibility preference, selecting the candidate synthesized speech with the highest and lowest fine-grained speech scores under the intelligibility preference as the sample preferred speech and the sample non-preferred speech corresponding to the intelligibility preference, respectively;

[0022] From the candidate synthesized speech corresponding to the similarity preference, the candidate synthesized speech with the highest and lowest fine-grained speech scores under the similarity preference are selected as the sample preferred speech and the sample non-preferred speech corresponding to the similarity preference, respectively.

[0023] According to a speech synthesis method provided by the present invention, determining a plurality of initial synthesized speech corresponding to the sample text and the prompt speech based on an initial speech synthesis model includes:

[0024] Performing phoneme conversion on the sample text to obtain a sample phoneme sequence;

[0025] Performing feature encoding on the prompt voice to obtain sample voice features;

[0026] Embedding the sample phoneme sequence and the sample speech feature, and concatenating the sample phoneme vector representation and the sample speech vector representation obtained based on the embedded coding to obtain a sample fusion vector representation;

[0027] Based on the sample fusion vector representation, the initial speech synthesis model is applied to perform speech synthesis to obtain multiple initial synthesized speech.

[0028] According to a speech synthesis method provided by the present invention, each speech synthesis model is trained based on the following steps:

[0029] Selecting a preference data set corresponding to any one preference from the plurality of preference data sets corresponding to the preferences as a target preference data set;

[0030] Determining, based on the initial speech synthesis model, sample synthesized speech corresponding to the sample text and prompt speech in each group of preference data in the target preference data set;

[0031] Determining a reward score for the sample synthesized speech corresponding to each set of preference data under the other preferences based on a reward model for other preferences; the reward model is trained based on the preference data set corresponding to the other preferences;

[0032] Based on the sample preferred speech and sample non-preferred speech in each group of preference data, the sample synthesized speech corresponding to each group of preference data, the reward score of the sample synthesized speech corresponding to other preferences, and the preference weight configuration corresponding to the initial speech synthesis model, the initial speech synthesis model is iterated on parameters to obtain a speech synthesis model.

[0033] According to a speech synthesis method provided by the present invention, based on the sample preferred speech and sample non-preferred speech in each set of preference data, the sample synthesized speech corresponding to each set of preference data, the reward score of the sample synthesized speech corresponding to other preferences, and the preference weight configuration corresponding to the initial speech synthesis model, the initial speech synthesis model is iterated to obtain the speech synthesis model, including:

[0034] Determining a first loss based on the sample preferred speech and the sample non-preferred speech in each set of preference data, the sample synthesized speech corresponding to each set of preference data, and the weight of any preference in the preference weight configuration corresponding to the initial speech synthesis model;

[0035] Determining a second loss based on the sample preferred speech and sample non-preferred speech in each set of preference data, the sample synthesized speech and its reward score corresponding to other preferences, and the weight of the other preferences in the preference weight configuration corresponding to the initial speech synthesis model;

[0036] Based on the first loss and the second loss, parameters of the initial speech synthesis model are iterated to obtain a speech synthesis model.

[0037] The present invention also provides a speech synthesis device, comprising:

[0038] a data determination unit, configured to determine the text to be synthesized and the user's speech synthesis preference;

[0039] a speech synthesis unit, configured to select a target speech synthesis model from a plurality of speech synthesis models based on the speech synthesis preference, and perform speech synthesis based on the text to be synthesized by applying the target speech synthesis model to obtain synthesized speech that meets the speech synthesis preference;

[0040] Among them, each speech synthesis model is trained based on preference data sets and preference weight configurations corresponding to multiple preferences. The preference weight configurations used in training different speech synthesis models are different. Each preference data set contains multiple sets of preference data corresponding to preferences, and each set of preference data includes sample text corresponding to the preference, sample preference speech and sample non-preferential speech.

[0041] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor implements any of the above-mentioned speech synthesis methods when executing the computer program.

[0042] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which implements any of the above-mentioned speech synthesis methods when executed by a processor.

[0043] The speech synthesis method, device, electronic device and storage medium provided by the present invention select a target speech synthesis model from multiple speech synthesis models according to the user's speech synthesis preference, and apply the target speech synthesis model to perform speech synthesis according to the text to be synthesized, so as to obtain synthesized speech that meets the speech synthesis preference; wherein, each speech synthesis model is trained based on the preference data set and preference weight configuration corresponding to multiple preferences, and the preference weight configurations used in the training of different speech synthesis models are different, which overcomes the defect that the human preference alignment method for speech synthesis models in traditional solutions can only align a single human preference and cannot meet the diverse human preference needs. The speech synthesis model is trained by a multi-objective direct preference optimization method, which not only makes the training process simpler and more efficient, but also can perform human preference alignment from multiple dimensions, and meet the diverse human preference needs by assigning weights to each preference. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0045] Figure 1 1 is a flow chart of the speech synthesis method provided by the present invention;

[0046] Figure 2 This is an example diagram of the framework of the model training process provided by the present invention;

[0047] Figure 3 It is a structural diagram of the speech synthesis device provided by the present invention;

[0048] Figure 4 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0049] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0050] Currently, there are two main methods for aligning speech synthesis models with human preferences. The first is based on RLHF (Reinforcement Learning from Human Feedback). This method trains a reward model on a dataset that reflects human preferences and applies a reinforcement learning algorithm to adjust the speech synthesis model, optimizing it towards generating speech with high reward scores while minimizing deviation from the original model. However, this method takes a long time to sample and generate speech, and the model training efficiency is low.

[0051] The other is based on the DPO (Direct Preference Optimization) method. Unlike traditional preference-driven reinforcement learning, this method directly optimizes the closed-form loss of preference data without relying on an explicit reward function, avoiding the problem of time-consuming online sampling. Compared with the RLHF method, training is more efficient and stable.

[0052] However, while both of these methods can adjust the speech synthesis model to align with human preferences, the preferences they align to are specific and difficult to meet the diverse needs of human preferences. In actual applications, people's preferences for generated speech vary in different scenarios. For example, in a sound reproduction scenario, people are more concerned about whether the model can generate speech that is highly similar to the timbre of the prompt speech; in an audiobook scenario, people are more concerned about whether the model can generate speech that is highly natural and expressive. Clearly, neither RLHF nor DPO can currently achieve this and cannot adapt to the diverse preferences in different scenarios.

[0053] Furthermore, to address the limitations of single-objective preference alignment, research is currently underway on using multi-objective RLHF to optimize speech synthesis models. This approach trains separate reward models for different objectives and performs weighted optimization on the scoring results of multiple reward models to achieve alignment of multiple human preferences. However, this approach also has drawbacks such as cumbersome processes and unstable training. In the field of speech synthesis, when models encounter diverse demands, the preferences of multiple objectives often exhibit certain conflicts. For example, the generated speech may not be highly natural while being highly intelligible. Therefore, how to achieve alignment of diverse human preferences while maintaining efficient and stable training remains an urgent problem in the field of speech synthesis.

[0054] In this regard, the present invention provides a speech synthesis method, which aims to solve the problem that RLHF and DPO optimization are difficult to meet the diversity of human preferences. The speech synthesis model is trained by multi-objective direct preference optimization, which abandons the dependence on RLHF, makes training simpler and more efficient, and satisfies diverse human preference needs by assigning weights to each target (preference).

[0055] Figure 1 Schematic diagram of the speech synthesis method provided by the present invention, such as Figure 1 As shown, the method includes:

[0056] Step 110, determining the text to be synthesized and the user's speech synthesis preference;

[0057] Step 120 , based on the speech synthesis preference, select a target speech synthesis model from multiple speech synthesis models, and perform speech synthesis using the target speech synthesis model based on the text to be synthesized to obtain synthesized speech that meets the speech synthesis preference;

[0058] Among them, each speech synthesis model is trained based on preference data sets and preference weight configurations corresponding to multiple preferences. The preference weight configurations used in training different speech synthesis models are different. Each preference data set contains multiple sets of preference data corresponding to preferences, and each set of preference data includes sample text corresponding to the preference, sample preference speech and sample non-preferential speech.

[0059] Specifically, considering that both current RLHF and DPO methods can only align to a single human preference, people's preferences for generated speech vary in practice, and their preferences for generated speech differ in different scenarios. For example, in the context of sound reproduction, people focus more on the similarity of the generated speech, while in the context of audiobooks, they prioritize the naturalness and expressiveness of the generated speech. Clearly, these two current methods, which can only align to a single human preference, are not suitable for different scenarios and cannot meet the diverse human preferences in different scenarios.

[0060] In view of this, in an embodiment of the present invention, it is proposed that multi-target human preference alignment can be performed. In order to avoid the limitations of the current multi-target RLHF, in an embodiment of the present invention, a preference data set corresponding to multiple preferences is selected for model training. During the training process, corresponding weights are assigned to different preferences to obtain a diversified speech synthesis model. Therefore, in actual application, according to the user's preferences, a model adapted to it can be selected for speech synthesis, and a synthesized speech that meets the user's expectations can be obtained, thereby fully meeting the diverse human preference needs.

[0061] It is understood that in actual applications, before performing speech synthesis, it is first necessary to determine the text to be synthesized. This text contains the text information corresponding to the desired synthesized speech. This text can be text directly entered by the user, text automatically generated by a computer during human-computer interaction, or text obtained by performing OCR (Optical Character Recognition) on an image acquired by an image acquisition device. This embodiment of the present invention does not specifically limit this. The image acquisition device here can be a scanner, mobile phone, camera, etc.

[0062] While determining the text to be synthesized, in order to make the synthesized speech more in line with user expectations, in an embodiment of the present invention, the user's speech synthesis preference can also be determined. The speech synthesis preference is used to represent the user's requirements / expectations for the synthesized speech obtained by this speech synthesis. For example, it can be a speech with a synthesized timbre as similar as possible, a speech that is more realistic and easier to understand, etc.

[0063] Furthermore, considering that the synthesized speech obtained by performing speech synthesis directly based on the text to be synthesized is often of poor quality, especially in terms of timbre and pronunciation habits, and has obvious differences from real speech, the speech synthesis effect is poor. To address this problem, in embodiments of the present invention, after obtaining the text to be synthesized and before performing speech synthesis, it is necessary to obtain a segment of real speech as a reference for the subsequent speech synthesis process. This real speech can be the user's voice or the voice of another person, and can be determined based on the usage scenario and the user's speech synthesis preferences, which are not specifically limited in embodiments of the present invention.

[0064] After this, speech synthesis can be performed to obtain synthesized speech that meets the user's expectations, that is, synthesized speech that matches the user's speech synthesis preferences. However, considering that current speech synthesis models only align with a single human preference and cannot meet the diverse human preference needs and synthesize synthesized speech that meets the user's expectations, in an embodiment of the present invention, a preference dataset corresponding to multiple preferences is pre-established and used for model training. Preference weight configuration is added during training. By assigning corresponding weights to different preferences, a variety of speech synthesis models can be trained. Based on this, speech synthesis can be achieved in various scenarios, which can meet the diverse human preference needs.

[0065] Specifically, here, based on the user's speech synthesis preference, a model that matches or conforms to the user's expectations / requirements can be selected from multiple pre-trained speech synthesis models as the target speech synthesis model. For example, when the user's speech synthesis preference is to synthesize highly natural and realistic speech, a speech synthesis model with a corresponding preference weight configuration that reflects a greater focus on naturalness and similarity can be selected from multiple speech synthesis models. Specifically, a model whose weights corresponding to the naturalness preference and similarity preference are higher than the weights of other preferences can be selected as the target speech synthesis model. For another example, when the user's speech synthesis preference is to synthesize speech that is easy to understand and read, a speech synthesis model that focuses on intelligibility can be selected from multiple speech synthesis models. Specifically, a model whose weight corresponding to the intelligibility preference is significantly higher than the weights of other preferences can be selected as the target speech synthesis model.

[0066] Afterwards, the target speech synthesis model can be applied to perform speech synthesis to obtain synthesized speech. Specifically, the target speech synthesis model can be applied to the text to be synthesized and real speech to perform speech synthesis, thereby obtaining synthesized speech that meets the user's speech synthesis preferences. Specifically, the text to be synthesized and the real speech are input into the target speech synthesis model, so that the target speech synthesis model references the real speech and performs speech synthesis according to the content of the text to be synthesized, obtaining synthesized speech that possesses the characteristics of real speech and is consistent with the content of the text to be synthesized, and outputting the synthesized speech, thereby obtaining synthesized speech that meets the user's speech synthesis preferences.

[0067] It is worth noting that before the text to be synthesized and the real speech are input into the target speech synthesis model, in an embodiment of the present invention, it is necessary to pre-train the speech synthesis model. In consideration of the problem that single-target preference alignment cannot meet the diverse preference requirements, in an embodiment of the present invention, a multi-target alignment method is adopted when training the speech synthesis model. Specifically, a preference data set corresponding to multiple preferences is first constructed. Here, the multiple preferences may include naturalness preference, intelligibility preference, similarity preference, completeness preference, fluency preference, etc. The preference data set corresponding to each preference contains multiple sets of preference data, and each set of preference data contains sample text, sample preferred speech and sample non-preferred speech corresponding to the preference. The sample preferred speech and sample non-preferred speech can be understood as positive samples and negative samples under the corresponding preference, specifically the speech that meets the corresponding preference and the speech that does not meet the corresponding preference corresponding to the sample text. The sample text can be understood as the text to be synthesized during the training process, which can be obtained from a public speech synthesis dataset.

[0068] After determining the preference datasets corresponding to multiple preferences, you can also assign a corresponding weight to each preference. The weight indicates the importance of that preference in the speech synthesis model training process. The larger the weight, the more emphasis the training process places on that particular capability. For example, the larger the weight for the naturalness preference, the more attention the model pays to the naturalness of the speech synthesis output during model training, with the goal of achieving more natural synthesized speech. Conversely, the lower the weight, the less emphasis the model places on that capability during training. By assigning different weights to different preferences, different weight combinations can help the model training focus on different aspects, thereby generating a variety of speech synthesis models. Each speech synthesis model corresponds to a unique set of weights, and different speech synthesis models use different weight combinations (the preference weight configuration composed of the weights corresponding to multiple preferences) when training.

[0069] In an embodiment of the present invention, based on preference data sets and preference weight configurations corresponding to multiple preferences, multiple speech synthesis models are trained to align multiple human preferences and meet diverse human preference needs in different scenarios. This not only avoids the problem of limited application and inability to meet diverse user needs of a single preference alignment solution, but also overcomes the defects of unstable training and conflicting speech synthesis in the multi-objective RLHF solution. Under the guidance of the preference weight configuration, the preference data sets corresponding to multiple preferences are used to train the model, so that the model can be optimized and updated in the direction guided by the weights. While highlighting one or more capabilities, the optimization and update of other items will not be ignored. Therefore, the quality of the synthesized speech can be guaranteed in the application stage, and the problem of poor quality of the synthesized speech can be avoided.

[0070] The speech synthesis method provided by the present invention selects a target speech synthesis model from multiple speech synthesis models according to the user's speech synthesis preference, and applies the target speech synthesis model to perform speech synthesis according to the text to be synthesized, so as to obtain synthesized speech that meets the speech synthesis preference; wherein, each speech synthesis model is trained based on the preference data set and preference weight configuration corresponding to multiple preferences, and the preference weight configurations used in the training of different speech synthesis models are different, which overcomes the defect that the human preference alignment method for speech synthesis models in traditional solutions can only align a single human preference and cannot meet the diverse human preference needs. The speech synthesis model is trained by a multi-objective direct preference optimization method, which not only makes the training process simpler and more efficient, but also can perform human preference alignment from multiple dimensions, and meet the diverse human preference needs by assigning weights to each preference.

[0071] Based on the above embodiment, the preference data set corresponding to each preference is determined based on the following steps:

[0072] Determine the sample text corresponding to the preference and the prompt voice corresponding to the sample text;

[0073] Determine multiple initial synthesized speech patterns corresponding to the sample text and the prompt speech patterns based on an initial speech synthesis model; the initial speech synthesis model is trained based on a pre-training dataset, the pre-training dataset including pre-trained text and corresponding pre-trained prompt speech patterns;

[0074] Determining a speech score of each initial synthesized speech under a corresponding preference, and determining a sample preferred speech and a sample non-preferred speech corresponding to the corresponding preference from each initial synthesized speech based on the speech score corresponding to each initial synthesized speech;

[0075] Based on the sample texts and prompt voices corresponding to the preferences, as well as the sample preferred voices and sample non-preferred voices, a preference dataset corresponding to the preferences is constructed.

[0076] Specifically, the preferred dataset used for model training can be determined by the following steps:

[0077] The following will take any preference as an example to illustrate the construction of the preference dataset:

[0078] First, you need to determine the sample text for your preference. This sample text can be obtained from a public speech synthesis dataset, and there can be multiple sample texts. After determining the sample text, you also need to determine the actual speech, which is referred to as the prompt speech corresponding to the sample text. This prompts the model during training to perform speech synthesis based on the sample text and output the synthesized speech.

[0079] Then, speech synthesis can be performed based on the sample text and prompt voice. Specifically, based on the sample text and prompt voice, the initial speech synthesis model is applied to perform speech synthesis, thereby obtaining multiple initial synthesized voices output by the initial speech synthesis model. Here, the initial speech synthesis model can be an initial model directly constructed for the training process, or it can be a model obtained after pre-training the initial model. In order to ensure the quality of the initial synthesized voice, in the embodiment of the present invention, the initial model is preferably pre-trained, specifically, the initial model is trained using the pre-training text and the corresponding pre-training prompt voice in the pre-training data set, and the pre-trained initial model is used as the initial speech synthesis model. The initial model is constructed on the Transformer architecture, its training method is autoregressive, and the loss function adopts cross entropy loss.

[0080] Pre-training specifically consists of two phases: model pre-training and supervised fine-tuning. The two phases utilize different training datasets. In short, the pre-training dataset includes training datasets from both phases, each containing pre-training text and corresponding pre-training prompts. The pre-training phase uses a large amount of data (tens of thousands of hours of speech data) to train the initial model, while the supervised fine-tuning phase uses datasets from downstream tasks (speech synthesis tasks in various scenarios) for supervised fine-tuning. After these two phases of training, a model with robust speech synthesis capabilities is obtained.

[0081] Afterwards, the speech score of each initial synthesized speech under the corresponding preference can be determined, that is, the score of each initial synthesized speech output by the initial speech synthesis model based on the preference can be determined. For example, when the preference is a naturalness preference, the score of each generated initial synthesized speech based on the naturalness preference needs to be determined. Here, the process of scoring each initial synthesized speech can be performed manually or implemented using a currently mature scoring algorithm. Furthermore, based on the speech score of each initial synthesized speech corresponding to the preference, sample preferred speech and sample non-preferred speech can be selected from each initial synthesized speech. Here, specifically, based on the speech score, the initial synthesized speech that meets the preference (e.g., the initial synthesized speech with a high score) can be selected as the sample preferred speech corresponding to the preference, and the initial synthesized speech that does not meet the preference (e.g., the initial synthesized speech with a low score) can be selected as the sample non-preferred speech corresponding to the preference.

[0082] Then, based on the selected sample preferred speech, sample non-preferred speech, and the pre-acquired sample text and corresponding prompt speech for the preference, a preference dataset corresponding to the preference can be constructed. It should be noted that the preference dataset corresponding to the preference often contains more than one set of sample texts and their corresponding data. Therefore, in the process of constructing the preference dataset, multiple sample texts can be obtained, and multiple initial synthesized speech can be generated for each sample text, so as to determine the sample preferred speech and sample non-preferred speech corresponding to each sample text, and then multiple sets of preference data can be constructed, based on which a complete preference dataset can be constructed.

[0083] Based on the above embodiment, the initial speech synthesis model is used to determine multiple initial synthesized speech patterns corresponding to the sample text and the prompt speech patterns, including:

[0084] Perform phoneme conversion on the sample text to obtain a sample phoneme sequence;

[0085] Perform feature encoding on the prompt voice to obtain sample voice features;

[0086] Embedding the sample phoneme sequence and the sample speech features, and concatenating the sample phoneme vector representation and the sample speech vector representation obtained based on the embedded coding to obtain a sample fusion vector representation;

[0087] Based on the sample fusion vector representation, an initial speech synthesis model is applied to perform speech synthesis to obtain multiple initial synthesized speech.

[0088] Specifically, the process of generating the initial synthesized speech may include:

[0089] Before inputting the data into the initial speech synthesis model for speech synthesis, it needs to be processed. Specifically, the sample text can be first subjected to phoneme conversion to convert the text sequence in the sample text into a phoneme sequence, thereby obtaining a sample phoneme sequence. Next, the prompt speech can be feature encoded, specifically using a pre-trained audio codec to extract discrete speech representations, namely, sample speech features.

[0090] Afterwards, the sample phoneme sequence and sample speech features can be embedded and encoded, and the sample phoneme vector representation and sample speech vector representation obtained by the embedded encoding can be concatenated to obtain a sample fusion vector representation. Specifically, the converted phoneme sequence and the extracted discrete speech representation are embedded and encoded, and the two are concatenated in the time series dimension. The concatenated features are used as input to the initial speech synthesis model for speech synthesis. In other words, the concatenated sample fusion vector representation can be input to the initial speech synthesis model, so that the initial speech synthesis model performs speech synthesis based on it and outputs multiple initial synthesized speech sounds.

[0091] Based on the above embodiment, determining a speech score of each initial synthesized speech under a corresponding preference, and determining a sample preferred speech and a sample non-preferred speech corresponding to the corresponding preference from each initial synthesized speech based on the speech score corresponding to each initial synthesized speech, includes:

[0092] determining a coarse-grained speech score for each initial synthesized speech under a plurality of preferences;

[0093] Based on the coarse-grained speech scores of the initial synthesized speech corresponding to the plurality of preferences, a plurality of candidate synthesized speech are screened from the initial synthesized speech;

[0094] Determining a fine-grained speech score for each candidate synthesized speech under corresponding preferences;

[0095] Based on the fine-grained speech scores corresponding to the candidate synthesized speech, a sample preferred speech and a sample non-preferred speech corresponding to the candidate synthesized speech are determined.

[0096] Specifically, the process of determining the speech score of each initial synthesized speech under the corresponding preference, and determining the sample preferred speech and the sample non-preferred speech corresponding to the corresponding preference from each initial synthesized speech based on the speech score, may specifically include:

[0097] After generating multiple initial synthesized speech sounds using the initial speech synthesis model, each initial synthesized speech sound can be scored. This scoring process can be divided into two steps. The first is a coarse scoring phase, in which each initial synthesized speech sound is scored across multiple dimensions to obtain its score under multiple preferences. This is to determine the coarse-grained speech score for each initial synthesized speech sound under multiple preferences. For example, multiple initial synthesized speech sounds generated under a naturalness preference can be scored under multiple preferences to obtain a score for each initial synthesized speech sound under naturalness, intelligibility, similarity, etc. This score is the coarse-grained speech score.

[0098] After that, data screening can be performed based on the coarse-grained speech score corresponding to each initial synthesized speech obtained from the coarse score, so as to filter out valid data from the generated initial synthesized speech; specifically, when the coarse-grained speech score corresponding to an initial synthesized speech indicates that it fails to meet the requirements in any one or more preferences among multiple preferences, that is, its coarse-grained speech score in this or these preferences is too low and lower than the set score threshold (for example, the full score is 5 points, and 3 points can be set as the score threshold), it is considered to be unqualified in this or these preferences. At this time, it can be regarded as invalid data and eliminated, thereby screening out candidate synthesized speech that meets all preferences from each initial synthesized speech.

[0099] Furthermore, after obtaining candidate synthesized speech, the second stage of scoring, namely, fine scoring, can begin. Specifically, the multiple candidate synthesized speech obtained in the previous step are scored based on the corresponding preferences. That is, if the preference dataset to be constructed is the preference dataset corresponding to the naturalness preference, then the multiple candidate synthesized speech can be scored based on the naturalness preference. Unlike the coarse scoring in the first stage, which targets all preferences, this stage only scores the desired preference dimension, resulting in a more accurate and detailed speech score, namely, a fine-grained speech score for each candidate synthesized speech under the corresponding preference. Then, based on this fine-grained speech score, sample preferred speech and sample non-preferred speech under the corresponding preference can be selected from each candidate synthesized speech. Specifically, based on the fine-grained speech score, candidate synthesized speech that meets the corresponding preference (e.g., candidate synthesized speech with a high fine-grained speech score) can be selected as the sample preferred speech under the corresponding preference, and candidate synthesized speech that does not meet the corresponding preference (e.g., candidate synthesized speech with a low fine-grained speech score) can be selected as the sample non-preferred speech.

[0100] Based on the above embodiment, the multiple preferences include naturalness preference, intelligibility preference, and similarity preference; based on the fine-grained speech score corresponding to each candidate synthesized speech, the sample preferred speech and the sample non-preferred speech corresponding to the corresponding preference are determined from each candidate synthesized speech, including:

[0101] From the candidate synthesized speech corresponding to the naturalness preference, the candidate synthesized speech with the highest and lowest fine-grained speech scores under the naturalness preference are selected as the sample preferred speech and sample non-preferred speech corresponding to the naturalness preference, respectively;

[0102] From the candidate synthesized speech corresponding to the intelligibility preference, the candidate synthesized speech with the highest and lowest fine-grained speech scores under the intelligibility preference are selected as the sample preferred speech and sample non-preferred speech corresponding to the intelligibility preference, respectively;

[0103] From the candidate synthesized speech corresponding to the similarity preference, the candidate synthesized speech with the highest and lowest fine-grained speech scores under the similarity preference are selected as the sample preferred speech and sample non-preferred speech corresponding to the similarity preference, respectively.

[0104] Specifically, the process of determining the corresponding preferred sample speech and the sample non-preferred speech from each candidate synthesized speech according to the fine-grained speech score includes:

[0105] In an embodiment of the present invention, the multiple preferences are naturalness preference, understanding preference and similarity preference. When determining the sample preferred speech and sample non-preferred speech under each preference, the fine-grained speech score under the preference can be directly used to select the highest-scoring candidate synthesized speech as the sample preferred speech and the lowest-scoring candidate as the sample non-preferred speech from multiple candidate synthesized speech. In this way, the sample preferred speech and sample non-preferred speech under each preference can be determined.

[0106] In detail, when determining the sample preferred speech and sample non-preferred speech under naturalness preference, the candidate synthesized speech with the highest fine-grained speech score under naturalness preference is directly selected from the candidate synthesized speech corresponding to the naturalness preference as the sample preferred speech corresponding to the naturalness preference. At the same time, the candidate synthesized speech with the lowest fine-grained speech score under naturalness preference is selected as the sample non-preferred speech corresponding to the naturalness preference.

[0107] When determining the sample preferred speech and sample non-preferred speech under the intelligibility preference, the candidate synthesized speech with the highest fine-grained speech score under the intelligibility preference is directly selected from the candidate synthesized speech corresponding to the intelligibility preference as the sample preferred speech corresponding to the intelligibility preference. At the same time, the candidate synthesized speech with the lowest fine-grained speech score under the intelligibility preference is selected as the sample non-preferred speech corresponding to the intelligibility preference.

[0108] When determining the sample preferred speech and sample non-preferred speech under similarity preference, the candidate synthesized speech with the highest fine-grained speech score under similarity preference is directly selected from the candidate synthesized speech corresponding to the similarity preference as the sample preferred speech corresponding to the similarity preference. At the same time, the candidate synthesized speech with the lowest fine-grained speech score under the similarity preference is selected as the sample non-preferred speech corresponding to the similarity preference.

[0109] It should be noted that, whether scoring the similarity preference in the coarse scoring stage or scoring the candidate synthesized speech under the similarity preference in the fine scoring stage, the data itself is not scored only, but the speech is compared with the prompt speech under the corresponding preference, and the similarity is judged by comparison.

[0110] In an embodiment of the present invention, a preference dataset is constructed from three dimensions: naturalness, intelligibility, and similarity. Based on the constructed preference dataset, model training is performed in combination with preference weight configuration, which can achieve multi-objective human preference alignment, thereby meeting the human preference needs of multiple samples in different scenarios.

[0111] Based on the above embodiment, each speech synthesis model is trained based on the following steps:

[0112] Selecting a preference data set corresponding to any one preference from the preference data sets corresponding to multiple preferences as a target preference data set;

[0113] Based on the initial speech synthesis model, determine the sample synthesized speech corresponding to the sample text and prompt speech in each group of preference data in the target preference data set;

[0114] Based on the reward model of other preferences, the reward score of the sample synthesized speech corresponding to each set of preference data under other preferences is determined; the reward model is trained based on the preference data set corresponding to other preferences;

[0115] Based on the sample preferred speech and sample non-preferred speech in each group of preference data, the sample synthesized speech corresponding to each group of preference data, the reward score of the sample synthesized speech corresponding to other preferences, and the preference weight configuration corresponding to the initial speech synthesis model, the parameters of the initial speech synthesis model are iterated to obtain the speech synthesis model.

[0116] Specifically, the application of reward models in single-preference alignment approaches also presents challenges. Current single-preference alignment schemes for speech synthesis models typically use speech recognition models, voiceprint recognition models, and other reward models. While speech recognition models can evaluate the intelligibility of synthesized speech, they often overlook mispronunciations in the generated speech, resulting in inaccurate and unreliable scoring. Voiceprint recognition models, while capable of distinguishing between different speakers, still struggle to differentiate between model-generated speech and real speech.

[0117] Based on this, in an embodiment of the present invention, it is proposed to use a constructed preference data set to train a reward model to overcome the defect that human preferences cannot be accurately reflected when speech recognition and voiceprint recognition models are directly used as reward models. The trained reward model is applied to the training process of the speech synthesis model to assist the training of the speech synthesis model, thereby better ensuring the model training effect and achieving multi-objective human preference alignment.

[0118] Specifically, here, any preference dataset corresponding to a plurality of preference datasets corresponding to the constructed preferences may be selected as the target preference dataset; that is, In the example, select any preferred dataset , as the target preference dataset, the target preference dataset is used to train the speech synthesis model, and other preference datasets are used to train the reward model. 、 and They represent the preference data set corresponding to naturalness preference, the preference data set corresponding to intelligibility preference, and the preference data set corresponding to similarity preference respectively.

[0119] Then, the initial speech synthesis model can be used to generate synthesized speech for each group of preference data in the target preference data set. That is, the initial speech synthesis model can be applied to perform speech synthesis based on the sample text and corresponding prompt speech in each group of preference data in the target preference data set, thereby obtaining the sample synthesized speech output by the initial speech synthesis model. Afterwards, the reward model corresponding to other preferences can be applied to score the sample synthesized speech corresponding to each group of preference data output by the initial speech synthesis model, and obtain the reward score of the sample synthesized speech corresponding to each group of preference data under other preferences. For example, when the target preference data set is When , it is necessary to first perform speech synthesis through the initial speech synthesis model to obtain The sample synthesis speech corresponding to each group of preference data in and The corresponding reward model is The sample synthesized speech corresponding to each group of preference data is scored to obtain The sample synthesized speech corresponding to each group of preference data is and Reward score corresponding to preference.

[0120] It is worth noting that before applying the reward model for other preferences for scoring, it is necessary to pre-train the reward model. Specifically, this can be done by applying the preference dataset corresponding to the other preferences to train the initial reward model, thereby obtaining the trained reward model for the other preferences. Here, the initial reward model is constructed based on the initial speech synthesis model and the linear projection layer. Specifically, a linear projection layer for classification is added to the initial speech synthesis model to construct the initial reward model. When training the initial reward model, the loss function used is a binary classification loss.

[0121] After that, the loss of the initial speech synthesis model on the speech synthesis task can be measured based on the sample preferred speech and sample non-preferred speech in each group of preference data in the target preference data set, the sample synthesized speech corresponding to each group of preference data, the reward score of the sample synthesized speech corresponding to other preferences, and the preference weight configuration corresponding to the initial speech synthesis model. The parameters of the initial speech synthesis model can be iterated based on this loss to obtain the speech synthesis model.

[0122] It should be noted that in the multi-objective direct preference optimization scheme provided by the present invention, during model training, the initial speech synthesis model can be set Set weight , each set of weights can be expressed as , That is, a set of weights, which is also the preference weight configuration of the model. 、 and The weights of the three preferences are respectively, and the weight of each preference in each set of weights is between 0 and 1, and the sum of the weights of the three preferences is 1. Different preference weight configurations can be trained to obtain Different speech synthesis models can better meet people's diverse preference needs.

[0123] It is worth noting that after selecting the preference dataset corresponding to one of the preferences as the target preference dataset, the preference datasets corresponding to other preferences are automatically used for reward model training. After the target preference dataset, the remaining preference datasets and Separately train the intelligibility reward model and the similarity reward model. No need to reselect them in reverse, that is, no need to reselect or is the target preference dataset to repeat the above training process.

[0124] Based on the above embodiment, based on the sample preferred speech and sample non-preferred speech in each set of preference data, the sample synthesized speech corresponding to each set of preference data, the reward score of the sample synthesized speech corresponding to other preferences, and the preference weight configuration corresponding to the initial speech synthesis model, the initial speech synthesis model is iterated on parameters to obtain a speech synthesis model, including:

[0125] Determining a first loss based on the sample preferred speech and the sample non-preferred speech in each set of preference data, the sample synthesized speech corresponding to each set of preference data, and the weight of the preference in the preference weight configuration corresponding to the initial speech synthesis model;

[0126] Determining a second loss based on the sample preferred speech and sample non-preferred speech in each set of preference data, the sample synthesized speech and its reward score corresponding to other preferences, and the weights of other preferences in the preference weight configuration corresponding to the initial speech synthesis model;

[0127] Based on the first loss and the second loss, parameters of the initial speech synthesis model are iterated to obtain a speech synthesis model.

[0128] Specifically, the process of iterating the parameters of the initial speech synthesis model to obtain the speech synthesis model may include:

[0129] In embodiments of the present invention, the overall loss of the initial speech synthesis model can be obtained through weighted optimization using multiple preferences. Specifically, the overall loss can be divided into two parts. One part is measured using the binary cross entropy loss (BCEloss). Specifically, based on the sample preferred and non-preferred speech samples in each preference data set within the target preference dataset, as well as the sample synthesized speech corresponding to each preference data set, the model's loss is determined, i.e., the first loss. The other part can be expressed as a "with margin" loss, i.e., the loss on the reward score, i.e., the second loss. Specifically, based on the sample preferred and non-preferred speech samples in each preference data set within the target preference dataset, as well as the sample synthesized speech samples and their corresponding reward scores for other preferences, the difference between the score of the sample synthesized speech output by the initial speech synthesis model and the score of the sample preferred and non-preferred speech samples by the reward model for other preferences is determined. Based on this, the weight of the other preferences in the preference weight configuration is used to calculate the second loss.

[0130] After that, the first loss and the second loss can be weighted to determine the overall loss, and the initial speech synthesis model can be updated according to the overall loss to obtain a speech synthesis model. Here, the parameters of the initial speech synthesis model can be adjusted according to the calculated overall loss, so that the sample synthesized speech output by the model after the parameter adjustment matches the corresponding preference weight configuration as much as possible, and is close to the corresponding sample preference speech, and away from the sample non-preferential speech. In this way, the model output can be aligned with the human preference represented by the corresponding preference weight configuration, and finally a trained speech synthesis model can be obtained. It should be noted that after the reward model is pre-trained using the preference data set corresponding to other preferences, the reward model is frozen during the training process of the initial speech synthesis model and does not participate in the update.

[0131] Figure 2 This is an example diagram of the framework of the model training process provided by the present invention, such as Figure 2 As shown, when When the target preference dataset is used, you can apply and Train the reward model for intelligibility separately and similarity reward model Furthermore, after obtaining the sample synthesized speech through the initial speech synthesis model, the output of the model, the sample preferred speech and the sample non-preferred speech, the reward score corresponding to other preferences of the model output, and the preference weight configuration can be combined. The first loss and the second loss are determined by BCE loss and with margin, and the overall loss of the model is measured accordingly to update the model parameters, and finally the speech synthesis model corresponding to the trained preference weight configuration is obtained.

[0132] In the embodiment of the present invention, during the training process, the above-mentioned method is used to optimize and update the model for each set of weights in the multiple preference weight configurations, so that multiple speech synthesis models aligned with human preferences with different weights in three dimensions can be obtained. Then, in the actual application process, according to different speech synthesis preferences, the corresponding speech synthesis model can be selected from the multiple trained speech synthesis models. For example, for users who have high requirements for the naturalness of synthesized speech, the preferred weight configuration can be selected. A large speech synthesis model can be used to perform speech synthesis, which can obtain synthesized speech that meets user expectations, meet the diverse speech synthesis needs of users in different scenarios, and has strong scene adaptability.

[0133] The speech synthesis device provided by the present invention is described below. The speech synthesis device described below and the speech synthesis method described above can be referenced to each other.

[0134] Figure 3 Schematic diagram of the structure of the speech synthesis device provided by the present invention. Figure 3 As shown, the device includes:

[0135] A data determination unit 310 is used to determine the text to be synthesized and the user's speech synthesis preference;

[0136] A speech synthesis unit 320 is configured to select a target speech synthesis model from a plurality of speech synthesis models based on the speech synthesis preference, and apply the target speech synthesis model to perform speech synthesis based on the text to be synthesized to obtain synthesized speech that meets the speech synthesis preference;

[0137] Among them, each speech synthesis model is trained based on preference data sets and preference weight configurations corresponding to multiple preferences. The preference weight configurations used in training different speech synthesis models are different. Each preference data set contains multiple sets of preference data corresponding to preferences, and each set of preference data includes sample text corresponding to the preference, sample preference speech and sample non-preferential speech.

[0138] The speech synthesis device provided by the present invention selects a target speech synthesis model from multiple speech synthesis models according to the user's speech synthesis preference, and applies the target speech synthesis model to perform speech synthesis according to the text to be synthesized, so as to obtain synthesized speech that meets the speech synthesis preference; wherein, each speech synthesis model is trained based on the preference data set and preference weight configuration corresponding to multiple preferences, and the preference weight configurations used in the training of different speech synthesis models are different, which overcomes the defect that the human preference alignment method for speech synthesis models in traditional solutions can only align a single human preference and cannot meet the diverse human preference needs. The speech synthesis model is trained by a multi-objective direct preference optimization method, which not only makes the training process simpler and more efficient, but also can perform human preference alignment from multiple dimensions, and meet the diverse human preference needs by assigning weights to each preference.

[0139] Based on the above embodiment, the apparatus further includes a data set construction unit, configured to:

[0140] Determining a sample text corresponding to a preference and a prompt voice corresponding to the sample text;

[0141] Determining, based on an initial speech synthesis model, a plurality of initial synthesized speech sounds corresponding to the sample text and the prompt speech sound; wherein the initial speech synthesis model is trained based on a pre-training data set, and the pre-training data set includes pre-trained text and corresponding pre-trained prompt speech sounds;

[0142] Determining a speech score of each initial synthesized speech under a corresponding preference, and determining, from each initial synthesized speech, a sample preferred speech and a sample non-preferred speech corresponding to the corresponding preference based on the speech score corresponding to each initial synthesized speech;

[0143] Based on the sample text and prompt voice corresponding to the preference, and the corresponding sample preferred voice and the sample non-preferred voice, a preference data set corresponding to the preference is constructed.

[0144] Based on the above embodiment, the data set construction unit is used to:

[0145] determining a coarse-grained speech score of each of the initial synthesized speech under the plurality of preferences;

[0146] Based on the coarse-grained speech scores of the initial synthesized speech corresponding to the plurality of preferences, screening the initial synthesized speech to obtain a plurality of candidate synthesized speech;

[0147] Determining a fine-grained speech score for each candidate synthesized speech under corresponding preferences;

[0148] Based on the fine-grained speech scores corresponding to the candidate synthesized speech, sample preferred speech and sample non-preferred speech corresponding to the candidate synthesized speech are determined from the candidate synthesized speech.

[0149] Based on the above embodiment, the multiple preferences include naturalness preference, understanding preference and similarity preference; the data set construction unit is used to:

[0150] From the candidate synthesized speech corresponding to the naturalness preference, selecting the candidate synthesized speech with the highest and lowest fine-grained speech scores under the naturalness preference as the sample preferred speech and the sample non-preferred speech corresponding to the naturalness preference, respectively;

[0151] From the candidate synthesized speech corresponding to the intelligibility preference, selecting the candidate synthesized speech with the highest and lowest fine-grained speech scores under the intelligibility preference as the sample preferred speech and the sample non-preferred speech corresponding to the intelligibility preference, respectively;

[0152] From the candidate synthesized speech corresponding to the similarity preference, the candidate synthesized speech with the highest and lowest fine-grained speech scores under the similarity preference are selected as the sample preferred speech and the sample non-preferred speech corresponding to the similarity preference, respectively.

[0153] Based on the above embodiment, the data set construction unit is used to:

[0154] Performing phoneme conversion on the sample text to obtain a sample phoneme sequence;

[0155] Performing feature encoding on the prompt voice to obtain sample voice features;

[0156] Embedding the sample phoneme sequence and the sample speech feature, and concatenating the sample phoneme vector representation and the sample speech vector representation obtained based on the embedded coding to obtain a sample fusion vector representation;

[0157] Based on the sample fusion vector representation, the initial speech synthesis model is applied to perform speech synthesis to obtain multiple initial synthesized speech.

[0158] Based on the above embodiment, the device further includes a model training unit, which is used to:

[0159] Selecting a preference data set corresponding to any one preference from the plurality of preference data sets corresponding to the preferences as a target preference data set;

[0160] Determining, based on the initial speech synthesis model, sample synthesized speech corresponding to the sample text and prompt speech in each group of preference data in the target preference data set;

[0161] Determining a reward score for the sample synthesized speech corresponding to each set of preference data under the other preferences based on a reward model for other preferences; the reward model is trained based on the preference data set corresponding to the other preferences;

[0162] Based on the sample preferred speech and sample non-preferred speech in each group of preference data, the sample synthesized speech corresponding to each group of preference data, the reward score of the sample synthesized speech corresponding to other preferences, and the preference weight configuration corresponding to the initial speech synthesis model, the initial speech synthesis model is iterated on parameters to obtain a speech synthesis model.

[0163] Based on the above embodiment, the model training unit is used to:

[0164] Determining a first loss based on the sample preferred speech and the sample non-preferred speech in each set of preference data, the sample synthesized speech corresponding to each set of preference data, and the weight of any preference in the preference weight configuration corresponding to the initial speech synthesis model;

[0165] Determining a second loss based on the sample preferred speech and sample non-preferred speech in each set of preference data, the sample synthesized speech and its reward score corresponding to other preferences, and the weight of the other preferences in the preference weight configuration corresponding to the initial speech synthesis model;

[0166] Based on the first loss and the second loss, parameters of the initial speech synthesis model are iterated to obtain a speech synthesis model.

[0167] Figure 4 An example of a physical structure diagram of an electronic device is shown below. Figure 4 As shown, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communication bus 440, wherein the processor 410, the communications interface 420, and the memory 430 communicate with each other via the communication bus 440. The processor 410 may call logic instructions in the memory 430 to execute a speech synthesis method, which includes: determining a text to be synthesized and a user's speech synthesis preferences; based on the speech synthesis preferences, selecting a target speech synthesis model from multiple speech synthesis models, and applying the target speech synthesis model to perform speech synthesis based on the text to be synthesized to obtain synthesized speech that meets the speech synthesis preferences; wherein each speech synthesis model is trained based on a preference data set and a preference weight configuration corresponding to multiple preferences, wherein different speech synthesis models are trained using different preference weight configurations, and each preference data set contains multiple sets of preference data corresponding to a preference, each set of preference data including sample text, sample preferred speech, and sample non-preferred speech corresponding to the preference.

[0168] Furthermore, the logic instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0169] On the other hand, the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the speech synthesis method provided by the above methods, the method including: determining the text to be synthesized and the user's speech synthesis preference; based on the speech synthesis preference, selecting a target speech synthesis model from multiple speech synthesis models, and applying the target speech synthesis model to perform speech synthesis based on the text to be synthesized to obtain a synthesized speech that meets the speech synthesis preference; wherein each speech synthesis model is trained based on a preference data set and preference weight configuration corresponding to multiple preferences, and the preference weight configurations used in training different speech synthesis models are different, and each preference data set contains multiple sets of preference data corresponding to preferences, and each set of preference data includes sample text, sample preference speech and sample non-preferential speech corresponding to the preference.

[0170] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the speech synthesis method provided by the above-mentioned methods, the method comprising: determining the text to be synthesized and the user's speech synthesis preference; based on the speech synthesis preference, selecting a target speech synthesis model from a plurality of speech synthesis models, and applying the target speech synthesis model to perform speech synthesis based on the text to be synthesized to obtain a synthesized speech that meets the speech synthesis preference; wherein each speech synthesis model is trained based on a preference data set and a preference weight configuration corresponding to a plurality of preferences, different speech synthesis models are trained using different preference weight configurations, each preference data set contains multiple sets of preference data corresponding to preferences, and each set of preference data includes sample text, sample preference speech and sample non-preferential speech corresponding to the preference.

[0171] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0172] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0173] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A speech synthesis method, characterized in that: include: Determine the text to be synthesized and the user's speech synthesis preferences; Based on the speech synthesis preference, a target speech synthesis model is selected from a plurality of speech synthesis models, and based on the text to be synthesized, the target speech synthesis model is applied to perform speech synthesis to obtain synthesized speech that meets the speech synthesis preference; Among them, each speech synthesis model is trained based on a plurality of preference data sets and preference weight configurations corresponding to preferences. Different speech synthesis models are trained using different preference weight configurations. Each preference data set contains multiple sets of preference data corresponding to preferences. Each set of preference data includes sample text corresponding to the preference, sample preferred speech, and sample non-preferred speech. The sample preferred speech and sample non-preferred speech under the corresponding preference are determined based on the following steps: Determining a sample text corresponding to a preference and a prompt voice corresponding to the sample text; Perform speech synthesis based on the sample text and the prompt speech to obtain multiple initial synthesized speech; determining a coarse-grained speech score of each initial synthesized speech under the plurality of preferences; Based on the coarse-grained speech scores of the initial synthesized speech corresponding to the plurality of preferences, screening the initial synthesized speech to obtain a plurality of candidate synthesized speech; Determining a fine-grained speech score for each candidate synthesized speech under corresponding preferences; Based on the fine-grained speech scores corresponding to the candidate synthesized speech, sample preferred speech and sample non-preferred speech corresponding to the candidate synthesized speech are determined from the candidate synthesized speech.

2. The speech synthesis method according to claim 1, wherein: The performing speech synthesis based on the sample text and the prompt speech to obtain a plurality of initial synthesized speech sounds includes: Determining, based on an initial speech synthesis model, a plurality of initial synthesized speech sounds corresponding to the sample text and the prompt speech sound; wherein the initial speech synthesis model is trained based on a pre-training data set, and the pre-training data set includes pre-trained text and corresponding pre-trained prompt speech sounds; The preference dataset corresponding to the preference is constructed based on the following steps: Based on the sample text and prompt voice corresponding to the preference, and the sample preferred voice and the sample non-preferred voice corresponding to the preference, a preference data set corresponding to the preference is constructed.

3. The speech synthesis method according to claim 1, wherein: The multiple preferences include a naturalness preference, an intelligibility preference, and a similarity preference; and determining, based on the fine-grained speech scores corresponding to the candidate synthesized speech, sample preferred speech and sample non-preferred speech corresponding to the corresponding preferences from the candidate synthesized speech, including: From the candidate synthesized speech corresponding to the naturalness preference, selecting the candidate synthesized speech with the highest and lowest fine-grained speech scores under the naturalness preference as the sample preferred speech and the sample non-preferred speech corresponding to the naturalness preference, respectively; From the candidate synthesized speech corresponding to the intelligibility preference, selecting the candidate synthesized speech with the highest and lowest fine-grained speech scores under the intelligibility preference as the sample preferred speech and the sample non-preferred speech corresponding to the intelligibility preference, respectively; From the candidate synthesized speech corresponding to the similarity preference, the candidate synthesized speech with the highest and lowest fine-grained speech scores under the similarity preference are selected as the sample preferred speech and the sample non-preferred speech corresponding to the similarity preference, respectively.

4. The speech synthesis method according to claim 2, wherein: The determining, based on the initial speech synthesis model, a plurality of initial synthesized speech corresponding to the sample text and the prompt speech includes: Performing phoneme conversion on the sample text to obtain a sample phoneme sequence; Performing feature encoding on the prompt voice to obtain sample voice features; Embedding the sample phoneme sequence and the sample speech feature, and concatenating the sample phoneme vector representation and the sample speech vector representation obtained based on the embedded coding to obtain a sample fusion vector representation; Based on the sample fusion vector representation, the initial speech synthesis model is applied to perform speech synthesis to obtain multiple initial synthesized speech.

5. The speech synthesis method according to claim 2, wherein: Each speech synthesis model is trained based on the following steps: Selecting a preference data set corresponding to any one preference from the plurality of preference data sets corresponding to the preferences as a target preference data set; Determining, based on the initial speech synthesis model, sample synthesized speech corresponding to the sample text and prompt speech in each group of preference data in the target preference data set; Based on the reward model of other preferences, determining the reward score of the sample synthesized speech corresponding to each set of preference data under the other preferences; The reward model is trained based on the preference data set corresponding to the other preferences; Based on the sample preferred speech and sample non-preferred speech in each group of preference data, the sample synthesized speech corresponding to each group of preference data, the reward score of the sample synthesized speech corresponding to other preferences, and the preference weight configuration corresponding to the initial speech synthesis model, the initial speech synthesis model is iterated on parameters to obtain a speech synthesis model.

6. The speech synthesis method according to claim 5, characterized in that: The method of performing parameter iteration on the initial speech synthesis model based on the sample preferred speech and sample non-preferred speech in each set of preference data, the sample synthesized speech corresponding to each set of preference data, the reward score of the sample synthesized speech corresponding to other preferences, and the preference weight configuration corresponding to the initial speech synthesis model to obtain the speech synthesis model includes: Determining a first loss based on the sample preferred speech and the sample non-preferred speech in each set of preference data, the sample synthesized speech corresponding to each set of preference data, and the weight of any preference in the preference weight configuration corresponding to the initial speech synthesis model; Determining a second loss based on the sample preferred speech and sample non-preferred speech in each set of preference data, the sample synthesized speech and its reward score corresponding to other preferences, and the weight of the other preferences in the preference weight configuration corresponding to the initial speech synthesis model; Based on the first loss and the second loss, parameters of the initial speech synthesis model are iterated to obtain a speech synthesis model.

7. A speech synthesis device, characterized in that: include: a data determination unit, configured to determine the text to be synthesized and the user's speech synthesis preference; a speech synthesis unit, configured to select a target speech synthesis model from a plurality of speech synthesis models based on the speech synthesis preference, and perform speech synthesis based on the text to be synthesized by applying the target speech synthesis model to obtain synthesized speech that meets the speech synthesis preference; Among them, each speech synthesis model is trained based on a plurality of preference data sets and preference weight configurations corresponding to preferences. Different speech synthesis models are trained using different preference weight configurations. Each preference data set contains multiple sets of preference data corresponding to preferences. Each set of preference data includes sample text corresponding to the preference, sample preferred speech, and sample non-preferred speech. The sample preferred speech and sample non-preferred speech under the corresponding preference are determined based on the following steps: Determining a sample text corresponding to a preference and a prompt voice corresponding to the sample text; Perform speech synthesis based on the sample text and the prompt speech to obtain multiple initial synthesized speech; determining a coarse-grained speech score of each initial synthesized speech under the plurality of preferences; Based on the coarse-grained speech scores of the initial synthesized speech corresponding to the plurality of preferences, screening the initial synthesized speech to obtain a plurality of candidate synthesized speech; Determining a fine-grained speech score for each candidate synthesized speech under corresponding preferences; Based on the fine-grained speech scores corresponding to the candidate synthesized speech, sample preferred speech and sample non-preferred speech corresponding to the candidate synthesized speech are determined from the candidate synthesized speech.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the speech synthesis method according to any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the speech synthesis method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Speech synthesis method and device, electronic equipment and storage medium

    CN116704998A

  • Personalized custom synthetic speech

    US20200265829A1