A speech synthesis method, apparatus, device, and storage medium

By acquiring text features and speaker voice data, and combining them with user-adjusted parameters to synthesize user-customized voices, the problem of a single voice tone not meeting user needs is solved, thus realizing personalized voice synthesis and improving user experience and enjoyment.

CN114708847BActive Publication Date: 2025-11-11UNIV OF SCI & TECH OF CHINA +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210252704.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-15
Publication Date
2025-11-11
Estimated Expiration
2042-03-15

AI Technical Summary

Technical Problem

Existing speech synthesis solutions can only synthesize a single fixed timbre, which cannot meet users' personalized needs for synthesized speech timbre.

Method used

By acquiring text features and the speech data of a specified speaker, speaker features are extracted as the original timbre feature vector. Combined with the timbre adjustment parameters and timbre stretching vector determined by the user, the timbre feature vector of the user-customized timbre is determined, and finally, the speech data of the user-customized timbre is synthesized.

Benefits of technology

It synthesizes voice data that matches the user's personal preferences, enhancing the user experience and the fun and interactivity of voice synthesis, and satisfying the user's personalized needs for synthesized voice timbre.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114708847B_ABST
    Figure CN114708847B_ABST
Patent Text Reader

Abstract

This invention provides a speech synthesis method, apparatus, device, and storage medium. The speech synthesis method includes: acquiring text features for speech synthesis and speech data of a specified speaker; extracting speaker features from the specified speaker's speech data as an original timbre feature vector; determining a timbre feature vector for a user-customized timbre based on the original timbre feature vector, user-determined timbre adjustment parameters under a set timbre dimension, and a timbre stretching vector under the set timbre dimension; and synthesizing speech data with the user-customized timbre based on the text features and the timbre feature vector of the user-customized timbre. The speech synthesis method provided by this invention can meet users' personalized needs for synthesized speech timbre. Furthermore, by allowing users to deeply participate in the selection of synthesized speech timbre, speech synthesis becomes more interesting and interactive, thereby improving the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech synthesis technology, and in particular to a speech synthesis method, apparatus, device and storage medium. Background Technology

[0002] Speech synthesis is an intelligent voice technology that converts text into speech, and it is one of the key technologies for realizing human-computer interaction. The timbre of the synthesized speech is one of the factors that affects the user's experience with speech synthesis products. A synthesized speech timbre that matches user preferences can bring a good product experience and enhance product value.

[0003] However, current speech synthesis solutions can only synthesize speech data with a single fixed timbre. It is understandable that different users usually have different preferences for synthesized speech timbre, and a single fixed timbre cannot be liked by all users. It is clear that current speech synthesis solutions cannot meet users' personalized needs for synthesized speech timbre. Summary of the Invention

[0004] In view of this, this application provides a speech synthesis method, apparatus, device, and storage medium to solve the problem that existing speech synthesis schemes can only synthesize speech data with a single fixed timbre, failing to meet users' personalized needs for synthesized speech timbre. The technical solution is as follows:

[0005] A speech synthesis method, comprising:

[0006] Acquire text features for speech synthesis and speech data of a specified speaker;

[0007] Speaker features are extracted from the speech data of the specified speaker and used as the original timbre feature vector.

[0008] Based on the original timbre feature vector, the timbre adjustment parameters determined by the user under the set timbre dimension, and the timbre stretching vector under the set timbre dimension, the timbre feature vector of the user-customized timbre is determined.

[0009] Based on the text features and the timbre feature vector of the user-customized timbre, speech data with the user-customized timbre is synthesized.

[0010] Optionally, determining the timbre feature vector of the user-customized timbre based on the original timbre feature vector, the timbre adjustment parameters determined by the user under the set timbre dimension, and the timbre stretching vector under the set timbre dimension includes:

[0011] Based on the timbre adjustment parameters under the set timbre dimension and the timbre stretching vector under the set timbre dimension, determine the timbre feature adjustment vector under the set timbre dimension;

[0012] The timbre feature vector of the user-customized timbre is determined based on the original timbre feature vector and the timbre feature adjustment vector under the set timbre dimension.

[0013] Optionally, the step of synthesizing speech data with a user-customized timbre based on the text features and the timbre feature vector of the user-customized timbre includes:

[0014] Based on the text features and the speaker features, obtain a frame-level feature vector sequence;

[0015] Based on the frame-level feature vector sequence and the timbre feature vector of the user-customized timbre, speech data with the user-customized timbre is synthesized.

[0016] Optionally, obtaining the frame-level feature vector sequence based on the text features and the speaker features includes:

[0017] Based on the text features, obtain the phoneme-level context feature vector;

[0018] Based on the phoneme-level context feature vector and the speaker features, predict the number of frames for the pronunciation duration of the phoneme;

[0019] Based on the number of frames for the pronunciation duration of the phoneme, the context feature vector at the phoneme level is expanded into a sequence of feature vectors at the frame level.

[0020] Optionally, the setting of timbre dimensions includes one or more timbre dimensions;

[0021] Determining the timbre stretch vector under the specified timbre dimension includes:

[0022] Obtain a speech data set, which includes speech data from multiple speakers, and each speaker's speech data is labeled with a timbre attribute value under the set timbre dimension;

[0023] Each timbre dimension included in the set timbre dimension is taken as the target timbre dimension. Based on the timbre attribute values ​​of the voice data labeled in the voice data set, voice data related to the target timbre dimension is obtained from the voice data set.

[0024] Based on the speech data related to the target timbre dimension, determine the timbre stretching vector under the target timbre dimension.

[0025] Optionally, the process of annotating the timbre attribute values ​​of a speaker's speech data under the defined timbre dimension includes:

[0026] Obtain the timbre attribute values ​​marked by multiple annotators on the timbre attributes under the set timbre dimension based on the speaker's speech data, so as to obtain the annotation results of multiple annotators;

[0027] Based on the annotation results of the multiple annotators, the timbre attribute value of the speaker's voice data under the set timbre dimension is determined, and the determined timbre attribute value is annotated for the speaker's voice data.

[0028] Optionally, the step of obtaining speech data related to the target timbre dimension from the speech data set based on the timbre attribute values ​​labeled in the speech data set includes:

[0029] Speech data with the attribute value of the first timbre attribute and the attribute value of the first target attribute are obtained from the speech data set to form a first speech dataset, wherein the first timbre attribute is a timbre attribute related to the target timbre dimension, and the first target attribute value is an attribute value related to the target timbre dimension;

[0030] Speech data with the attribute value of the second timbre attribute and the attribute value of the second target attribute are obtained from the speech data set to form a second speech dataset, wherein the second timbre attribute is a timbre attribute related to the first timbre attribute, and the second target attribute value is an attribute value related to the first timbre attribute.

[0031] Optionally, determining the timbre stretching vector under the target timbre dimension based on the speech data related to the target timbre dimension includes:

[0032] Speaker features are extracted from the speech data of each speaker in the first speech dataset, and the mean of several speaker features extracted from the speech data in the first speech dataset is calculated. The calculated mean is used as the first average timbre vector under the target timbre dimension.

[0033] Speaker features are extracted from the speech data of each speaker in the second speech dataset, and the mean of several speaker features extracted from the speech data in the second speech dataset is calculated. The calculated mean is used as the second average timbre vector under the target timbre dimension.

[0034] The timbre stretching vector under the target timbre dimension is determined based on the first average timbre vector and the second average timbre vector under the target timbre dimension.

[0035] Optionally, extracting speaker features from the speech data of the specified speaker includes:

[0036] Using a pre-built speech generation module, speaker features are extracted from the speech data of the specified speaker;

[0037] The process of synthesizing user-customized voice data based on the text features and the timbre feature vector of the user-customized voice includes:

[0038] Based on the speech generation module, the text features, and the timbre feature vector of the user-customized timbre, speech data with the user-customized timbre is synthesized.

[0039] Optionally, the speech generation module is a speech synthesis model, which is trained using multiple training speech data from multiple speakers and training texts corresponding to the multiple training speech data.

[0040] The process of synthesizing user-customized voice data based on the speech generation module, the text features, and the timbre feature vector of the user-customized voice includes:

[0041] The text encoding module based on the speech synthesis model encodes the text features to obtain a phoneme-level context feature vector.

[0042] The duration prediction module based on the speech synthesis model predicts the number of frames for the pronunciation of a phoneme based on the context feature vector at the phoneme level and the speaker features.

[0043] Based on the length adjustment module of the speech synthesis model, the phoneme-level context feature vector is expanded into a frame-level feature vector sequence according to the number of frames of pronunciation duration of the phoneme.

[0044] The decoding module based on the speech synthesis model predicts spectral features based on the frame-level feature vector sequence and the timbre feature vector of the user-customized timbre.

[0045] Based on the spectral characteristics, user-customized voice data is synthesized.

[0046] Optionally, the step of synthesizing speech data with a user-customized timbre based on the speech synthesis model, the text features, and the timbre feature vector of the user-customized timbre includes:

[0047] The text encoding module based on the speech synthesis model encodes the text features to obtain a phoneme-level context feature vector.

[0048] The duration prediction module based on the speech synthesis model predicts the number of frames for the pronunciation of a phoneme based on the context feature vector at the phoneme level and the speaker features.

[0049] Based on the length adjustment module of the speech synthesis model, the phoneme-level context feature vector is expanded into a frame-level feature vector sequence according to the number of frames of pronunciation duration of the phoneme.

[0050] Based on the decoding module in the speech synthesis model, the spectral features are predicted according to the frame-level feature vector sequence and the timbre feature vector of the user-customized timbre.

[0051] Based on the spectral characteristics, user-customized voice data is synthesized.

[0052] A speech synthesis device includes: a data acquisition module, a speaker feature extraction module, a timbre feature vector determination module, and a speech synthesis module;

[0053] The data acquisition module is used to acquire text features for speech synthesis and speech data of a specified speaker.

[0054] The speaker feature extraction module is used to extract speaker features from the speech data of the specified speaker as the original timbre feature vector;

[0055] The timbre feature vector determination module is used to determine the timbre feature vector of the user-customized timbre based on the original timbre feature vector, the timbre adjustment parameters determined by the user under the set timbre dimension, and the timbre stretching vector under the set timbre dimension.

[0056] The speech synthesis module is used to synthesize speech data with a user-customized timbre based on the text features and the timbre feature vector of the user-customized timbre.

[0057] A speech synthesis device includes: a memory and a processor;

[0058] The memory is used to store programs;

[0059] The processor is configured to execute the program to implement each step of the speech synthesis method described above.

[0060] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the preceding speech synthesis methods.

[0061] The speech synthesis method, apparatus, device, and storage medium provided in this application first acquire text features for speech synthesis and speech data of a specified speaker. Then, speaker features are extracted from the specified speaker's speech data to obtain an original timbre feature vector. Next, based on the original timbre feature vector, user-determined timbre adjustment parameters under a set timbre dimension, and a timbre stretching vector under the set timbre dimension, a timbre feature vector for the user-customized timbre is determined. Finally, based on the text features and the timbre feature vector of the user-customized timbre, speech data with the user-customized timbre is synthesized. The speech synthesis method provided in this application can synthesize speech data with deeply customized user timbres, resulting in speech data that better matches user preferences. Therefore, the speech synthesis method provided in this application can meet users' personalized needs for synthesized speech timbres. Furthermore, allowing users to deeply participate in the selection of synthesized speech timbres makes speech synthesis more interesting and interactive, thereby improving the user experience. Attached Figure Description

[0062] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0063] Figure 1 A schematic flowchart illustrating the speech synthesis method provided in this application embodiment;

[0064] Figure 2 A schematic diagram illustrating the process of implementing speech synthesis based on a speech synthesis model, as provided in an embodiment of this application;

[0065] Figure 3 A schematic diagram of the structure of the speech synthesis model provided in the embodiments of this application;

[0066] Figure 4 A schematic diagram illustrating the process of determining the timbre stretching vector under a set timbre dimension, provided in an embodiment of this application;

[0067] Figure 5 This is a schematic diagram of the structure of the speech synthesis device provided in the embodiments of this application;

[0068] Figure 6 This is a schematic diagram of the structure of the speech synthesis device provided in the embodiments of this application. Detailed Implementation

[0069] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0070] Since current speech synthesis schemes can only synthesize speech data with a single fixed timbre, they cannot meet users' personalized needs for synthesized speech timbre. In view of this, the applicant conducted research. The initial idea was to provide users with a sound library with multiple timbres in a scenario, allowing users to choose according to their personal preferences. However, in practical applications, the available timbres are usually limited, which means that users have a limited range of timbre choices and can only passively accept the provided timbres.

[0071] To address the shortcomings of the aforementioned approach, the applicant realized that allowing users to deeply participate in the selection of synthesized speech timbres (i.e., enabling users to deeply customize their preferred synthesized speech timbres) would greatly enhance the user experience. This deep user involvement would, on the one hand, result in synthesized speech that better matches the user's personal preferences, and on the other hand, increase the fun and interactivity of the speech synthesis product. Following this line of thought, the applicant continued its research and, through continuous efforts, ultimately provided a speech synthesis method capable of synthesizing speech with user-customized timbres.

[0072] The speech synthesis method provided in this application can be applied to electronic devices with processing capabilities. These devices can be network-side servers (which can be a single server, a server cluster consisting of multiple servers, or a cloud computing server center) or user-side terminals (terminals can be, but are not limited to, PCs, laptops, smartphones, in-vehicle terminals, smart home devices, wearable devices, etc.). The network-side server or the user-side terminal can synthesize speech with a timbre more suited to the user's personal preferences using the speech synthesis method provided in this application. Those skilled in the art should understand that the servers and terminals listed above are merely examples. Other existing or future servers and terminals that are applicable to this application should also be included within the scope of protection of this application and are hereby incorporated by reference.

[0073] The speech synthesis method provided in this application will be described below through the following embodiments.

[0074] First Embodiment

[0075] Please see Figure 1 The diagram illustrates a flowchart of a speech synthesis method provided in an embodiment of this application. The method may include:

[0076] Step S101: Obtain text features for speech synthesis and speech data of the specified speaker.

[0077] The text features used for speech synthesis are obtained from the text used for speech synthesis. These text features may include phoneme information, tone information, and prosodic word segmentation information. It should be noted that tone information and prosodic word segmentation information are both phoneme-level information, obtained by processing word-level prosodic word segmentation information to the phoneme level. In general, the text features used for speech synthesis in this embodiment contain phoneme-level information.

[0078] Step S102: Extract speaker features from the speech data of the specified speaker as the original timbre feature vector.

[0079] In this embodiment, the speaker features extracted from the speech data of a specified speaker are used as the original timbre feature vector, which is the feature vector of the original timbre of the specified speaker.

[0080] Step S103: Based on the original timbre feature vector, the timbre adjustment parameters determined by the user under the set timbre dimension, and the timbre stretching vector under the set timbre dimension, determine the timbre feature vector of the user-customized timbre.

[0081] The timbre dimension setting is a user-adjustable timbre dimension, which can be set according to a person's perception of timbre. The timbre dimension setting can include one or multiple timbre dimensions, and the specific timbre dimensions and their number can be set according to the specific application scenario. For example, the timbre dimension setting can include a gender-independent "nasal quality" dimension, a gender-related "sweet female voice" dimension, and a "deep male voice" dimension.

[0082] The timbre stretching vector under the timbre dimension is the basic vector used to adjust the timbre. The timbre stretching vector under the timbre dimension is predetermined, that is, determined before the actual speech synthesis. The method of determining the timbre stretching vector under the timbre dimension will be introduced in subsequent embodiments.

[0083] Among them, the timbre adjustment parameters are used to adjust the timbre. They are determined by the user. Optionally, a timbre adjustment interface can be displayed. The user can change the timbre adjustment parameters based on the timbre adjustment interface, and then adjust the timbre based on the timbre adjustment parameters in combination with the timbre stretching vector.

[0084] Step S104: Based on text features and the timbre feature vector of user-customized timbre, synthesize speech data with user-customized timbre.

[0085] Optionally, the process of synthesizing user-customized voice data based on text features and the timbre feature vector of the user-customized voice may include: obtaining a frame-level feature vector sequence based on text features and speaker features; and synthesizing user-customized voice data based on the frame-level feature vector sequence and the timbre feature vector of the user-customized voice.

[0086] The process of obtaining a frame-level feature vector sequence based on text features and speaker features includes: obtaining a phoneme-level context feature vector based on text features; predicting the number of frames for the pronunciation duration of a phoneme based on the phoneme-level context feature vector and the speaker features; and expanding the phoneme-level context feature vector into a frame-level feature vector sequence based on the number of frames for the pronunciation duration of a phoneme.

[0087] Since the final synthesized speech data is based on the timbre feature vector of the user-customized timbre, speech data with the user-customized timbre can be synthesized based on text features and the timbre feature vector of the user-customized timbre. The speech data with the user-customized timbre is speech data with a timbre that matches the user's personal preferences.

[0088] The speech synthesis method provided in this application first acquires text features for speech synthesis and speech data of a specified speaker. Then, it extracts speaker features from the specified speaker's speech data to obtain an original timbre feature vector. Next, based on the original timbre feature vector, user-determined timbre adjustment parameters under a set timbre dimension, and a timbre stretching vector under the set timbre dimension, it determines a timbre feature vector for a user-customized timbre. Finally, based on the text features and the timbre feature vector of the user-customized timbre, it synthesizes speech data with the user-customized timbre. The speech synthesis method provided in this application can synthesize speech data with deeply customized timbres, resulting in speech data that better matches user preferences. Therefore, the speech synthesis method provided in this application can meet users' personalized needs for synthesized speech timbres. Furthermore, allowing users to deeply participate in the selection of synthesized speech timbres makes speech synthesis more interesting and interactive, thereby improving the user experience.

[0089] Second Embodiment

[0090] The speech synthesis method provided in this application can be implemented based on a pre-built speech generation module. Optionally, the speech generation module can be a speech synthesis model. Of course, this embodiment is not limited to this. In addition to being a model, the speech generation module can also be other types of modules, such as modules based on speech generation rules. This embodiment does not limit the specific form of the speech generation module.

[0091] The speech synthesis model is trained using training speech data and corresponding training text. The training text is the labeled text of the training speech data. Preferably, in order to synthesize speech data from different speakers, the training speech data used to train the speech synthesis model uses speech data from multiple (e.g., thousands) different speakers (e.g., speaker A's speech data, speaker B's speech data, speaker C's speech data, etc.), with multiple speech data points for each speaker (e.g., more than 100). Preferably, the duration of each speech data point for each speaker is around ten seconds.

[0092] It should be noted that, in order for the speech synthesis model to effectively synthesize speech data with customized timbres based on user-defined timbres, the timbre dimensions need to be considered when collecting training speech data. For example, if the timbre dimensions include "sweet female voice," "deep male voice," and "nasal quality," then speech data from speakers with a balanced male-to-female ratio (e.g., collecting speech data from 600 male speakers and 600 female speakers) and simultaneously possessing the characteristics of "nasal quality," "sweet female voice," and "deep male voice" should be collected. Furthermore, it should be ensured that there are at least fifty speakers with the characteristics of "nasal quality," "sweet female voice," and "deep male voice." For instance, all collected speech data should include speech data from 60 speakers with "nasal quality," 65 speakers with "sweet female voice," and 60 speakers with "deep male voice."

[0093] The following section uses the speech generation module as an example to introduce the process of speech synthesis based on the speech synthesis model.

[0094] Please see Figure 2 The diagram illustrates the process of speech synthesis based on a speech synthesis model, which may include:

[0095] Step S201: Obtain text features for speech synthesis and speech data of the specified speaker.

[0096] The text features used for speech synthesis are obtained from the text used for speech synthesis. These text features may include phoneme information, tone information, prosodic word segmentation information, etc.

[0097] Step S202: Based on the speech synthesis model, extract speaker features from the speech data of a specified speaker, and use the extracted speaker features as the original timbre feature vector.

[0098] Specifically, such as Figure 3As shown, the speech synthesis model may include a speaker encoding module 301, which can extract speaker features from the speech data of a specified speaker based on the speaker encoding module of the speech synthesis model, and use the extracted speaker features as the original timbre feature vector.

[0099] Step S203: Based on the original timbre feature vector, the timbre adjustment parameters determined by the user under the set timbre dimension, and the timbre stretching vector under the set timbre dimension, determine the timbre feature vector of the user-customized timbre.

[0100] Specifically, the process of determining the timbre feature vector of a user-customized timbre, based on the original timbre feature vector, the timbre adjustment parameters determined by the user under the set timbre dimension, and the timbre stretching vector under the set timbre dimension, can include:

[0101] Step S2031: Determine the timbre feature adjustment vector under the set timbre dimension based on the timbre adjustment parameters and timbre stretching vector.

[0102] Among them, the timbre feature adjustment vector under the set timbre dimension is a vector used to adjust the original timbre feature vector under the set timbre dimension.

[0103] Specifically, the process of determining the timbre feature adjustment vector under a given timbre dimension, based on the timbre adjustment parameters and timbre stretching vector, can include: multiplying the timbre adjustment parameters under the given timbre dimension by the timbre stretching vector, and using the result as the timbre feature adjustment vector under the given timbre dimension. It should be noted that when the given timbre dimension includes multiple timbre dimensions, the timbre adjustment parameters under the same dimension are multiplied by the timbre stretching vector to obtain the timbre stretching vector for each timbre dimension.

[0104] Step S2032: Determine the timbre feature vector of the user-customized timbre based on the original timbre feature vector and the timbre feature adjustment vector under the set timbre dimension.

[0105] Specifically, the process of determining the timbre feature vector of a user-customized timbre based on the original timbre feature vector and the timbre feature adjustment vector under the set timbre dimension may include: summing the original timbre feature vector with the timbre feature adjustment vector under the set timbre dimension, and using the summation result as the timbre feature vector of the user-customized timbre.

[0106] For example, the timbre dimensions are set to include "sweet female voice", "deep male voice", and "nasal quality", where the timbre adjustment parameter under the "sweet female voice" timbre dimension is λ. 甜美 The timbre adjustment parameter for the "deep male voice" timbre dimension is λ. 浑厚 The timbre adjustment parameter for the "nasal quality" timbre dimension is λ. 鼻音The timbre stretching vector under the timbre dimension of "sweet female voice" is d. 甜美 The timbre stretching vector under the timbre dimension of "deep male voice" is d. 浑厚 The timbre stretching vector under the timbre dimension of "nasal quality" is d 鼻音 Then, the timbre feature adjustment vector under the timbre dimension of "sweet female voice" is λ. 甜美 ×d 甜美 The timbre feature adjustment vector under the timbre dimension of "deep male voice" is λ. 浑厚 ×d 浑厚 The timbre feature adjustment vector under the timbre dimension of "nasality" is λ. 鼻音 ×d 鼻音 User-customized timbre timbre feature vector s 定制 for:

[0107] s 定制 =s 原始 +λ 甜美 ×d 甜美 +λ 浑厚 ×d 浑厚 +λ 鼻音 ×d 鼻音 (1)

[0108] Among them, s 原始 λ represents the original timbre feature vector. 甜美 , λ 浑厚 , λ 鼻音 It is a continuous value, taking values ​​in the range [0,1]. It should be noted that when λ... 甜美 , λ 浑厚 , λ 鼻音 When both λ and λ are 0, the timbre of the synthesized speech is the original timbre of the specified speaker; when λ... 甜美 , λ 浑厚 , λ 鼻音 When any of λ is non-zero, the timbre of the synthesized speech is no longer the original timbre of the specified speaker, but a new timbre. The user can adjust λ. 甜美 , λ 浑厚 , λ 鼻音 It can generate different timbres to easily meet your personalized needs.

[0109] Step S204: Based on the speech synthesis model, text features, and the timbre feature vector of the user-customized timbre, synthesize speech data with the user-customized timbre.

[0110] Specifically, such as Figure 3As shown, in addition to the speaker encoding module 301 mentioned above, the speech synthesis model also includes a text encoding module 302, a duration prediction module 303, a length adjustment module 304, and a decoding module 305. Therefore, the process of synthesizing speech data with a user-customized timbre based on the speech synthesis model, text features, and the timbre feature vector of the user-customized timbre can include:

[0111] Step S2041: The text encoding module 302 based on the speech synthesis model encodes the text features to obtain the phoneme-level context feature vector.

[0112] Specifically, the text features are input into the text encoding module 302 of the speech synthesis model for encoding, and the text encoding module 302 outputs a phoneme-level context feature vector.

[0113] Step S2042: Based on the duration prediction module 303 in the speech synthesis model, the phoneme pronunciation duration frame number is predicted based on the phoneme-level context feature vector and speaker features.

[0114] Specifically, the phoneme-level contextual feature vector output by the text encoding module 302 and the speaker features output by the speaker encoding module 301 are input into the duration prediction module 303 in the speech synthesis model. The duration prediction module 303 in the speech synthesis model outputs the predicted pronunciation duration in frames for each phoneme. It should be noted that, in order to preserve the duration, rate, and prosodic features of a specified speaker, the speaker features extracted directly from the speech data of the specified speaker are input into the duration prediction module.

[0115] Step S2043: Based on the length adjustment module 304 in the speech synthesis model, the phoneme-level context feature vector is expanded into a frame-level feature vector sequence according to the phoneme pronunciation duration frame number.

[0116] Step S2044: Based on the decoding module 305 in the speech synthesis model, predict the spectral features using the frame-level feature vector sequence and the timbre feature vector of the user-customized timbre.

[0117] The predicted spectral features are the spectral features of the speech to be synthesized.

[0118] Step S2045: Based on the spectral characteristics, synthesize speech data with user-customized timbre.

[0119] Specifically, the spectral characteristics output by the decoding section 305 can be input into the vocoder to obtain synthesized speech with a user-customized timbre.

[0120] The following section describes the process of training the speech synthesis model using training speech and the corresponding training text.

[0121] The process of training a speech synthesis model using training speech data and corresponding training text can include:

[0122] Step a1: Obtain training speech data and corresponding training text, and extract text features based on the training text. Use the extracted text features as training text features.

[0123] The training text features include phoneme information, tone information, and prosodic word segmentation information.

[0124] Step a2: Extract speaker features from the training speech data based on the speech synthesis model, and use the extracted speaker features as training speaker features.

[0125] Specifically, the spectral features of the training speech are obtained and input into the speaker coding module of the speech synthesis model. The speaker coding module extracts sentence-level speaker representation features, i.e. speaker features, from the spectral features of the training speech data.

[0126] Step a3: Based on the speech synthesis model, training text features, and training speaker features, predict spectral features, and synthesize speech data according to the predicted spectral features.

[0127] Specifically, the process of predicting spectral features based on speech synthesis models, training text features, and training speaker features includes:

[0128] Step a31: The text encoding module based on the speech synthesis model encodes the features of the training text to obtain the phoneme-level context feature vector corresponding to the training text.

[0129] Specifically, the training text features are input into the text encoding module of the speech synthesis model for encoding, and the text encoding module outputs the phoneme-level context feature vector corresponding to the training text.

[0130] Step a32: Based on the length adjustment module in the speech synthesis model, the phoneme-level context feature vectors corresponding to the training text are expanded into frame-level feature vector sequences based on the actual pronunciation duration of the phonemes in frames, so as to obtain the frame-level feature vector sequences corresponding to the training text.

[0131] Among them, the actual pronunciation duration frame number of a phoneme refers to the actual pronunciation duration of each phoneme obtained from the training text. The actual pronunciation duration frame number of a phoneme can be determined based on the training speech data.

[0132] Step a33: Based on the decoding part of the speech synthesis model, predict the spectral features according to the frame-level feature vector sequence corresponding to the training text and the training speaker features, and synthesize speech data according to the predicted spectral features.

[0133] Step a4: Determine the prediction loss of the speech synthesis model based on the predicted spectral features, the actual spectral features, the actual phoneme information, and the predicted phoneme information.

[0134] The actual spectral features are the spectral features of the training speech data, the actual phoneme information is the phoneme information obtained from the training text, such as phoneme duration information and phoneme state duration obtained from the training text, and the predicted phoneme information is the phoneme information obtained from the text corresponding to the synthesized speech data, such as phoneme duration information and phoneme state duration obtained from the text corresponding to the synthesized speech data.

[0135] Specifically, the prediction loss L of the speech synthesis model can be determined based on the following formula:

[0136]

[0137] Where T is the total number of frames in the training speech data, y i This represents the spectral features of the i-th frame of the training speech data. Let d be the spectral features of the i-th frame of the synthesized speech data, P be the total number of phonemes obtained from the training text, and d be the spectral features of the i-th frame of the synthesized speech data. j To represent the phoneme information of the j-th phoneme obtained from the training text, for example, the phoneme duration information of the j-th phoneme obtained from the training text, i.e., the actual phoneme duration information of the j-th phoneme, This represents the phoneme information of the j-th phoneme obtained from the text corresponding to the synthesized speech data, such as the phoneme duration information of the j-th phoneme obtained from the text corresponding to the synthesized speech data.

[0138] Step a5: Update the parameters of the speech synthesis model based on the prediction loss of the speech synthesis model.

[0139] The speech synthesis model is trained multiple times according to steps a1 to a5 above until the training termination condition is met.

[0140] Third Embodiment

[0141] As mentioned in the first embodiment above, in order to synthesize user-customized voice data, it is necessary to use the original timbre feature vector, the timbre adjustment parameters determined by the user under the set timbre dimension, and the timbre stretching vector under the set timbre dimension. This embodiment focuses on introducing the specific implementation process of determining the timbre stretching vector under the set timbre dimension.

[0142] Please see Figure 4 This illustrates a flowchart for determining the timbre stretch vector under a given timbre dimension, which may include:

[0143] Step S401: Obtain the total set of voice data.

[0144] The speech data set in this embodiment includes the speech data used to train the speech synthesis model, after labeling the speech data with timbre attribute values. That is, the speech data set includes the speech data of multiple speakers (e.g., the speech data of speaker A, speaker B, speaker C, etc.). Each speaker's speech data is labeled with the timbre attribute value of a set timbre dimension. For example, the set timbre dimension includes "sweet female voice", "deep male voice", and "nasal sound". The timbre attributes under the set timbre dimension include "gender attribute", "sweet attribute", "deep attribute", and "nasal attribute". Then, each speaker's speech data in the speech data set is labeled with the attribute value of "gender attribute", the attribute value of "sweet attribute", the attribute value of "deep attribute", and the attribute value of "nasal sound".

[0145] When annotating the timbre attributes of speech data from multiple speakers, the timbre attribute values ​​for each speaker's speech data can be annotated according to the timbre attributes under a set timbre dimension. In order to obtain robust and consistent timbre attribute values, this embodiment preferably adopts an integrated method of multiple annotation results. That is, for each speaker's speech data, firstly, the timbre attribute values ​​annotated by multiple annotators on the timbre attributes of the speaker's speech data under the set timbre dimension are obtained; then, the final timbre attribute value of the speaker's speech data under the set timbre dimension is determined based on the annotation results of multiple annotators; finally, the final timbre attribute value determined is annotated for the speaker's speech data.

[0146] The following example illustrates the process of setting timbre attribute values ​​for the timbre dimension in the annotation of a speaker's speech data:

[0147] The timbre dimension is defined as "sweet female voice," "deep male voice," and "nasal quality." The timbre attributes within this dimension include "gender attribute," "sweetness attribute," "deep voice attribute," and "nasal quality attribute." When annotating a speaker's voice data, for the "gender attribute," if the voice is female, the timbre attribute value is labeled "female," and if the voice is male, the timbre attribute value is labeled "male." For the "sweetness attribute," if... If the voice is female and sweet, the timbre attribute value of "sweetness" will be marked as "1"; otherwise, the timbre attribute value of "sweetness" will be marked as "0". For the "resonant" attribute, if the voice is male and resonant, the timbre attribute value of "resonant" will be marked as "1"; otherwise, the timbre attribute value of "resonant" will be marked as "0". For the "nasal" attribute, if the voice has a nasal tone, the attribute value of "nasal" will be marked as "1"; otherwise, the attribute value of "nasal" will be marked as "0".

[0148] For the speech data of speaker A, annotation results from 6 annotators (preferably 3 males and 3 females) can be obtained, and the final annotation result is determined based on the annotation results of the 6 annotators. Optionally, one speech data point can be extracted from all the speech data of speaker A, and each annotator can annotate based on that speech data point. It should be noted that the number of annotators, 6, mentioned above is only an example, and the number of annotators can also be other, such as 8, 10, etc.

[0149] Specifically, for the "gender attribute," the timbre attribute values ​​of speaker A's speech data annotated by six annotators are obtained. Based on the annotation results of the six annotators for speaker A's speech data in the "gender attribute," the final annotation result for speaker A's speech data in the "gender attribute" is determined. Specifically, if the annotation results of the six annotators for speaker A's speech data in the "gender attribute" are inconsistent, the final attribute value of speaker A's speech data in the "gender attribute" is marked as "invalid," i.e., l gender =invalid. If all six annotators give the same "male" annotation for speaker A's speech data in the "gender attribute" category, then the final attribute value for speaker A's speech data in the "gender attribute" category will be annotated as "male". gender =Male. If all six annotators give the same gender attribute annotation for speaker A's speech data, and all are "female", then the final gender attribute value of speaker A's speech data will be annotated as female, i.e., l gender =Female.

[0150] For the "sweetness attribute," obtain the timbre attribute values ​​of speaker A's speech data labeled with the "sweetness attribute" by six annotators. Based on the annotation results of the six annotators on speaker A's speech data for the "sweetness attribute," determine the final annotation result of speaker A's speech data for the "sweetness attribute." Specifically, if the annotation results of the six annotators on speaker A's speech data for the "sweetness attribute" are consistent and all are "1," then the final attribute value of speaker A's speech data for the "sweetness attribute" is labeled as 1, i.e., l. 甜美女声 =1. If the annotation results for the "sweetness attribute" of speaker A's speech data by the six annotators are consistent and all are 0, or if the annotation results for the "sweetness attribute" of speaker A's speech data by the six annotators are inconsistent, then the final attribute value of speaker A's speech data for the "sweetness attribute" will be annotated as 0, i.e., l 甜美女声 =0.

[0151] For the "resonance attribute," obtain the timbre attribute values ​​of speaker A's speech data labeled with the "resonance attribute" by six annotators. Based on the annotation results of the six annotators on speaker A's speech data in the "resonance attribute," determine the final annotation result of speaker A's speech data in the "resonance attribute." Specifically, if the annotation results of the six annotators on speaker A's speech data in the "resonance attribute" are consistent and all are "1," then the final timbre attribute value of speaker A's speech data in the "resonance attribute" is labeled as "1," i.e., l 浑厚男声 =1. If the annotation results for the "resonance attribute" of speaker A's voice data by the six annotators are consistent and all are "0", or if the annotation results for the "resonance attribute" of speaker A's voice data by the six annotators are inconsistent, then the final timbre attribute value of speaker A's voice data in the "resonance attribute" will be annotated as "0", i.e., l 浑厚男声 =0.

[0152] For the "nasal sound attribute," obtain the timbre attribute values ​​of speaker A's speech data labeled with the "nasal sound attribute" by six annotators. Based on the annotation results of the six annotators for speaker A's speech data in the "nasal sound attribute," determine the final annotation result of speaker A's speech data in the "nasal sound attribute." Specifically, if the annotation results of the six annotators for speaker A's speech data in the "nasal sound attribute" are consistent and all are "1," then the final timbre attribute value of speaker A's speech data in the "nasal sound attribute" is labeled as "1," i.e., l 鼻音 =1. If the annotation results for speaker A's speech data on the "nasal attribute" by the six annotators are consistent and all are "0", or if the annotation results for speaker A's speech data on the "nasal attribute" by the six annotators are inconsistent, then the final timbre attribute value of speaker A's speech data on the "nasal attribute" will be annotated as "0", i.e., l鼻音 =0.

[0153] Step S402: Take each timbre dimension included in the set timbre dimension as the target timbre dimension, and obtain the speech data related to the target timbre dimension from the speech data set according to the timbre attribute values ​​of the speech data annotation in the speech data set.

[0154] Specifically, based on the timbre attribute values ​​annotated in the speech data set, the process of obtaining speech data related to the target timbre dimension from the speech data set may include: obtaining speech data whose first timbre attribute value is the first target attribute value from the speech data set, forming a first speech dataset; obtaining speech data whose second timbre attribute value is the second target attribute value from the speech data set, forming a second speech dataset; the speech data in the first speech dataset and the speech data in the second speech dataset are used as the speech data related to the target timbre dimension. Optionally, when obtaining speech data whose first timbre attribute value is the first target attribute value from the speech data set, only one speech data point may be obtained for each speaker; similarly, when obtaining speech data whose second timbre attribute value is the second target attribute value from the speech data set, only one speech data point may be obtained for each speaker.

[0155] Among them, the first timbre attribute is a timbre attribute related to the target timbre dimension, and the first timbre attribute has multiple optional timbre attribute values. The first target attribute value is the attribute value related to the target timbre dimension among the multiple optional timbre attribute values ​​of the first timbre attribute. The second timbre attribute is a timbre attribute related to the first timbre attribute, and the second timbre attribute has multiple optional timbre attribute values. The second target attribute value is the attribute value related to the first timbre attribute among the multiple optional timbre attribute values ​​of the second timbre attribute.

[0156] For example, if the target timbre dimension is "sweet female voice", then the first timbre attribute is "sweet attribute", the second timbre attribute is "gender attribute", the first target attribute value is "1", and the second target attribute value is "female". The process of obtaining speech data related to "sweet female voice" from the speech data set includes: obtaining the attribute value of "sweet attribute" as "1" (i.e., l) from the speech data set. 甜美女声 The first speech dataset is composed of speech data of type 1 (=1). From the speech data set, the attribute value of "gender" is obtained as "female" (i.e., l). gender =Female) voice data, forming the second voice dataset.

[0157] For example, if the target timbre dimension is "deep male voice", then the first timbre attribute is "deep voice attribute", the second timbre attribute is "gender attribute", the first target attribute value is "1", and the second target attribute value is "male". The process of obtaining speech data related to "deep male voice" from the speech data set includes: obtaining the attribute value of "deep voice attribute" as "1" (i.e., l) from the speech data set. 浑厚男声 The first speech dataset is composed of speech data of type 1 (=1). From this dataset, the attribute value for "gender" is "male" (i.e., l). gender The voice data of male (=male) are used to form the second voice dataset.

[0158] For example, if the target timbre dimension is "nasality", then the first timbre attribute is "nasality attribute", the second timbre attribute is "gender attribute", the first target attribute value is "1", and the second target attribute value is "male" and "female". The process of obtaining speech data related to "nasality" from the speech data set includes: obtaining the attribute value of "nasality attribute" as "1" (i.e., l) from the speech data set. 鼻音 The first speech dataset is composed of speech data of type 1 (=1). From this dataset, the attribute value for "gender" is "male" (i.e., l). gender =Male) voice data and the attribute value of "gender attribute" is "female" (i.e., l gender =Female), forming the second speech dataset.

[0159] Step S403: Determine the timbre stretching vector under the target timbre dimension based on the speech data related to the target timbre dimension.

[0160] Specifically, based on the first and second speech datasets, the timbre stretching vector under the target timbre dimension is determined.

[0161] More specifically, the process of determining the timbre stretching vector under the target timbre dimension based on the first and second speech datasets may include: extracting speaker features from the speech data of each speaker in the first speech dataset (for example, if the first speech dataset includes speech data of 100 speakers, then extracting speaker features from the speech data of each speaker will yield 100 speaker features), and calculating the mean of several speaker features extracted from the speech data of the first speech dataset, with the calculated mean serving as the first average timbre vector under the target timbre dimension; extracting speaker features from the speech data of each speaker in the second speech dataset (for example, if the second speech dataset includes speech data of 150 speakers, then extracting speaker features from the speech data of each speaker will yield 150 speaker features), and calculating the mean of several speaker features extracted from the speech data of the second speech dataset, with the calculated mean serving as the second average timbre vector under the target timbre dimension; and determining the timbre stretching vector under the target timbre dimension based on the first average timbre vector and the second average timbre vector under the target timbre dimension.

[0162] For example, the target timbre dimension is "sweet female voice", and the first speech dataset includes l 甜美女声 =1 speech data, the second speech dataset includes l gender =For female voice data, the first average timbre vector s under the timbre dimension of "sweet female voice" is... 甜美女声 Second average timbre vector s f It can be represented as:

[0163]

[0164]

[0165] N in equation (3) 甜美女声 N represents the number of speakers involved in the speech data in the first speech dataset. For example, if the first speech dataset includes speech data from 100 speakers, then N in equation (3) 甜美女声 =100, s in equation (3) i N represents the number of speech data involved in the first speech dataset. 甜美女声 The speaker characteristics of the i-th speaker among the speakers, N in equation (4) f N represents the number of speakers involved in the speech data in the second speech dataset. For example, if the second speech dataset includes speech data from 150 speakers, then N is... f =150, s in equation (4) j N represents the number of speech data involved in the second speech dataset. fSpeaker characteristics of the j-th speaker among 3 speakers.

[0166] The first average timbre vector s under the timbre dimension of "sweet female voice" was obtained. 甜美女声 Second average timbre vector s f Then, the timbre stretching vector d under the timbre dimension of "sweet female voice" can be determined according to the following formula. 甜美 :

[0167] d 甜美 =s 甜美女声 -s f (5)

[0168] For example, the target timbre dimension is "deep male voice", and the first speech dataset includes l 浑厚男声 =1 speech data, the second speech dataset includes l gender = Male voice data, then the first average timbre vector s under the timbre dimension of "deep male voice" 浑厚男声 Second average timbre vector s m It can be represented as:

[0169]

[0170]

[0171] N in equation (6) 浑厚男声 s represents the number of speakers involved in the speech data in the first speech dataset, where s in equation (6) i N represents the number of speech data involved in the first speech dataset. 浑厚男声 The speaker characteristics of the i-th speaker among the speakers, N in equation (7) m The s in equation (7) represents the number of speakers involved in the speech data in the second speech dataset. j N represents the number of speech data involved in the second speech dataset. m Speaker characteristics of the j-th speaker among 3 speakers.

[0172] The first average timbre vector s under the timbre dimension of "deep male voice" was obtained. 浑厚男声 Second average timbre vector s m Then, the timbre stretching vector d under the timbre dimension of "deep male voice" can be determined according to the following formula. 浑厚 :

[0173] d 浑厚 =s 浑厚男声 -s m (8)

[0174] For example, the target timbre dimension is "nasal quality," and the first speech dataset includes l鼻音 =1 speech data, the second speech dataset includes l gender =Male voice data and l gender =For female voice data, the first average timbre vector s under the timbre dimension of "nasal quality" is... 鼻音 Second average timbre vector s avg It can be represented as:

[0175]

[0176]

[0177] N in equation (9) 鼻音 s represents the number of speakers involved in the speech data in the first speech dataset, where s in equation (9) i N represents the number of speech data involved in the first speech dataset. 鼻音 The speaker features of the i-th speaker among the speakers, N in equation (10) represents the number of speakers involved in the speech data in the second speech dataset, and s in equation (10) j This represents the speaker feature of the j-th speaker among the N speakers involved in the speech data in the second speech dataset.

[0178] The first average timbre vector s is obtained under the timbre dimension of "nasal quality". 鼻音 Second average timbre vector s avg Then, the timbre stretching vector d under the timbre dimension of "nasal resonance" can be determined according to the following formula. 鼻音 :

[0179] d 鼻音 =s 鼻音 -s avg (11)

[0180] The method provided in this embodiment can determine the timbre stretching vector under a set timbre dimension, such as the timbre stretching vector under the timbre dimension of "sweet female voice", the timbre stretching vector under the timbre dimension of "deep male voice", and the timbre stretching vector under the timbre dimension of "nasal sound".

[0181] Fourth embodiment

[0182] This application also provides a speech synthesis device. The speech synthesis device provided in this application is described below. The speech synthesis device described below can be referred to in correspondence with the speech synthesis method described above.

[0183] Please see Figure 5The diagram shows a schematic of the speech synthesis device provided in the embodiment of this application, which may include: a data acquisition module 501, a speaker feature extraction module 502, a timbre feature vector determination module 503, and a speech synthesis module 504.

[0184] The data acquisition module 501 is used to acquire text features for speech synthesis and speech data of a specified speaker.

[0185] The speaker feature extraction module 502 is used to extract speaker features from the speech data of the specified speaker as the original timbre feature vector.

[0186] The timbre feature vector determination module 503 is used to determine the timbre feature vector of the user-customized timbre based on the original timbre feature vector, the timbre adjustment parameters determined by the user under the set timbre dimension, and the timbre stretching vector under the set timbre dimension.

[0187] The speech synthesis module 504 is used to synthesize speech data with a user-customized timbre based on the text features and the timbre feature vector of the user-customized timbre.

[0188] Optionally, when determining the timbre feature vector of a user-customized timbre based on the original timbre feature vector, the timbre adjustment parameters determined by the user under the set timbre dimension, and the timbre stretching vector under the set timbre dimension, the timbre feature vector determination module 503 is specifically used for:

[0189] Based on the timbre adjustment parameters under the set timbre dimension and the timbre stretching vector under the set timbre dimension, determine the timbre feature adjustment vector under the set timbre dimension;

[0190] The timbre feature vector of the user-customized timbre is determined based on the original timbre feature vector and the timbre feature adjustment vector under the set timbre dimension.

[0191] Optionally, when the speech synthesis module 504 synthesizes speech data with a user-customized timbre based on the text features and the timbre feature vector of the user-customized timbre, it is specifically used for:

[0192] Based on the text features and the speaker features, obtain a frame-level feature vector sequence;

[0193] Based on the frame-level feature vector sequence and the timbre feature vector of the user-customized timbre, speech data with the user-customized timbre is synthesized.

[0194] Optionally, when the speech synthesis module 504 obtains a frame-level feature vector sequence based on the text features and the speaker features, it specifically performs the following:

[0195] Based on the text features, obtain the phoneme-level context feature vector;

[0196] Based on the phoneme-level context feature vector and the speaker features, predict the number of frames for the pronunciation duration of the phoneme;

[0197] Based on the number of frames for the pronunciation duration of the phoneme, the context feature vector at the phoneme level is expanded into a sequence of feature vectors at the frame level.

[0198] Optionally, the setting of timbre dimensions includes one or more timbre dimensions.

[0199] Optionally, the speech synthesis apparatus provided in this application embodiment further includes: a timbre stretching vector determination module for determining the timbre stretching vector under the set timbre dimension. The timbre stretching vector determination module, when determining the timbre stretching vector under the set timbre dimension, is specifically used for:

[0200] Obtain a speech data set, which includes speech data from multiple speakers, and each speaker's speech data is labeled with a timbre attribute value under the set timbre dimension;

[0201] Each timbre dimension included in the set timbre dimension is taken as the target timbre dimension. Based on the timbre attribute values ​​of the voice data labeled in the voice data set, voice data related to the target timbre dimension is obtained from the voice data set.

[0202] Based on the speech data related to the target timbre dimension, determine the timbre stretching vector under the target timbre dimension.

[0203] Optionally, the speech synthesis device provided in this application embodiment further includes: a timbre attribute value annotation module.

[0204] When the timbre attribute value annotation module annotates the timbre attribute values ​​under the defined timbre dimension for a speaker's speech data, it is specifically used for:

[0205] Obtain the timbre attribute values ​​marked by multiple annotators on the timbre attributes under the set timbre dimension based on the speaker's speech data, so as to obtain the annotation results of multiple annotators;

[0206] Based on the annotation results of the multiple annotators, the timbre attribute value of the speaker's voice data under the set timbre dimension is determined, and the determined timbre attribute value is annotated for the speaker's voice data.

[0207] Optionally, when the timbre stretching vector determination module obtains speech data related to the target timbre dimension from the speech data set based on the timbre attribute values ​​labeled in the speech data set, it is specifically used for:

[0208] Speech data with the attribute value of the first timbre attribute and the attribute value of the first target attribute are obtained from the speech data set to form a first speech dataset, wherein the first timbre attribute is a timbre attribute related to the target timbre dimension, and the first target attribute value is an attribute value related to the target timbre dimension;

[0209] Speech data with the attribute value of the second timbre attribute and the attribute value of the second target attribute are obtained from the speech data set to form a second speech dataset, wherein the second timbre attribute is a timbre attribute related to the first timbre attribute, and the second target attribute value is an attribute value related to the first timbre attribute.

[0210] Optionally, when the timbre stretching vector determination module determines the timbre stretching vector under the target timbre dimension based on the speech data related to the target timbre dimension, it is specifically used for:

[0211] Speaker features are extracted from the speech data of each speaker in the first speech dataset, and the mean of several speaker features extracted from the speech data in the first speech dataset is calculated. The calculated mean is used as the first average timbre vector under the target timbre dimension.

[0212] Speaker features are extracted from the speech data of each speaker in the second speech dataset, and the mean of several speaker features extracted from the speech data in the second speech dataset is calculated. The calculated mean is used as the second average timbre vector under the target timbre dimension.

[0213] The timbre stretching vector under the target timbre dimension is determined based on the first average timbre vector and the second average timbre vector under the target timbre dimension.

[0214] Optionally, when extracting speaker features from the speech data of the specified speaker, the speaker feature extraction module 502 is specifically used for:

[0215] Using a pre-built speech generation module, speaker features are extracted from the speech data of the specified speaker.

[0216] When the speech synthesis module 504 synthesizes speech data with a user-customized timbre based on the text features and the timbre feature vector of the user-customized timbre, it is specifically used for:

[0217] Based on the speech generation module, the text features, and the timbre feature vector of the user-customized timbre, speech data with the user-customized timbre is synthesized.

[0218] Optionally, the speech generation module is a speech synthesis model, which is trained using multiple training speech data from multiple speakers and training texts corresponding to the multiple training speech data.

[0219] When the speech synthesis module 504 synthesizes speech data with a user-customized timbre based on the speech generation module, the text features, and the timbre feature vector of the user-customized timbre, it is specifically used for:

[0220] The text encoding module based on the speech synthesis model encodes the text features to obtain a phoneme-level context feature vector.

[0221] The duration prediction module based on the speech synthesis model predicts the number of frames for the pronunciation of a phoneme based on the context feature vector at the phoneme level and the speaker features.

[0222] Based on the length adjustment module of the speech synthesis model, the phoneme-level context feature vector is expanded into a frame-level feature vector sequence according to the number of frames of pronunciation duration of the phoneme.

[0223] Based on the decoding module in the speech synthesis model, the spectral features are predicted according to the frame-level feature vector sequence and the timbre feature vector of the user-customized timbre.

[0224] Based on the spectral characteristics, user-customized voice data is synthesized.

[0225] The speech synthesis apparatus provided in this application first acquires text features for speech synthesis and speech data of a specified speaker. Then, it extracts speaker features from the specified speaker's speech data to obtain an original timbre feature vector. Next, based on the original timbre feature vector, user-determined timbre adjustment parameters under a set timbre dimension, and a timbre stretching vector under the set timbre dimension, it determines a timbre feature vector for a user-customized timbre. Finally, based on the text features and the timbre feature vector of the user-customized timbre, it synthesizes speech data with the user-customized timbre. The speech synthesis apparatus provided in this application can synthesize speech data with deeply customized timbres, resulting in speech data that better matches user preferences. Therefore, the speech synthesis apparatus provided in this application can meet users' personalized needs for synthesized speech timbres. Furthermore, allowing users to deeply participate in the selection of synthesized speech timbres makes speech synthesis more interesting and interactive, thereby improving the user experience.

[0226] Fifth Embodiment

[0227] This application also provides a speech synthesis device. Please refer to [link to relevant documentation]. Figure 6 The diagram shows the structure of the speech synthesis device, which may include: at least one processor 601, at least one communication interface 602, at least one memory 603 and at least one communication bus 604.

[0228] In this embodiment of the application, the number of processor 601, communication interface 602, memory 603 and communication bus 604 is at least one, and processor 601, communication interface 602 and memory 603 communicate with each other through communication bus 604.

[0229] The processor 601 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention.

[0230] The memory 603 may include high-speed RAM, and may also include non-volatile memory, such as at least one disk storage device;

[0231] The memory stores a program, which the processor can call. The program is used for:

[0232] Acquire text features for speech synthesis and speech data of a specified speaker;

[0233] Speaker features are extracted from the speech data of the specified speaker and used as the original timbre feature vector.

[0234] Based on the original timbre feature vector, the timbre adjustment parameters determined by the user under the set timbre dimension, and the timbre stretching vector under the set timbre dimension, the timbre feature vector of the user-customized timbre is determined.

[0235] Based on the text features and the timbre feature vector of the user-customized timbre, speech data with the user-customized timbre is synthesized.

[0236] Optionally, the refined and extended functions of the program can be found in the description above.

[0237] Sixth Embodiment

[0238] This application also provides a computer-readable storage medium that can store a program suitable for execution by a processor, the program being used for:

[0239] Acquire text features for speech synthesis and speech data of a specified speaker;

[0240] Speaker features are extracted from the speech data of the specified speaker and used as the original timbre feature vector.

[0241] Based on the original timbre feature vector, the timbre adjustment parameters determined by the user under the set timbre dimension, and the timbre stretching vector under the set timbre dimension, the timbre feature vector of the user-customized timbre is determined.

[0242] Based on the text features and the timbre feature vector of the user-customized timbre, speech data with the user-customized timbre is synthesized.

[0243] Optionally, the refined and extended functions of the program can be found in the description above.

[0244] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0245] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0246] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A speech synthesis method, characterized in that, include: Acquire text features for speech synthesis and speech data of a specified speaker; Speaker features are extracted from the speech data of the specified speaker and used as the original timbre feature vector. Based on the original timbre feature vector, the timbre adjustment parameters determined by the user under the set timbre dimension, and the timbre stretching vector under the set timbre dimension, the timbre feature vector of the user-customized timbre is determined, and the timbre stretching vector under the set timbre dimension is the basic vector used to adjust the timbre. Based on the text features and the timbre feature vector of the user-customized timbre, speech data with the user-customized timbre is synthesized. The process of determining the timbre stretching vector under the set timbre dimension includes: Obtain a speech data set, which includes speech data from multiple speakers, and each speaker's speech data is labeled with a timbre attribute value under the set timbre dimension; Based on the timbre attribute values ​​of the voice data labeled in the speech data set, obtain the voice data related to the set timbre dimension from the speech data set; Based on the speech data related to the set timbre dimension, determine the timbre stretching vector under the set timbre dimension; The process of determining the timbre feature vector of the user-customized timbre based on the original timbre feature vector, the timbre adjustment parameters determined by the user under the set timbre dimension, and the timbre stretching vector under the set timbre dimension includes: Based on the timbre adjustment parameters under the set timbre dimension and the timbre stretching vector under the set timbre dimension, determine the timbre feature adjustment vector under the set timbre dimension; The timbre feature vector of the user-customized timbre is determined based on the original timbre feature vector and the timbre feature adjustment vector under the set timbre dimension.

2. The speech synthesis method according to claim 1, characterized in that, The process of synthesizing user-customized voice data based on the text features and the timbre feature vector of the user-customized voice includes: Based on the text features and the speaker features, obtain a frame-level feature vector sequence; Based on the frame-level feature vector sequence and the timbre feature vector of the user-customized timbre, speech data with the user-customized timbre is synthesized.

3. The speech synthesis method according to claim 2, characterized in that, The step of obtaining a frame-level feature vector sequence based on the text features and the speaker features includes: Based on the text features, obtain the phoneme-level context feature vector; Based on the phoneme-level context feature vector and the speaker features, the phoneme pronunciation duration in frames is predicted. Based on the number of frames for the pronunciation duration of the phoneme, the context feature vector at the phoneme level is expanded into a sequence of feature vectors at the frame level.

4. The speech synthesis method according to claim 1, characterized in that, The defined timbre dimension includes one or more timbre dimensions; The step of obtaining speech data related to the set timbre dimension from the speech data set based on the timbre attribute values ​​labeled in the speech data set, and the step of determining the timbre stretching vector under the set timbre dimension based on the speech data related to the set timbre dimension, includes: Each timbre dimension included in the set timbre dimension is taken as the target timbre dimension. Based on the timbre attribute values ​​of the voice data labeled in the voice data set, voice data related to the target timbre dimension is obtained from the voice data set. Based on the speech data related to the target timbre dimension, determine the timbre stretching vector under the target timbre dimension.

5. The speech synthesis method according to claim 4, characterized in that, The process of labeling the timbre attribute values ​​of a speaker's speech data under the defined timbre dimension includes: Obtain the timbre attribute values ​​marked by multiple annotators on the timbre attributes under the set timbre dimension based on the speaker's speech data, so as to obtain the annotation results of multiple annotators; Based on the annotation results of the multiple annotators, the timbre attribute value of the speaker's voice data under the set timbre dimension is determined, and the determined timbre attribute value is annotated for the speaker's voice data.

6. The speech synthesis method according to claim 4, characterized in that, The step of obtaining speech data related to the target timbre dimension from the speech data set based on the timbre attribute values ​​labeled in the speech data set includes: Speech data with the attribute value of the first timbre attribute and the attribute value of the first target attribute are obtained from the speech data set to form a first speech dataset, wherein the first timbre attribute is a timbre attribute related to the target timbre dimension, and the first target attribute value is an attribute value related to the target timbre dimension; Speech data with the attribute value of the second timbre attribute and the attribute value of the second target attribute are obtained from the speech data set to form a second speech dataset, wherein the second timbre attribute is a timbre attribute related to the first timbre attribute, and the second target attribute value is an attribute value related to the first timbre attribute.

7. The speech synthesis method according to claim 6, characterized in that, The step of determining the timbre stretching vector under the target timbre dimension based on speech data related to the target timbre dimension includes: Speaker features are extracted from the speech data of each speaker in the first speech dataset, and the mean of several speaker features extracted from the speech data in the first speech dataset is calculated. The calculated mean is used as the first average timbre vector under the target timbre dimension. Speaker features are extracted from the speech data of each speaker in the second speech dataset, and the mean of several speaker features extracted from the speech data in the second speech dataset is calculated. The calculated mean is used as the second average timbre vector under the target timbre dimension. The timbre stretching vector under the target timbre dimension is determined based on the first average timbre vector and the second average timbre vector under the target timbre dimension.

8. The speech synthesis method according to claim 1, characterized in that, The step of extracting speaker features from the speech data of the specified speaker includes: Using a pre-built speech generation module, speaker features are extracted from the speech data of the specified speaker; The process of synthesizing user-customized voice data based on the text features and the timbre feature vector of the user-customized voice includes: Based on the speech generation module, the text features, and the timbre feature vector of the user-customized timbre, speech data with the user-customized timbre is synthesized.

9. The speech synthesis method according to claim 8, characterized in that, The speech generation module is a speech synthesis model, which is trained using multiple training speech data from multiple speakers and training texts corresponding to the multiple training speech data. The process of synthesizing user-customized voice data based on the speech generation module, the text features, and the timbre feature vector of the user-customized voice includes: The text encoding module based on the speech synthesis model encodes the text features to obtain a phoneme-level context feature vector. The duration prediction module based on the speech synthesis model predicts the number of frames for the pronunciation of a phoneme based on the context feature vector at the phoneme level and the speaker features. Based on the length adjustment module of the speech synthesis model, the phoneme-level context feature vector is expanded into a frame-level feature vector sequence according to the number of frames of pronunciation duration of the phoneme. The decoding module based on the speech synthesis model predicts spectral features based on the frame-level feature vector sequence and the timbre feature vector of the user-customized timbre. Based on the spectral characteristics, user-customized voice data is synthesized.

10. A speech synthesis device, characterized in that, include: The module includes a data acquisition module, a speaker feature extraction module, a timbre feature vector determination module, and a speech synthesis module. The data acquisition module is used to acquire text features for speech synthesis and speech data of a specified speaker. The speaker feature extraction module is used to extract speaker features from the speech data of the specified speaker as the original timbre feature vector; The timbre feature vector determination module is used to determine the timbre feature vector of the user-customized timbre based on the original timbre feature vector, the timbre adjustment parameters determined by the user under the set timbre dimension, and the timbre stretching vector under the set timbre dimension. The timbre stretching vector under the set timbre dimension is the basic vector used to adjust the timbre. The speech synthesis module is used to synthesize speech data with a user-customized timbre based on the text features and the timbre feature vector of the user-customized timbre. The speech synthesis device also includes a timbre stretching vector determination module; The timbre stretching vector determination module is used to obtain a speech data set, which includes speech data of multiple speakers. Each speaker's speech data is labeled with a timbre attribute value under the set timbre dimension. Based on the timbre attribute values ​​labeled on the speech data in the speech data set, speech data related to the set timbre dimension is obtained from the speech data set. Based on the speech data related to the set timbre dimension, the timbre stretching vector under the set timbre dimension is determined. When determining the timbre feature vector of a user-customized timbre based on the original timbre feature vector, the timbre adjustment parameters determined by the user under the set timbre dimension, and the timbre stretching vector under the set timbre dimension, the timbre feature vector determination module is specifically used to determine the timbre feature adjustment vector under the set timbre dimension based on the timbre adjustment parameters under the set timbre dimension and the timbre stretching vector under the set timbre dimension, and to determine the timbre feature vector of the user-customized timbre based on the original timbre feature vector and the timbre feature adjustment vector under the set timbre dimension.

11. A speech synthesis device, characterized in that, include: Memory and processor; The memory is used to store programs; The processor is configured to execute the program to implement the various steps of the speech synthesis method as described in any one of claims 1 to 9.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements each step of the speech synthesis method as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Speech synthesis method, device and equipment and storage medium

    CN112735373A

  • Voice synthesis method and device, server and storage medium

    CN113096634A