Speech synthesis model training method, speech synthesis method, speech synthesis device and electronic equipment
By introducing style dimension data and a step-by-step training method into the speech synthesis model, the problem of low accuracy of the existing model is solved, and higher speech synthesis accuracy and style matching are achieved.
Patent Information
- Application Number
- CN202510841503.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-09-23
AI Technical Summary
Existing speech synthesis models consider fewer features during training, resulting in low accuracy.
By obtaining style description sample text, input sample text, output sample speech and style dimension data in the training data, the speech synthesis model is trained, and the backbone network and style control network are used for step-by-step training to increase the number of features considered.
Improved the accuracy of speech synthesis models and enhanced the matching between style control and speech generation.
Smart Images

Figure CN120690171A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to technical fields such as deep learning, natural language processing, speech technology, and large models, and in particular to a training method for a speech synthesis model, a speech synthesis method, a device, and an electronic device. Background Art
[0002] Current speech synthesis models are primarily trained by combining style description sample text, input sample text, and output speech sample text. The loss function of the speech synthesis model is determined based on the output speech sample and the predicted speech sample, taking fewer features into account, resulting in low accuracy in the trained speech synthesis model. Summary of the Invention
[0003] The present disclosure provides a training method for a speech synthesis model, a speech synthesis method, a device, and an electronic device.
[0004] According to one aspect of the present disclosure, a method for training a speech synthesis model is provided, the method comprising: obtaining training data; wherein the training samples in the training data comprise: style description sample text, input sample text, output sample speech, and style dimension data of the output sample speech; obtaining an initial speech synthesis model; and training the speech synthesis model based on the style description sample text, the input sample text, the output sample speech, and the style dimension data to obtain a trained speech synthesis model.
[0005] According to another aspect of the present disclosure, a speech synthesis method is provided, the method comprising: obtaining input text and style description text; inputting the style description text into a style control network in a speech synthesis model, and obtaining a style vector output by the style control network; the speech synthesis model is trained based on style description sample text, input sample text, output sample speech, and style dimension data of the output sample speech; inputting the style vector and the input text into a backbone network in the speech synthesis model, and obtaining an output speech corresponding to the input text output by the backbone network.
[0006] According to another aspect of the present disclosure, a training device for a speech synthesis model is provided, the device comprising: a first acquisition module for acquiring training data; training samples in the training data comprising: style description sample text, input sample text, output sample speech, and style dimension data of the output sample speech; a second acquisition module for acquiring an initial speech synthesis model; and a training processing module for performing training processing on the speech synthesis model based on the style description sample text, the input sample text, the output sample speech, and the style dimension data to obtain a trained speech synthesis model.
[0007] According to another aspect of the present disclosure, a speech synthesis device is provided, comprising: a first acquisition module for acquiring input text and style description text; a second acquisition module for inputting the style description text into a style control network in a speech synthesis model to acquire a style vector output by the style control network; the speech synthesis model is trained based on style description sample text, input sample text, output sample speech, and style dimension data of the output sample speech; and a third acquisition module for inputting the style vector and the input text into a backbone network in the speech synthesis model to acquire output speech corresponding to the input text output by the backbone network.
[0008] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the training method of the speech synthesis model proposed above in the present disclosure; or, to execute the speech synthesis method proposed above in the present disclosure.
[0009] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the training method of the speech synthesis model proposed above in the present disclosure; or, to execute the speech synthesis method proposed above in the present disclosure.
[0010] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the steps of the training method of the speech synthesis model proposed above in the present disclosure; or, implements the steps of the speech synthesis method proposed above in the present disclosure.
[0011] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0013] Figure 1 is a schematic diagram according to a first embodiment of the present disclosure;
[0014] Figure 2 is a schematic diagram according to a second embodiment of the present disclosure;
[0015] Figure 3 is a schematic diagram according to a third embodiment of the present disclosure;
[0016] Figure 4 This is a training diagram of the style control network;
[0017] Figure 5 This is a training diagram of the backbone network;
[0018] Figure 6 is a schematic diagram according to a fourth embodiment of the present disclosure;
[0019] Figure 7 is a schematic diagram according to a fifth embodiment of the present disclosure;
[0020] Figure 8 It is a block diagram of an electronic device used to implement the training method of the speech synthesis model or the speech synthesis method of the embodiment of the present disclosure. DETAILED DESCRIPTION
[0021] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0022] Current speech synthesis models are primarily trained by combining style description sample text, input sample text, and output speech sample text. The loss function of the speech synthesis model is determined based on the output speech sample and the predicted speech sample, taking fewer features into account, resulting in low accuracy in the trained speech synthesis model.
[0023] In response to the above problems, the present disclosure proposes a training method for a speech synthesis model, a speech synthesis method, a device, and an electronic device.
[0024] Figure 1This is a schematic diagram according to the first embodiment of the present disclosure. It should be noted that the training method of the speech synthesis model of the embodiment of the present disclosure can be applied to a training device for the speech synthesis model, and the device can be configured in an electronic device so that the electronic device can perform the training function of the speech synthesis model.
[0025] Among them, the electronic device can be any device with computing capabilities, such as a personal computer (PC), a mobile terminal, a server, etc. The mobile terminal can be, for example, a vehicle-mounted device, a mobile phone, a tablet computer, a personal digital assistant, a wearable device, a smart speaker, a server, a server cluster, and other hardware devices with various operating systems, touch screens and / or display screens.
[0026] The speech synthesis model training device may also be software in an electronic device, such as speech synthesis model training software, etc. In the following embodiments, the execution subject is an electronic device as an example for description.
[0027] like Figure 1 As shown, the training method of the speech synthesis model may include the following steps:
[0028] Step 101: Acquire training data. The training samples in the training data include: style description sample text, input sample text, output sample speech, and style dimension data of the output sample speech.
[0029] In the embodiment of the present disclosure, the style dimension data of the output sample speech may include a sample style sequence in at least one style dimension and / or a sample style content in at least one style dimension.
[0030] In one example, the style dimension data of the output speech sample may include a sample style sequence in at least one style dimension. In another example, the style dimension data of the output speech sample may include sample style content in at least one style dimension. In another example, the style dimension data of the output speech sample may include a sample style sequence and sample style content in at least one style dimension.
[0031] In which case, when the style dimension data of the output sample speech includes sample style content on at least one style dimension, the sample style content on each style dimension can be determined in at least one of the following ways: manually labeled, or identified using a style recognition model corresponding to the style dimension.
[0032] In which, when the style dimension data of the output sample speech includes a sample style sequence and sample style content on at least one style dimension, in one example, the process of the electronic device executing step 101 may be, for example, obtaining a style description sample text, inputting a sample text, and outputting a sample speech; performing style extraction processing on at least one style dimension on each speech frame in the output sample speech to obtain a sample style sequence on at least one style dimension; and determining the sample style content on at least one style dimension based on the sample style sequence on at least one style dimension.
[0033] In an embodiment of the present disclosure, the style dimension includes at least one of the following: pitch, volume, speaking speed, pitch variance, and emotion type.
[0034] Taking pitch as an example, the sample style sequence and sample style content of the output sample speech in the pitch dimension can be obtained by, for example, performing fundamental frequency extraction processing on the output sample speech to obtain a fundamental frequency signal sequence; the fundamental frequency signal sequence includes the fundamental frequency of each speech frame in the output sample speech; performing quantile discretization processing on the fundamental frequency signal sequence to obtain a sample style sequence in the pitch dimension, i.e., a pitch sequence in the pitch dimension; and determining the sample style content in the pitch dimension, i.e., the pitch value, based on the pitch sequence in the pitch dimension. Furthermore, the pitch variance value in the pitch variance dimension can also be determined based on the pitch sequence in the pitch dimension.
[0035] The respective tone values in the tone sequence may include, for example, at least one of the following: very low, low, medium, high, and very high.
[0036] Among them, taking volume as an example, the method for obtaining the sample style sequence and sample style content of the output sample speech in the volume dimension can be, for example, performing energy calculation processing on each speech frame in the output sample speech to obtain the energy on each speech frame, and then obtaining an energy time series sequence; performing quantile discretization processing on the energy time series sequence to obtain the sample style sequence in the volume dimension, that is, the volume sequence in the volume dimension; according to the volume sequence in the volume dimension, determining the sample style content in the volume dimension, that is, the volume value.
[0037] The volume values in the volume sequence may include, for example, at least one of the following: very low, low, medium, high, and very high.
[0038] Among them, taking speaking speed as an example, the method of obtaining the sample style sequence and sample style content of the output sample speech in the speaking speed dimension can be, for example, to perform phoneme-level forced alignment on the output sample speech, calculate the duration of each phoneme, and construct a rhythm feature sequence; according to the rhythm feature sequence, determine the speaking speed of each speech frame in the output sample speech, and then obtain the speaking speed time series sequence; perform quantile discretization on the speaking speed time series sequence to obtain the sample style sequence in the speaking speed dimension, that is, the speaking speed sequence in the speaking speed dimension; according to the speaking speed sequence in the speaking speed dimension, determine the sample style content in the speaking speed dimension, that is, the speaking speed value.
[0039] The speech rate values in the speech rate sequence may include, for example, at least one of the following: very low, low, medium, high, and very high.
[0040] Among them, taking emotion type as an example, the method of obtaining the sample style sequence and sample style content of the output sample speech in the emotion type dimension can be, for example, performing emotion feature vector extraction processing on each speech frame in the output sample speech to obtain an emotion feature vector sequence; performing spectral clustering processing on the emotion feature vector sequence to obtain a sample style sequence in the emotion type dimension, that is, an emotion type sequence in the emotion type dimension; and determining the sample style content in the emotion type dimension, that is, the emotion type, based on the emotion type sequence in the emotion type dimension.
[0041] Among them, the emotion type may include, for example, at least one of the following: joy, anger, sorrow, fear, surprise, disgust, sadness, heaviness, warmth, humor, vividness, etc., which can be set according to actual needs and are not specifically limited here.
[0042] Among them, style extraction processing is performed on each speech frame in the output sample speech on at least one style dimension, so that the electronic device can combine the sample style sequence and / or sample style content on at least one style dimension to train the speech synthesis model, so that when training the speech synthesis model, it can be combined with dimensional data on at least one style dimension for training processing, and more features are considered, thereby improving the accuracy of the trained speech synthesis model.
[0043] Among them, the flexible setting of multiple style dimensions allows the electronic device to select the required style dimensions according to actual needs for training and processing the speech synthesis model, thereby improving the training flexibility of the speech synthesis model.
[0044] Step 102: Obtain an initial speech synthesis model.
[0045] In an embodiment of the present disclosure, a speech synthesis model may include: a backbone network and a style control network; the style control network is used to perform style extraction processing on the input style description text to obtain a style feature vector; the backbone network is used to perform speech synthesis processing on the input text and the style feature vector to obtain output speech.
[0046] Among them, the setting and processing of the backbone network and the style control network enable the electronic device to perform targeted training on the backbone network and the style control network respectively in combination with the training data, thereby improving the accuracy of the backbone network and the accuracy of the style control network, and further improving the accuracy of the trained speech synthesis model.
[0047] In an embodiment of the present disclosure, the backbone network may include an encoding network and a decoding network; a hierarchical cross-attention mechanism is provided in the encoding network for determining a predicted style sequence on at least one style dimension based on an intermediate feature vector and a second predicted style vector; the intermediate feature vector is an intermediate vector in the encoding network processing process.
[0048] Among them, the setting of the hierarchical cross-attention mechanism enables the encoding network to extract more features from the second predicted style vector, thereby further improving the accuracy of the trained speech synthesis model.
[0049] The backbone network may be, for example, a fish-speech text-to-speech model, and the style control network may be, for example, a BERT model.
[0050] Step 103 : Training the speech synthesis model based on the style description sample text, the input sample text, the output sample speech, and the style dimension data to obtain a trained speech synthesis model.
[0051] In an embodiment of the present disclosure, in one example, the process of the electronic device executing step 103 may be, for example, inputting the style description sample text into the style control network in the speech synthesis model to obtain the predicted style vector output by the style control network; inputting the predicted style vector and the input sample text into the backbone network in the speech synthesis model to obtain the predicted style dimension data and the output predicted speech output by the backbone network; and combining the predicted style dimension data, the style dimension data, the output predicted speech and the output sample speech to perform parameter adjustment processing on the speech synthesis model to obtain the trained speech synthesis model.
[0052] The training method of the speech synthesis model of the embodiment of the present disclosure is through obtaining training data; the training samples in the training data include: style description sample text, input sample text, output sample speech and style dimension data of the output sample speech; obtaining an initial speech synthesis model; training the speech synthesis model according to the style description sample text, input sample text, output sample speech and style dimension data to obtain a trained speech synthesis model; wherein, the electronic device can train the speech synthesis model in combination with the style dimension data, so that more features can be considered when training the speech synthesis model, thereby improving the accuracy of the trained speech synthesis model.
[0053] Among them, in order to further improve the accuracy of the trained speech synthesis model, the electronic device can perform step-by-step training on the backbone network and style control network in the speech synthesis model in combination with the training samples. First, the style control network is trained, and then the backbone network is trained based on the trained style control network. Figure 2 As shown, Figure 2 is a schematic diagram according to a second embodiment of the present disclosure, Figure 2 The illustrated embodiment may include the following steps:
[0054] Step 201: Acquire training data. The training samples in the training data include: style description sample text, input sample text, output sample speech, and style dimension data of the output sample speech. The style dimension data includes the style content of the output sample speech in at least one style dimension.
[0055] Step 202: Obtain an initial speech synthesis model.
[0056] Step 203 : Training a style control network in the speech synthesis model based on the style description sample text and the sample style content of the output sample speech in at least one style dimension to obtain a trained style control network.
[0057] In the embodiment of the present disclosure, the process of the electronic device executing step 203 may, for example, include inputting the style description sample text into the style control network to obtain a first predicted style vector output by the style control network; determining the predicted style content in at least one style dimension based on the first predicted style vector; and performing parameter adjustment processing on the style control network based on the predicted style content in at least one style dimension and the sample style content to obtain a trained style control network.
[0058] The electronic device may input the first predicted style vector into a classification network of at least one style dimension to obtain predicted style content on at least one style dimension.
[0059] Among them, the electronic device can determine the value of the loss function based on the predicted style content, sample style content and loss function of the style control network in at least one style dimension; adjust the parameters of the style control network according to the value of the loss function to obtain a trained style control network.
[0060] The style control network is trained in combination with sample style content in at least one style dimension, so that the style control network can learn accurate style vectors from style description sample texts, which are used to determine the style content in at least one style dimension, so that the style control network can learn more style-related features, thereby further improving the accuracy of the trained style control network.
[0061] Step 204 : training the backbone network in the speech synthesis model based on the style description sample text, the input sample text, the output sample speech, and the trained style control network to obtain the trained backbone network.
[0062] In an embodiment of the present disclosure, in one example, the process of the electronic device executing step 204 may be, for example, inputting the style description sample text into the trained style control network to obtain the second predicted style vector output by the style control network; inputting the second predicted style vector and the input sample text into the backbone network to obtain the output predicted speech output by the backbone network; and performing parameter adjustment processing on the backbone network according to the output predicted speech and the output sample speech to obtain the trained backbone network.
[0063] The electronic device can specifically determine the loss value based on the output predicted speech and the output sample speech; and perform parameter adjustment processing on the backbone network according to the loss value to obtain the trained backbone network.
[0064] In another example, the process of the electronic device executing step 204 can be, for example, inputting the style description sample text into the trained style control network to obtain the second predicted style vector output by the style control network; inputting the second predicted style vector and the input sample text into the backbone network to obtain the predicted style sequence and output predicted speech on at least one style dimension output by the backbone network; and performing parameter adjustment processing on the backbone network based on the predicted style sequence, sample style sequence, output predicted speech and output sample speech on at least one style dimension to obtain the trained backbone network.
[0065] Among them, the electronic device can specifically determine a first loss value based on a predicted style sequence and a sample style sequence on at least one style dimension; determine a second loss value based on the output predicted speech and the output sample speech; and perform parameter adjustment processing on the backbone network based on the first loss value and the second loss value to obtain a trained backbone network.
[0066] Among them, the electronic device combines the predicted style sequence and the sample style sequence on at least one style dimension to adjust the parameters of the backbone network, so that the backbone network can learn style-related knowledge and thus learn more features; the electronic device adjusts the parameters of the backbone network according to the first loss function value and the second loss function value, so that the backbone network can simultaneously learn style-related knowledge and speech-related knowledge, thereby further improving the accuracy of the trained backbone network.
[0067] It should be noted that the details of steps 201 to 202 can be found in Figure 1 Steps 101 to 102 in the illustrated embodiment will not be described in detail here.
[0068] The training method of the speech synthesis model of the embodiment of the present disclosure is through obtaining training data; the training samples in the training data include: style description sample text, input sample text, output sample speech and style dimension data of the output sample speech; the style dimension data includes the style content of the output sample speech in at least one style dimension; obtaining an initial speech synthesis model; training the style control network in the speech synthesis model according to the style description sample text and the sample style content of the output sample speech in at least one style dimension to obtain a trained style control network; training the backbone network in the speech synthesis model according to the style description sample text, the input sample text, the output sample speech and the trained style control network to obtain a trained backbone network; wherein, the electronic device performs step-by-step training on the backbone network and the style control network in the speech synthesis model in combination with the training samples; firstly train the style control network, and then train the backbone network based on the trained style control network, so as to further improve the accuracy of the trained speech synthesis model.
[0069] Figure 3 This is a schematic diagram according to the third embodiment of the present disclosure. It should be noted that the speech synthesis method of the embodiment of the present disclosure can be applied to a speech synthesis device, which can be configured in an electronic device so that the electronic device can perform a speech synthesis function.
[0070] Among them, the electronic device can be any device with computing capabilities, such as a personal computer (PC), a mobile terminal, a server, etc. The mobile terminal can be, for example, a vehicle-mounted device, a mobile phone, a tablet computer, a personal digital assistant, a wearable device, a smart speaker, a server, a server cluster, and other hardware devices with various operating systems, touch screens and / or display screens.
[0071] The speech synthesis device may also be software in an electronic device, such as speech synthesis software, etc. In the following embodiments, the execution subject is an electronic device as an example for description.
[0072] like Figure 3 As shown, the speech synthesis method may include the following steps:
[0073] Step 301: Obtain input text and style description text.
[0074] The style description text is used to describe the voice style. Examples of the style description text include "expressing sad emotions," "using a heavy tone," "adopting a vivid tone," and "reading the message in a warm tone."
[0075] Step 302: Input the style description text into the style control network in the speech synthesis model to obtain the style vector output by the style control network; the speech synthesis model is trained based on the style description sample text, input sample text, output sample speech, and style dimension data of the output sample speech.
[0076] In an embodiment of the present disclosure, a speech synthesis model may include: a backbone network and a style control network; the style control network is used to perform style extraction processing on the input style description text to obtain a style feature vector; the backbone network is used to perform speech synthesis processing on the input text and the style feature vector to obtain output speech.
[0077] Among them, the training method of the speech synthesis model can be referred to Figures 1 to 2 The training method of the speech synthesis model in the embodiment will not be introduced in detail here.
[0078] The style dimension data of the output sample speech may include a sample style sequence in at least one style dimension and / or a sample style content in at least one style dimension. The sample style sequence in a style dimension may include the sample style content in the style dimension of each speech frame in the output sample speech.
[0079] Step 303: Input the style vector and the input text into the backbone network of the speech synthesis model, and obtain the output speech corresponding to the input text output by the backbone network.
[0080] In an embodiment of the present disclosure, the backbone network may include an encoding network and a decoding network; a hierarchical cross-attention mechanism is provided in the encoding network for determining a predicted style sequence on at least one style dimension based on an intermediate feature vector and a second predicted style vector; the intermediate feature vector may be an intermediate vector generated by the encoding network during the processing of the input text.
[0081] The backbone network can be, for example, a fish-speech text-to-speech model. The style control network can be, for example, a BERT model. In the fish-speech model, the encoding network can include a slow autoregressive network and a fast autoregressive network. The slow autoregressive network and the fast autoregressive network can each be equipped with a layered cross-attention mechanism.
[0082] Among them, the slow autoregressive network focuses on modeling long-range dependencies across language contexts and is used to extract high-dimensional features; the fast autoregressive network is responsible for the autoregressive generation of low-dimensional acoustic features and is used to extract coarse-grained acoustic features.
[0083] The speech synthesis method of the embodiment of the present disclosure obtains input text and style description text; inputs the style description text into the style control network in the speech synthesis model to obtain the style vector output by the style control network; the speech synthesis model is trained based on the style description sample text, input sample text, output sample speech and style dimension data of the output sample speech; inputs the style vector and the input text into the backbone network in the speech synthesis model to obtain the output speech corresponding to the input text output by the backbone network; wherein the speech synthesis model can take into account style-related features, thereby improving the style matching degree between the determined output speech and the style description text, and thereby improving the accuracy of the output speech.
[0084] The following examples are given to illustrate this. Figure 4 The following is a training diagram of the style control network. Figure 4 In , the input of BERT (i.e., style control network) can be the style description text or the text vector corresponding to the style description text. Figure 4 In the example, CLS instructs BERT to output a style embedding; T1 through TM represent the individual characters in the style description text. Combining the style embedding allows us to determine the predicted style content across five style dimensions: emotion, pitch, volume, speaking rate, and pitch variation. BERT is then fine-tuned based on the sample style content across these five dimensions.
[0085] The following examples are given to illustrate this. Figure 5 The following is a diagram of the backbone network training. Figure 5In the BERT model, the Natural Language Description (i.e., the style description text) is fed into BERT to obtain the Style Embedding (style vector). The Input text (i.e., the input text) is fed into the Slow Transformer + Fast Transformer (the encoding network in the backbone network) to obtain the predicted style sequence (i.e., the 4 style dimensions) output by the Slow Transformer. Figure 5 The backbone network is parameterized based on the predicted style sequences in the four style dimensions.
[0086] Among them, SlowTransformer determines the predicted style sequence on the four style dimensions based on the processed Hidden States (intermediate feature vector), StyleEmbedding and layered cross attention mechanism. FastTransformer performs further feature extraction processing based on the processed intermediate feature vector, StyleEmbedding and layered cross attention mechanism, and then performs speech synthesis processing. Among them, Figure 5 The duration in can represent a speech rate sequence.
[0087] In order to implement the above embodiment, the present disclosure also provides a training device for a speech synthesis model. Figure 6 As shown, Figure 6 6 is a schematic diagram of the fourth embodiment of the present disclosure. The speech synthesis model training device 60 may include: a first acquisition module 601 , a second acquisition module 602 and a training processing module 603 .
[0088] Among them, the first acquisition module 601 is used to obtain training data; the training samples in the training data include: style description sample text, input sample text, output sample speech and style dimension data of the output sample speech; the second acquisition module 602 is used to obtain the initial speech synthesis model; the training processing module 603 is used to train the speech synthesis model according to the style description sample text, the input sample text, the output sample speech and the style dimension data to obtain a trained speech synthesis model.
[0089] As a possible implementation method of an embodiment of the present disclosure, the speech synthesis model includes: a backbone network and a style control network; the style control network is used to perform style extraction processing on the input style description text to obtain a style feature vector; the backbone network is used to perform speech synthesis processing on the input text and the style feature vector to obtain output speech.
[0090] As a possible implementation method of an embodiment of the present disclosure, the style dimension data includes: the sample style content of the output sample speech in at least one style dimension; the training processing module 603 is specifically used to train the style control network in the speech synthesis model according to the style description sample text and the sample style content of the output sample speech in at least one style dimension to obtain a trained style control network; and train the backbone network in the speech synthesis model according to the style description sample text, the input sample text, the output sample speech and the trained style control network to obtain a trained backbone network.
[0091] As a possible implementation method of an embodiment of the present disclosure, the training processing module 603 is further specifically used to input the style description sample text into the style control network to obtain a first predicted style vector output by the style control network; determine the predicted style content in at least one style dimension based on the first predicted style vector; and perform parameter adjustment processing on the style control network based on the predicted style content in at least one style dimension and the sample style content to obtain the trained style control network.
[0092] As a possible implementation method of an embodiment of the present disclosure, the style dimension data also includes: a sample style sequence of the output sample speech in at least one style dimension; the sample style sequence includes the style content of each speech frame in the output sample speech; the training processing module 603 is specifically further used to input the style description sample text into the trained style control network to obtain the second predicted style vector output by the style control network; input the second predicted style vector and the input sample text into the backbone network to obtain the predicted style sequence and output predicted speech in at least one style dimension output by the backbone network; according to the predicted style sequence, the sample style sequence, the output predicted speech and the output sample speech in at least one style dimension, the backbone network is parameter adjusted to obtain the trained backbone network.
[0093] As a possible implementation method of an embodiment of the present disclosure, the training processing module 603 is further specifically used to determine a first loss value based on the predicted style sequence and the sample style sequence on at least one style dimension; determine a second loss value based on the output predicted speech and the output sample speech; and perform parameter adjustment processing on the backbone network based on the first loss value and the second loss value to obtain the trained backbone network.
[0094] As a possible implementation method of an embodiment of the present disclosure, the backbone network includes an encoding network and a decoding network; a hierarchical cross-attention mechanism is provided in the encoding network and / or the decoding network, for determining a predicted style sequence on at least one style dimension based on an intermediate feature vector and the second predicted style vector; the intermediate feature vector is an intermediate vector in the processing process of the encoding network and / or the decoding network.
[0095] As a possible implementation method of an embodiment of the present disclosure, the style dimension data includes: sample style content and sample style sequence of the output sample speech in at least one style dimension; the first acquisition module is specifically used to obtain the style description sample text, the input sample text and the output sample speech; perform style extraction processing on each speech frame in the output sample speech in at least one style dimension to obtain a sample style sequence in at least one style dimension; and determine the sample style content in at least one style dimension based on the sample style sequence in at least one style dimension.
[0096] As a possible implementation of an embodiment of the present disclosure, the style dimension includes at least one of the following: pitch, volume, speaking speed, pitch variance, and emotion type.
[0097] The training device of the speech synthesis model of the embodiment of the present disclosure obtains training data; the training samples in the training data include: style description sample text, input sample text, output sample speech and style dimension data of the output sample speech; obtains an initial speech synthesis model; trains the speech synthesis model according to the style description sample text, input sample text, output sample speech and style dimension data to obtain a trained speech synthesis model; wherein, the electronic device can train the speech synthesis model in combination with the style dimension data, so that more features can be considered when training the speech synthesis model, thereby improving the accuracy of the trained speech synthesis model.
[0098] In order to implement the above embodiment, the present disclosure also provides a speech synthesis device. Figure 7 As shown, Figure 7 Schematic diagram of the fifth embodiment of the present disclosure. The speech synthesis device 70 may include: a first acquisition module 701 , a second acquisition module 702 , and a third acquisition module 703 .
[0099] Among them, the first acquisition module 701 is used to obtain input text and style description text; the second acquisition module 702 is used to input the style description text into the style control network in the speech synthesis model to obtain the style vector output by the style control network; the speech synthesis model is trained based on the style description sample text, input sample text, output sample speech and style dimension data of the output sample speech; the third acquisition module 703 is used to input the style vector and the input text into the backbone network in the speech synthesis model to obtain the output speech corresponding to the input text output by the backbone network.
[0100] The speech synthesis device of the embodiment of the present disclosure obtains input text and style description text; inputs the style description text into the style control network in the speech synthesis model to obtain the style vector output by the style control network; the speech synthesis model is trained based on the style description sample text, input sample text, output sample speech and style dimension data of the output sample speech; inputs the style vector and the input text into the backbone network in the speech synthesis model to obtain the output speech corresponding to the input text output by the backbone network; wherein the speech synthesis model can take into account style-related features, thereby improving the style matching degree between the determined output speech and the style description text, and thereby improving the accuracy of the output speech.
[0101] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information are all carried out with the user's consent, comply with relevant laws and regulations, and do not violate public order and good morals.
[0102] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0103] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0104] like Figure 8As shown, the device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. Various programs and data required for the operation of the device 800 can also be stored in the RAM 803. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0105] Various components in device 800 are connected to I / O interface 805, including an input unit 806, such as a keyboard, mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, optical disk, etc.; and a communication unit 809, such as a network card, modem, wireless communication transceiver, etc. The communication unit 809 allows device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0106] The computing unit 801 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 801 performs the various methods and processes described above, such as the training method of the speech synthesis model or the speech synthesis method. For example, in some embodiments, the training method of the speech synthesis model or the speech synthesis method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the training method of the speech synthesis model or the speech synthesis method described above can be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to execute the training method of the speech synthesis model or the speech synthesis method in any other appropriate manner (for example, by means of firmware).
[0107] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0108] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0109] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0110] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0111] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0112] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0113] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.
[0114] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A method for training a speech synthesis model, the method comprising: Get training data; The training samples in the training data include: style description sample text, input sample text, output sample speech and style dimension data of the output sample speech; Obtain the initial speech synthesis model; The speech synthesis model is trained according to the style description sample text, the input sample text, the output sample speech and the style dimension data to obtain a trained speech synthesis model.
2. The method according to claim 1, wherein The speech synthesis model includes: a backbone network and a style control network; The style control network is used to perform style extraction processing on the input style description text to obtain a style feature vector; The backbone network is used to perform speech synthesis processing on the basis of the input text and the style feature vector to obtain output speech.
3. The method according to claim 1 or 2, wherein The style dimension data includes: sample style content of the output sample speech in at least one style dimension; the training process of the speech synthesis model based on the style description sample text, the input sample text, the output sample speech, and the style dimension data to obtain a trained speech synthesis model includes: Training a style control network in the speech synthesis model based on the style description sample text and the sample style content of the output sample speech in at least one style dimension to obtain a trained style control network; The backbone network in the speech synthesis model is trained based on the style description sample text, the input sample text, the output sample speech and the trained style control network to obtain the trained backbone network.
4. The method according to claim 3, wherein: The training process of the style control network in the speech synthesis model according to the style description sample text and the sample style content of the output sample speech in at least one style dimension to obtain a trained style control network includes: Inputting the style description sample text into the style control network to obtain a first predicted style vector output by the style control network; determining predicted style content in at least one style dimension according to the first predicted style vector; According to the predicted style content and the sample style content in at least one style dimension, parameter adjustment processing is performed on the style control network to obtain the trained style control network.
5. The method according to claim 3, wherein: The style dimension data further includes: a sample style sequence of the output sample speech in at least one style dimension; the sample style sequence includes the style content of each speech frame in the output sample speech; the training process of the backbone network in the speech synthesis model based on the style description sample text, the input sample text, the output sample speech, and the trained style control network to obtain the trained backbone network includes: Inputting the style description sample text into the trained style control network to obtain a second predicted style vector output by the style control network; Inputting the second predicted style vector and the input sample text into the backbone network, obtaining a predicted style sequence on at least one style dimension output by the backbone network and outputting predicted speech; According to the predicted style sequence, the sample style sequence, the output predicted speech, and the output sample speech in at least one style dimension, parameter adjustment processing is performed on the backbone network to obtain the trained backbone network.
6. The method according to claim 5, wherein: The step of performing parameter adjustment processing on the backbone network according to the predicted style sequence, the sample style sequence, the output predicted speech, and the output sample speech in at least one style dimension to obtain the trained backbone network includes: determining a first loss value according to the predicted style sequence and the sample style sequence in at least one style dimension; Determining a second loss value according to the output predicted speech and the output sample speech; Parameter adjustment processing is performed on the backbone network according to the first loss value and the second loss value to obtain the trained backbone network.
7. The method according to claim 5, wherein: The backbone network includes an encoding network; the encoding network is provided with a hierarchical cross attention mechanism for determining a predicted style sequence in at least one style dimension based on the intermediate feature vector and the second predicted style vector; The intermediate feature vector is an intermediate vector in the encoding network processing process.
8. The method according to claim 1, wherein The style dimension data includes: sample style content and sample style sequence of the output sample speech in at least one style dimension; the acquiring training data includes: Obtaining the style description sample text, the input sample text, and the output sample speech; Performing style extraction processing on each speech frame in the output sample speech in at least one style dimension to obtain a sample style sequence in at least one style dimension; The sample style content in at least one style dimension is determined according to the sample style sequence in at least one style dimension.
9. The method according to claim 2 or 8, wherein: The style dimension includes at least one of the following: pitch, volume, speaking speed, pitch variance, and emotion type.
10. A speech synthesis method, comprising: Get the input text and style description text; Inputting the style description text into a style control network in a speech synthesis model to obtain a style vector output by the style control network; the speech synthesis model is trained based on the style description sample text, input sample text, output sample speech, and style dimension data of the output sample speech; The style vector and the input text are input into a backbone network in the speech synthesis model, and an output speech corresponding to the input text output by the backbone network is obtained.
11. A training device for a speech synthesis model, comprising: A first acquisition module is used to acquire training data; The training samples in the training data include: style description sample text, input sample text, output sample speech and style dimension data of the output sample speech; The second acquisition module is used to obtain an initial speech synthesis model; The training processing module is used to train the speech synthesis model according to the style description sample text, the input sample text, the output sample speech and the style dimension data to obtain a trained speech synthesis model.
12. A speech synthesis device, comprising: A first acquisition module is used to acquire input text and style description text; A second acquisition module is configured to input the style description text into a style control network in a speech synthesis model to obtain a style vector output by the style control network; the speech synthesis model is trained based on the style description sample text, input sample text, output sample speech, and style dimension data of the output sample speech; The third acquisition module is used to input the style vector and the input text into the backbone network in the speech synthesis model, and obtain the output speech corresponding to the input text output by the backbone network.
13. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 9; Alternatively, the method according to claim 10 is performed.
14. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable the computer to execute the method according to any one of claims 1 to 9; or to execute the method according to claim 10.
15. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method according to any one of claims 1 to 9; or implements the method according to claim 10.