Method for training a speech synthesis model, speech synthesis method, apparatus, and electronic device
The speech synthesis model training method addresses high training costs and low generality by using semantic coding and decoding networks to fuse style and timbre features, enhancing efficiency and accuracy in voice synthesis.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-12-19
- Publication Date
- 2026-03-17
AI Technical Summary
Current voice synthesis models require high training costs due to limited expressiveness of style tags or style prompt texts, leading to low generality in handling new styles or prompts.
A method for training a speech synthesis model that includes acquiring training data with style and timbre sample voices and texts, using a semantic coding network and decoding network to extract and fuse features, reducing the need for individual training on each style or timbre, and incorporating an autoregressive semantic feature extraction module for improved accuracy and efficiency.
The method reduces training costs and enhances the model's generality in handling new styles and timbres, improving the efficiency and accuracy of voice synthesis.
Smart Images

Figure 2026048914000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, particularly to technical fields such as deep learning, natural language processing, voice technology, large-scale models, etc., and particularly to a training method, a voice synthesis method, an apparatus, and an electronic device for a voice synthesis model.
Background Art
[0002] In current voice synthesis models, the model structure is a backbone network + a style control module. Here, the input of the style control module is a style tag or a style prompt text. Here, the style tag or the style prompt text has limited expressiveness, and in the model training process, it is necessary to perform training processing for each of the style tags or the style prompt texts, resulting in a high training cost. Therefore, the voice synthesis model obtained through training has low generality in new style tags or style prompt texts.
Summary of the Invention
Problems to be Solved by the Invention
[0003] The present disclosure provides a training method, a voice synthesis method, an apparatus, and an electronic device for a voice synthesis model.
[0004] According to one aspect of the present disclosure, a method for training a speech synthesis model is provided, the method comprising: acquiring training data, wherein the training samples in the training data include style sample speech, timbre sample speech, input sample text, and output sample speech; acquiring an initial speech synthesis model, wherein the speech synthesis model includes a semantic coding network and a semantic decoding network; performing a training process on the semantic coding network and the semantic decoding network based on the style sample speech, the timbre sample speech, the input sample text, and the output sample speech, respectively, to acquire a trained semantic coding network and a trained semantic decoding network; and acquiring a trained speech synthesis model based on the trained semantic coding network and the trained semantic decoding network.
[0005] Another aspect of the present disclosure provides a speech synthesis method comprising: acquiring a style voice, a timbre voice, and an input text to be processed; acquiring a speech synthesis model, wherein a semantic coding network and a semantic decoding network in the speech synthesis model are acquired by training on a style sample voice, a timbre sample voice, an input sample text, and an output sample voice, respectively; inputting the style voice, the timbre voice, and the input text into the semantic coding network in the speech synthesis model to acquire a semantic feature vector sequence output by the semantic coding network; and inputting the semantic feature vector sequence into the semantic decoding network to acquire an output voice corresponding to the input text, which is output by the semantic decoding network.
[0006] According to another aspect of the present disclosure, a speech synthesis model training apparatus is provided, the apparatus comprising: a first acquisition module for acquiring training data, wherein the training samples in the training data include style sample speech, timbre sample speech, input sample text, and output sample speech; a second acquisition module for acquiring an initial speech synthesis model, wherein the speech synthesis model includes a semantic coding network and a semantic decoding network; a training processing module for performing training on the semantic coding network and the semantic decoding network based on the style sample speech, the timbre sample speech, the input sample text, and the output sample speech, respectively, and acquiring a trained semantic coding network and a trained semantic decoding network; and a third acquisition module for acquiring a trained speech synthesis model based on the trained semantic coding network and the trained semantic decoding network.
[0007] According to another aspect of the present disclosure, a speech synthesis apparatus is provided, the apparatus comprising: a first acquisition module for acquiring style speech, timbre speech and input text to be processed; a second acquisition module for acquiring a speech synthesis model, wherein the semantic coding network and semantic decoding network in the speech synthesis model are acquired by training them based on style sample speech, timbre sample speech, input sample text and output sample speech, respectively; a third acquisition module for inputting the style speech, timbre speech and input text into the semantic coding network in the speech synthesis model and acquiring a semantic feature vector sequence output by the semantic coding network; and a fourth acquisition module for inputting the semantic feature vector sequence into the semantic decoding network and acquiring an output speech corresponding to the input text output by the semantic decoding network.
[0008] According to another aspect of the present disclosure, an electronic device is provided, comprising at least one processor and a memory communicably connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor so that the at least one processor can perform the speech synthesis model training method or speech synthesis method described above in the present disclosure.
[0009] According to another aspect of the present disclosure, a non-temporary computer-readable storage medium is provided which stores computer instructions, the computer instructions causing the computer to execute the above-described method for training a speech synthesis model or a speech synthesis method of the present disclosure.
[0010] According to another aspect of the present disclosure, a computer program is provided which, when executed by a processor, implements the steps of the above-described steps of the training method for a speech synthesis model or a speech synthesis method.
[0011] Furthermore, the content described in this section does not identify any essential or important features of the embodiments of this disclosure, nor does it limit the scope of this disclosure. Other features of this disclosure are more readily apparent from the following specification. [Brief explanation of the drawing]
[0012] The drawings are provided to better understand this solution and do not limit the scope of this disclosure. [Figure 1] This is a schematic diagram relating to the first embodiment of the present disclosure. [Figure 2] This is a schematic diagram relating to a second embodiment of the present disclosure. [Figure 3] This is a schematic diagram relating to a third embodiment of the present disclosure. [Figure 4] This is a schematic diagram relating to the fourth embodiment of the present disclosure. [Figure 5] This is a schematic block diagram of the speech synthesis model. [Figure 6] This is a schematic diagram relating to the fifth embodiment of the present disclosure. [Figure 7] This is a schematic diagram relating to the sixth embodiment of the present disclosure. [Figure 8] This is a block diagram of an electronic device for a speech synthesis model training method or speech synthesis method for realizing an embodiment of the present disclosure. [Modes for carrying out the invention]
[0013] Illustrative embodiments of the present disclosure are described below in conjunction with the drawings, and various details of the embodiments of the present disclosure are included herefor the sake of clarity, and these details should be considered illustrative. Thus, various changes and modifications can be made to the embodiments described herein, as will be understood by those skilled in the art, without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0014] In current speech synthesis models, the model structure consists of a backbone network plus a style control module. Here, the input to the style control module is a style tag or style-presented text. However, since style tags or style-presented texts have limited expressive power, training processing must be performed for each style tag or style-presented text during model training, resulting in a high training cost and low generality of the speech synthesis model obtained from training with new style tags or style-presented texts.
[0015] To address the above issues, this disclosure provides a method for training a speech synthesis model, a speech synthesis method, an apparatus, and an electronic device.
[0016] Figure 1 is a schematic diagram relating to a first embodiment of the present disclosure. The speech synthesis model training method of the embodiment of the present disclosure is applicable to a speech synthesis model training device, and by placing the device in an electronic device, the electronic device can perform the speech synthesis model training function.
[0017] The electronic device may be any device with computing capabilities, for example, a personal computer (PC), a mobile terminal, a server, etc. The mobile terminal may be, for example, an in-vehicle device, a mobile phone, a tablet, a personal digital assistant, a wearable device, a smart speaker, a server, a server cluster, etc., and may be a hardware device having various operating systems, a touch screen and / or a display.
[0018] The training device for the voice synthesis model may be software within the electronic device, for example, training software for the voice synthesis model. In the following embodiments, the case where the execution entity is an electronic device will be described as an example.
[0019] As shown in FIG. 1, the training method for the voice synthesis model may include the following steps 101 to 104.
[0020] In step 101, training data is acquired. The training samples in the training data include style sample voices, timbre sample voices, input sample texts, and output sample voices.
[0021] In an example of the embodiments of the present disclosure, the process in which the electronic device executes step 101 may be, for example, to acquire style sample voices, timbre sample voices, and input sample texts, to acquire a reference style voice and a reference timbre voice, where the reference style voice has the same style as the style of the style sample voice, the reference timbre voice has the same timbre as the timbre of the timbre sample voice, and the text corresponding to the reference timbre voice is the input sample text, and to generate an output sample voice based on the reference style voice and the reference timbre voice.
[0022] Here, the process of obtaining the reference style voice may be, for example, determining a style description text corresponding to the style sample voice, and selecting a voice from a plurality of voices having the style described by the style description text as the reference style voice. Here, the selection condition is that the content of the text obtained by performing voice recognition on the selected voice is different from the content of the text obtained by performing voice recognition on the style sample voice.
[0023] Here, the process of obtaining the reference timbre voice may be, for example, determining timbre information corresponding to the reference timbre voice, and selecting a candidate voice from a plurality of voices having the timbre information as the reference timbre voice. Here, the selection condition is that the content of the text obtained by performing voice recognition on the selected voice is the same as the content of the input sample text.
[0024] The process by which the electronic device generates an output sample voice based on the reference style voice and the reference timbre voice may be, for example, inputting the reference style voice and the reference timbre voice into a voice fusion model, and obtaining the output sample voice generated by the voice fusion model. Here, the voice fusion model may be, for example, a voice conversion model (Voice Conversion model, VC) or the like.
[0025] Here, performing the generation process of the output sample voice by combining the reference style voice and the reference timbre voice can improve the efficiency of preparing the training sample while reducing the cost of preparing the training sample.
[0026] In embodiments of the present disclosure, to further reduce the cost of preparing training samples, the electronic device can obtain matching style sample audio by querying a style audio library based on style description sample text. Alternatively, the process by which the electronic device obtains style sample audio, timbre sample audio, and input sample text may involve, for example, obtaining style description sample text, timbre sample audio, and input sample text, obtaining a style audio library which includes at least one style description text and style audio corresponding to the style description text, querying the style audio library based on the style description sample text, obtaining a first style description sample that matches the style description sample text within the style audio library, and determining the style audio corresponding to the first style description sample as the style sample audio.
[0027] Matching a style description sample text to a first style description sample may mean at least one of the following: the style description sample text is the same as the first style description sample; the similarity between the style description sample text and the first style description sample is greater than or equal to a similarity threshold; or the similarity between the style described by the style description sample text and the style described by the first style description sample is greater than or equal to a similarity threshold.
[0028] The style could be, for example, bright and humorous yet logical, or warm and comforting like a friend. For example, if the style is bright and humorous yet logical, the corresponding style description sample text could be, for example, "speaks in a bright and humorous yet logical tone." For example, if the style is warm and comforting like a friend, the corresponding style description sample text could be, for example, "speaks in a warm and comforting tone like a friend."
[0029] In another example, the process by which an electronic device performs step 101 may involve, for example, obtaining a style sample voice, a timbre sample voice, and an input sample text, and then selecting a voice from a large number of candidate voices based on the style sample voice, timbre sample voice, and input sample text to be used as an output sample voice. Here, the selection criteria may be that the text corresponding to the selected voice is the input sample text, the style of the voice is the same as the style of the style sample voice, and the timbre of the voice is the same as the timbre of the timbre sample voice.
[0030] In step 102, an initial speech synthesis model is obtained, which includes a semantic coding network and a semantic decoding network.
[0031] In embodiments of this disclosure, a semantic coding network is used to extract a semantic feature vector sequence based on a style sample speech, a timbre sample speech, and an input sample text to obtain a predicted semantic feature vector sequence. A semantic decoding network is used to perform semantic decoding on the predicted semantic feature vector sequence to obtain a predicted output speech.
[0032] In step 103, the semantic coding network and semantic decoding network are trained based on the style sample audio, timbre sample audio, input sample text, and output sample audio, respectively, and the trained semantic coding network and trained semantic decoding network are obtained.
[0033] In embodiments of the present disclosure, the process by which the electronic device performs step 103 may, for example, involve training a semantic coding network based on style sample audio, timbre sample audio, input sample text, and output sample audio, obtaining the trained semantic coding network, and training a semantic decoding network based on style sample audio, timbre sample audio, input sample text, output sample audio, and the trained semantic coding network, obtaining the trained semantic decoding network.
[0034] Here, by setting the style sample audio and timbre sample audio, the semantic coding network can extract and fuse style and timbre features by combining the style sample audio and timbre sample audio. By performing semantic feature extraction, the trained semantic coding network can then extract and fuse style and timbre features for new styles and / or timbres by combining them with the extracted audio, thus avoiding prior training for new styles and / or timbres. This improves style and timbre transfer capabilities and supports generalization of new styles and / or timbres.
[0035] In step 104, a trained speech synthesis model is obtained based on the trained semantic coding network and the trained semantic decoding network.
[0036] The method for training a speech synthesis model in the embodiments of this disclosure involves acquiring training data, the training samples in the training data including style sample speech, timbre sample speech, input sample text, and output sample speech; acquiring an initial speech synthesis model, the speech synthesis model including a semantic coding network and a semantic decoding network; training the semantic coding network and the semantic decoding network based on the style sample speech, timbre sample speech, input sample text, and output sample speech, respectively; acquiring a post-trained semantic coding network and a post-trained semantic decoding network; and acquiring a post-trained speech synthesis model based on the post-trained semantic coding network and a post-trained semantic decoding network. Here, the style sample speech and timbre sample speech are expressive and, in combination with semantic feature extraction of the semantic coding network, more features can be extracted and fused into the output speech obtained by synthesis, thereby avoiding training processing for each style and / or timbre, reducing the training cost of the speech synthesis model, and improving the generality of the post-trained speech synthesis model in new styles and / or timbres.
[0037] To further improve the training efficiency of the speech synthesis model and the accuracy of the speech synthesis model after training, an autoregressive semantic feature extraction module may be set in the semantic coding network of the speech synthesis model, and the sample text of the input speech coding network may be concatenated text, which is obtained by concatenating a timbre sample text corresponding to a timbre sample speech with the input sample text, thereby realizing autoregressive processing. As shown in Figure 2, Figure 2 is a schematic diagram relating to a second embodiment of the present disclosure, and the embodiment shown in Figure 2 includes the following steps 201 to 207.
[0038] In step 201, training data is acquired, and the training samples in the training data include style sample audio, timbre sample audio, input sample text, and output sample audio.
[0039] In step 202, an initial speech synthesis model is obtained, which includes a semantic coding network and a semantic decoding network, and an autoregressive semantic feature extraction module is installed in the semantic coding network.
[0040] In step 203, determine the tone sample text that corresponds to the tone sample audio.
[0041] In embodiments of this disclosure, the process by which the electronic device performs step 203 may involve inputting a timbre sample audio into a speech recognition model and obtaining timbre sample text output by the speech recognition model.
[0042] In step 204, the tone sample text and the input sample text are concatenated to obtain the concatenated sample text.
[0043] To adapt to the autoregressive semantic feature extraction module and extract a semantic feature vector sequence corresponding to the input sample text in combination with the timbre sample text, the electronic device can concatenate the timbre sample text before the input sample text, thereby obtaining the concatenated sample text.
[0044] In step 205, the semantic coding network is trained based on the style sample audio, timbre sample audio, concatenated sample text, and output sample audio.
[0045] In embodiments of this disclosure, the process by which the electronic device performs step 205 may, for example, involve performing a semantic feature extraction process on the output sample speech to obtain an output semantic feature vector sequence, inputting the style sample speech, timbre sample speech, and concatenated sample text into a semantic coding network to obtain an intermediate semantic feature vector sequence corresponding to the concatenated sample text output by the semantic coding network, extracting a predicted semantic feature vector sequence corresponding to the input sample text from the intermediate semantic feature vector sequence, and performing a parameter tuning process on the semantic coding network based on the predicted semantic feature vector sequence and the output semantic feature vector sequence to obtain a trained semantic coding network.
[0046] Based on the predicted semantic feature vector sequence and the output semantic feature vector sequence, the semantic coding network is subjected to parameter tuning to gradually reduce the difference between the predicted semantic features output by the semantic coding network and the output semantic feature vector sequence, thereby improving the accuracy of the trained semantic coding network.
[0047] In embodiments of this disclosure, the semantic coding network may include a style coding module, a timbre coding module, a text coding module, and an autoregressive semantic feature extraction module, the autoregressive semantic feature extraction module being connected to the style coding module, the timbre coding module, and the text coding module, respectively, and used to perform semantic feature extraction processing based on the style representation vector output by the style coding module, the timbre representation vector output by the timbre coding module, and the text representation vector sequence output by the text coding module, in order to obtain a semantic feature vector sequence.
[0048] The process by which an electronic device inputs style sample audio, timbre sample audio, and concatenated sample text into a semantic coding network and obtains an intermediate semantic feature vector sequence corresponding to the concatenated sample text output by the semantic coding network may, for example, involve performing style coding on the style sample audio via a style coding module to obtain a style representation vector, performing timbre coding on the timbre sample audio via a timbre coding module to obtain a timbre representation vector, performing text coding on the concatenated sample text via a text coding module to obtain a text representation vector sequence, and performing semantic feature extraction on the input style representation vector, timbre representation vector, and text representation vector sequence via an autoregressive semantic feature extraction module to obtain an intermediate semantic feature vector sequence.
[0049] The autoregressive semantic feature extraction module performs semantic feature extraction based on the style representation vector output by the style coding module, the timbre representation vector output by the timbre coding module, and the text representation vector sequence output by the text coding module. This allows for the fusion of more style and timbre features into the extracted semantic feature vector sequence, resulting in a deeper fusion of style and timbre features with text features, thereby further improving the accuracy of the trained semantic coding network.
[0050] In embodiments of the present disclosure, the position vector of each character may be considered when determining the text display vector sequence in order to improve style consistency between paragraphs in the concatenated sample text and / or to improve style consistency between characters in the concatenated sample text. In contrast, the processing method of the concatenated sample text by the text encoding module includes the steps of: performing character encoding on each character in the concatenated sample text to obtain a character vector sequence; performing character position encoding and / or paragraph position encoding on each character in the concatenated sample text to obtain a position vector sequence; and performing concatenation on the character vectors in the character vector sequence and the position vectors in the position vector sequence according to the characters to obtain the text display vector sequence output by the text encoding module.
[0051] A character vector sequence may include a character vector corresponding to each character in the concatenated sample text, and a position vector sequence may include a position vector corresponding to each character in the concatenated sample text. Here, the electronic device can obtain a display vector corresponding to each character in the concatenated sample text by concatenating the character vector and position vector corresponding to that character, and then perform a combination process on the display vectors corresponding to each character to obtain a text display vector sequence.
[0052] In embodiments of the present disclosure, in order to further improve the accuracy of the determined predicted semantic feature vector sequence, the representation vectors of each character preceding the character can be determined when determining the semantic feature vector corresponding to each character. In contrast, the autoregressive semantic feature extraction module includes sequentially connected sample context coding modules, a style-gated attention module, and a semantic output projector head module. The sample context coding module is used to determine the context representation vector corresponding to the current character for each current character to be predicted in the concatenated sample text, based on the representation vector of the current character and the representation vectors of each character preceding the current character. The style-gated attention module and the semantic output projector head module are used to predict the semantic feature vector of the current character based on the context representation vector, style representation vector, and timbre representation vector.
[0053] The process by which an electronic device predicts an intermediate semantic feature vector sequence may involve obtaining, for each current character to be predicted in the concatenated sample text, the representation vector of the current character and the representation vectors of each character preceding the current character; performing context feature extraction on the representation vector of the current character and the representation vectors of each character via a sample context coding module to obtain the context representation vector corresponding to the current character; and performing semantic feature extraction on the context representation vector, style representation vector, and timbre representation vector via a style-gated attention module and a semantic output projector head module to obtain the semantic feature vector of the current character.
[0054] If the current character is the first character in the input sample text, the contextual display vector corresponding to the first character can be determined by combining the display vector of the first character with the display vectors of each character in the concatenated sample text's tone sample text. This avoids using the display vector of the first character to perform the semantic feature vector prediction process, thereby further improving the accuracy of the semantic feature vector of the first character obtained through prediction.
[0055] In embodiments of the present disclosure, in order to ensure style tendency consistency between the semantic feature vectors of each character and to naturally vary the style of each semantic feature vector in the predicted semantic feature vector sequence, a history semantic memory coding module, an attention adjustment module, and a style adjustment gated unit are provided sequentially connected in a style-gated attention module, which is used to perform style tendency extraction processing based on an existing semantic feature vector sequence output by an autoregressive semantic feature extraction module, and the extracted style tendency is used to predict the semantic feature vector of the current character.
[0056] The input to the history semantic memory coding module may be an existing semantic feature vector sequence output by the autoregressive semantic feature extraction module, and the output may be a history memory state. Here, the input to the attention adjustment module may be a contextual representation vector corresponding to the current character and the history memory state, and the output may be a style tendency. The input to the style adjustment gated unit may be a style tendency, and it determines a style adjustment policy for the current character and, in combination with the style adjustment policy, predicts the semantic feature vector of the current character. Here, the style adjustment policy may be, for example, continuation, reinforcement, or transfer.
[0057] In embodiments of this disclosure, in order to improve the training speed of the semantic coding network and the accuracy of the semantic coding network after training, the electronic device can determine a loss function value in combination with the predicted semantic feature vector sequence and the output semantic feature vector sequence in combination with a sub-loss function of at least one dimension, and then perform parameter tuning on the semantic coding network. Alternatively, the process of performing parameter tuning on the semantic coding network based on the electronic device's predicted semantic feature vector sequence and the output semantic feature vector sequence may involve, for example, determining a value for a sub-loss function of at least one dimension among the style dimension, timbre dimension, semantic dimension, and style consistency dimension based on the predicted semantic feature vector sequence and the output semantic feature vector sequence, and then performing parameter tuning on the semantic coding network based on the value of the sub-loss function of at least one dimension to obtain the semantic coding network after training.
[0058] In the embodiments of this disclosure, the numerical value of the style-dimension sub-loss function is determined based on the style representation vector output by the style coding module and the predicted style representation vector, the predicted style representation vector is obtained by performing style extraction on the predicted semantic feature vector sequence.
[0059] The process by which an electronic device obtains a predicted style representation vector may, for example, involve inputting a predicted semantic feature vector sequence into a style extraction module and obtaining the predicted style representation vector output by the style extraction module. Here, the electronic device can determine the vector similarity based on the style representation vector and the predicted style representation vector, thereby determining the difference between the number 1 and the vector similarity as the numerical value of the style dimension sub-loss function.
[0060] The higher the vector similarity between the style representation vector and the predicted style representation vector, the smaller the value of the style dimension sub-loss function.
[0061] Determining the numerical values of the style-dimension sub-loss function and tuning the parameters of the semantic coding network can guide the network to utilize more style features.
[0062] In the embodiments of this disclosure, the numerical value of the timbre dimension sub-loss function is determined based on a timbre representation vector and a predicted timbre representation vector corresponding to a timbre sample speech, the predicted timbre representation vector is obtained by performing timbre extraction on a predicted semantic feature vector sequence.
[0063] The process by which an electronic device obtains a predicted timbre representation vector may, for example, involve inputting a predicted semantic feature vector sequence into a timbre extraction module and obtaining the predicted timbre representation vector output by the timbre extraction module. Here, the electronic device can determine vector similarity based on the timbre representation vector and the predicted timbre representation vector, and further determine the difference between the number 1 and the vector similarity as a numerical value of the timbre dimension sub-loss function.
[0064] The higher the vector similarity between the timbre representation vector and the predicted timbre representation vector, the smaller the value of the timbre dimension sub-loss function.
[0065] Determining the numerical values of the timbre-dimensional sub-loss function and tuning the parameters of the semantic coding network can guide the semantic coding network to utilize more timbre features.
[0066] The numerical values of the semantic dimension sub-loss function were determined based on the predicted semantic feature vector sequence and the output semantic feature vector sequence.
[0067] An electronic device can determine vector sequence similarity based on the predicted semantic feature vector sequence and the output semantic feature vector sequence, and further determine the difference between the number 1 and the vector similarity as a numerical value of the semantic dimension sub-loss function.
[0068] The higher the vector sequence similarity between the predicted semantic feature vector sequence and the output semantic feature vector sequence, the smaller the value of the sub-loss function for the semantic dimension.
[0069] The numerical value of the sub-loss function for the style consistency dimension is determined based on the predicted semantic feature vectors corresponding to adjacent characters in the predicted semantic feature vector sequence.
[0070] The electronic device can perform style feature extraction processing based on predicted semantic feature vectors corresponding to adjacent characters to obtain style features corresponding to adjacent characters. This allows it to determine the degree of difference between the style features corresponding to adjacent characters, and based on each degree of difference, determine the numerical value of the sub-loss function for the style consistency dimension.
[0071] The degree of difference is positively correlated with the numerical value of the sub-loss function for the style consistency dimension.
[0072] Numerical values of the sub-loss function for the style consistency dimension, as well as parameter tuning for the semantic coding network, can guide the semantic coding network to generate predictive semantic feature vector sequences with consistent styles.
[0073] In step 206, the semantic decoding network is trained based on the style sample audio, timbre sample audio, concatenated sample text, output sample audio, and the trained semantic coding network to obtain the trained semantic decoding network.
[0074] In step 207, a trained speech synthesis model is obtained based on the trained semantic coding network and the trained semantic decoding network.
[0075] For detailed information on steps 201-202 and 207, please refer to steps 101-102 and 104 in the embodiment shown in Figure 1; a detailed explanation is omitted here.
[0076] The method for training a speech synthesis model in the embodiments of this disclosure involves acquiring training data, the training samples in the training data including style sample speech, timbre sample speech, input sample text, and output sample speech; acquiring an initial speech synthesis model, the speech synthesis model including a semantic coding network and a semantic decoding network, an autoregressive semantic feature extraction module installed within the semantic coding network, determining timbre sample text corresponding to timbre sample speech, concatenating the timbre sample text and input sample text to obtain concatenated sample text, and training the semantic coding network based on the style sample speech, timbre sample speech, concatenated sample text, and output sample speech, and style sample Based on the input voice, timbre sample voice, concatenated sample text, output sample voice, and the trained semantic coding network, a training process is performed on the semantic decoding network to obtain the trained semantic decoding network. Based on the trained semantic coding network and the trained semantic decoding network, a trained speech synthesis model is obtained. Here, the semantic coding network of the speech synthesis model is equipped with an autoregressive semantic feature extraction module, and the sample text of the input speech coding network is the text obtained by concatenating the timbre sample text corresponding to the timbre sample voice and the input sample text. This further improves the training efficiency of the speech synthesis model and further improves the accuracy of the trained speech synthesis model.
[0077] To further improve the training efficiency of the speech synthesis model and the accuracy of the speech synthesis model after training, parameter adjustment processing can be performed on the semantic decoding network by combining the output prediction speech output by the post-trained semantic coding network with the output sample speech. As shown in Figure 3, Figure 3 is a schematic diagram relating to a third embodiment of the present disclosure, and the embodiment shown in Figure 3 may include the following steps 301 to 307.
[0078] In step 301, training data is acquired, and the training samples in the training data include style sample audio, timbre sample audio, input sample text, and output sample audio.
[0079] In step 302, an initial speech synthesis model is obtained, which includes a semantic coding network and a semantic decoding network.
[0080] In step 303, the semantic coding network is trained based on the style sample audio, timbre sample audio, input sample text, and output sample audio to obtain the trained semantic coding network.
[0081] In step 304, the style sample audio, timbre sample audio, and input sample text are input to the trained semantic coding network to obtain the predicted semantic feature vector sequence output by the trained semantic coding network.
[0082] In the embodiments of this disclosure, an alternative solution is to replace the input sample text in step 304 with concatenated sample text. Here, the concatenated sample text is a text obtained by concatenating the timbre sample text corresponding to the timbre sample sound with the input sample text.
[0083] In step 305, the predicted semantic feature vector sequence is input to the semantic decoding network to obtain the output predicted speech output by the semantic decoding network.
[0084] In embodiments of this disclosure, in order to improve the consistency between the style features used in the semantic coding network processing and the style features used in the semantic decoding network processing, the process by which the electronic device performs step 305 may, for example, involve obtaining a style representation vector corresponding to a style sample speech, inputting the style representation vector and the predicted semantic feature vector sequence into a semantic decoding network, and obtaining the output predicted speech output by the semantic decoding network.
[0085] The semantic decoding network includes a semantic decoding module and a vocoder. The semantic decoding module is used to decode a semantic feature vector sequence and a style representation vector to obtain a predicted phonetic feature sequence, while the vocoder is used to perform speech generation processing based on the phonetic feature sequence to obtain a predicted output speech.
[0086] The format of the predicted phonetic feature sequence may be, for example, a MEL spectral acoustic feature sequence. Furthermore, different formats of predicted acoustic feature sequences may be processed using different vocoders for speech generation.
[0087] In step 306, parameter tuning is performed on the semantic decoding network based on the predicted output speech and the output sample speech to obtain the trained semantic decoding network.
[0088] In embodiments of this disclosure, in order to improve the training speed of the semantic decoding network and the accuracy of the semantic decoding network after training, the electronic device can output predicted speech and output sample speech in combination with numerical values of a sub-loss function of at least one dimension to determine the loss function value, thereby performing parameter tuning on the semantic coding network. Alternatively, the process by which the electronic device performs step 306 may be, for example, to determine numerical values of a sub-loss function of at least one dimension among speech dimension, speech rhythm dimension, and speech consistency dimension based on the output predicted speech and output sample speech, and then perform parameter tuning on the semantic decoding network based on the numerical values of the sub-loss function of at least one dimension to obtain the semantic decoding network after training.
[0089] In the embodiments of this disclosure, the numerical value of the speech-dimension sub-loss function is determined based on the speech similarity between the output predicted speech and the output sample speech.
[0090] The speech similarity between the predicted output speech and the output sample speech can be determined based on the degree of overlap of the speech frames between the predicted output speech and the output sample speech, or based on the similarity between the speech vector of the predicted output speech and the speech vector of the output sample speech.
[0091] In the embodiments of this disclosure, the numerical value of the sub-loss function for the speech rhythm dimension is determined based on predicted rhythm data and sample rhythm data, where the predicted rhythm data is obtained by rhythm extraction from the output predicted speech, and the sample rhythm data is obtained by rhythm extraction from the output sample speech.
[0092] The rhythm data may include parameter data for at least one of the following parameters: a baseband parameter, a duration parameter, a speech intensity parameter, etc. Here, the baseband parameter is, for example, the fundamental frequency. The duration parameter is, for example, the length of pronunciation of a phoneme, syllable, or word in speech. The speech intensity parameter is, for example, the amplitude or energy in speech.
[0093] The electronic device can determine the loss value for each of the above parameters based on the parameter data, and then perform a weighted addition process on the loss values for each parameter to obtain the value of the sub-loss function for the audio rhythm dimension.
[0094] In the embodiments of this disclosure, the numerical value of the sub-loss function for the speech consistency dimension is determined based on the differences between adjacent speech frames in the output prediction speech.
[0095] The electronic device can determine the smoothness between adjacent audio frames based on adjacent audio frames in the output prediction audio, determine this smoothness as the difference between adjacent audio frames, and further determine the numerical value of the sub-loss function of the audio consistency dimension based on each difference.
[0096] In step 307, a trained speech synthesis model is obtained based on the trained semantic coding network and the trained semantic decoding network.
[0097] For detailed information on steps 301-302 and 307, please refer to steps 101-102 and 104 in the embodiment shown in Figure 1; a detailed explanation is omitted here.
[0098] The method for training a speech synthesis model in an embodiment of the present disclosure includes: acquiring training data, the training samples in the training data include style sample speech, timbre sample speech, input sample text, and output sample speech; acquiring an initial speech synthesis model, the speech synthesis model includes a semantic coding network and a semantic decoding network; training the semantic coding network based on the style sample speech, timbre sample speech, input sample text, and output sample speech to acquire a trained semantic coding network; inputting the style sample speech, timbre sample speech, and input sample text into the trained semantic coding network; and predicting the semantic feature vector output by the trained semantic coding network. A sequence is obtained, the predicted semantic feature vector sequence is input to the semantic decoding network, the output predicted speech output by the semantic decoding network is obtained, the parameters of the semantic decoding network are adjusted based on the output predicted speech and output sample speech to obtain the trained semantic decoding network, the trained speech synthesis model is obtained based on the trained semantic coding network and the trained semantic decoding network, and here, the parameters of the semantic decoding network are adjusted based on the output predicted speech and output sample speech output by the trained semantic coding network to further improve the training efficiency of the speech synthesis model and the accuracy of the trained speech synthesis model.
[0099] Figure 4 may be a schematic diagram relating to a fourth embodiment of the present disclosure. The speech synthesis method of the embodiment of the present disclosure may be applied to a speech synthesis device, and the device may be arranged in an electronic device so that the electronic device can perform the speech synthesis function.
[0100] An electronic device may be any device with computing power, such as a personal computer (PC), a mobile terminal, or a server. A mobile terminal may be a hardware device with various operating systems, touchscreens, and / or displays, such as an in-car device, mobile phone, tablet, personal digital assistant, wearable device, smart speaker, server, or server cluster.
[0101] The speech synthesis device may be software within an electronic device, such as speech synthesis software. In the following embodiments, the case where the execution entity is an electronic device will be described as an example.
[0102] As shown in Figure 4, the speech synthesis method may include the following steps 401 to 404.
[0103] In step 401, the style voice, tone voice, and input text to be processed are obtained.
[0104] In the embodiments of this disclosure, the speech synthesis method of this disclosure can be applied to at least one of the following scenes: intelligent dialogue, voice content creation, games and virtual humans, assistive technologies, mental health and education.
[0105] The intelligent dialogue scene includes, for example, AI assistants, virtual companion applications, customer service dialogues, enterprise digital employees, and voice assistants. The voice content creation scene includes, for example, novel audio systems, movie script dubbing, and short video monologues. The assistive technology scene includes, for example, assistive applications for the visually impaired and assistance for people with speech impairments. The mental health and education scene includes, for example, language learning and psychoeducation.
[0106] Style voice, tone voice, and input text to be processed are data that needs to be processed during speech synthesis in each of the above scenes.
[0107] The style of a styled voice can be, for example, bright and humorous yet logical, or warm and comforting like a friend. A styled voice may also be an audio recording obtained by reading a text using the above style.
[0108] The timbre of a timbre voice may, for example, be the timbre of a specific contrasting object. The timbre voice may also be an audio recording obtained when an object with a specific timbre reads a text aloud.
[0109] In step 402, a speech synthesis model is obtained, and the semantic coding network and semantic decoding network in the speech synthesis model are acquired by training them based on style sample speech, timbre sample speech, input sample text, and output sample speech, respectively.
[0110] In embodiments of this disclosure, a semantic coding network may be used to extract a semantic feature vector sequence based on a style sample speech, a timbre sample speech, and an input sample text to obtain a predicted semantic feature vector sequence, and a semantic decoding network may be used to perform semantic decoding on the predicted semantic feature vector sequence to obtain a predicted output speech.
[0111] In embodiments of this disclosure, the semantic coding network may include a style coding module, a timbre coding module, a text coding module, and an autoregressive semantic feature extraction module, the autoregressive semantic feature extraction module being connected to the style coding module, the timbre coding module, and the text coding module, respectively, and used to perform semantic feature extraction processing based on the style representation vector output by the style coding module, the timbre representation vector output by the timbre coding module, and the text representation vector sequence output by the text coding module, in order to obtain a semantic feature vector sequence.
[0112] In embodiments of the present disclosure, the autoregressive semantic feature extraction module includes sequentially connected sample context coding modules, a style-gated attention module, and a semantic output projector head module. The sample context coding module is used to determine a context representation vector corresponding to the current character for each current character to be predicted in a concatenated sample text, based on the representation vector of the current character and the representation vectors of each character preceding the current character. The style-gated attention module and the semantic output projector head module are used to predict the semantic feature vector of the current character based on the context representation vector, style representation vector, and timbre representation vector.
[0113] In embodiments of the present disclosure, in order to ensure consistency of style tendencies between the semantic feature vectors of each character and to naturally vary the style of each semantic feature vector in the predicted semantic feature vector sequence, a sequentially connected historical semantic memory coding module, an attention adjustment module, and a style adjustment gated unit are provided in the style-gated attention module, which is used to perform style tendency extraction processing based on an existing semantic feature vector sequence output by an autoregressive semantic feature extraction module, and the extracted style tendencies are used to predict the semantic feature vector of the current character.
[0114] The process by which the electronic device trains the training semantic coding network and semantic decoding network, respectively, based on style sample audio, timbre sample audio, input sample text, and output sample audio, can be found in the embodiments shown in Figures 1 to 3, and a detailed explanation is omitted here.
[0115] In step 403, the style voice, timbre voice, and input text are input to the semantic coding network in the speech synthesis model, and the semantic feature vector sequence output by the semantic coding network is obtained.
[0116] To adapt to an autoregressive semantic feature extraction module and extract a semantic feature vector sequence corresponding to an input sample text in combination with a timbre sample text, the electronic device can concatenate the timbre sample text before the input sample text, thereby obtaining the concatenated sample text. The style speech, timbre speech, and concatenated sample text are then input into a semantic coding network within a speech synthesis model to obtain the semantic feature vector sequence output by the semantic coding network.
[0117] In step 404, the semantic feature vector sequence is input to the semantic decoding network to obtain the output speech corresponding to the input text, which is output by the semantic decoding network.
[0118] The speech synthesis method of the embodiments of this disclosure acquires a style voice, a timbre voice, and an input text to be processed, acquires a speech synthesis model, and acquires a semantic coding network and a semantic decoding network in the speech synthesis model that are trained on a style sample voice, a timbre sample voice, an input sample text, and an output sample voice, respectively. The style voice, timbre voice, and input text are input to the semantic coding network in the speech synthesis model to acquire a semantic feature vector sequence output by the semantic coding network, and the semantic feature vector sequence is input to the semantic decoding network to acquire an output voice corresponding to the input text output by the semantic decoding network. The semantic coding network and the semantic decoding network are acquired based on training on a style sample voice, a timbre sample voice, an input sample text, and an output sample voice, respectively. As a result, the semantic coding network can extract more features from the style voice and timbre voice and fuse them into the semantic feature vector sequence, thereby decoding to obtain an output voice, and thus applying it to a new style and / or timbre to improve speech synthesis efficiency.
[0119] The following will provide an example to illustrate this point. Figure 5 is a schematic block diagram of a speech synthesis model. In Figure 5, the speech synthesis model includes a semantic coding network and a semantic decoding network. The semantic coding network includes a style coding module, a text coding module, a timbre coding module, and an autoregressive semantic feature extraction module. The semantic decoding network includes a semantic decoding module and a vocoder.
[0120] The style encoding module is used to encode style sample audio and obtain a style representation vector. The text encoding module is used to encode sample text and obtain a text representation vector. The timbre encoding module is used to encode timbre sample audio and obtain a timbre representation vector.
[0121] The autoregressive semantic feature extraction module is used to perform semantic feature extraction on style representation vectors, text representation vectors, and timbre representation vectors to obtain a predicted semantic feature vector sequence.
[0122] The semantic decoding module is used to decode the predicted semantic feature vector sequence and style representation vector to obtain the predicted acoustic feature sequence, and the vocoder is used to perform speech generation processing based on the predicted acoustic feature sequence to obtain the predicted output speech.
[0123] The loss function of the semantic coding network may be constructed based on the output semantic feature vector sequence obtained by extracting the output sample speech and the predicted semantic feature vector sequence output by the semantic coding network.
[0124] To realize the above embodiments, the present disclosure further provides a speech synthesis model training apparatus. As shown in Figure 6, Figure 6 is a schematic diagram relating to a fifth embodiment of the present disclosure. The speech synthesis model training apparatus 60 may include a first acquisition module 601, a second acquisition module 602, a training processing module 603, and a third acquisition module 604.
[0125] The first acquisition module 601 is used to acquire training data, the training samples in the training data include style sample voice, timbre sample voice, input sample text, and output sample voice; the second acquisition module 602 is used to acquire an initial speech synthesis model, the speech synthesis model includes a semantic coding network and a semantic decoding network; the training processing module 603 is used to perform training on the semantic coding network and the semantic decoding network based on the style sample voice, the timbre sample voice, the input sample text, and the output sample voice, respectively, and to acquire the trained semantic coding network and the trained semantic decoding network; and the third acquisition module 604 is used to acquire a trained speech synthesis model based on the trained semantic coding network and the trained semantic decoding network.
[0126] As one possible embodiment of the embodiments of the present disclosure, the first acquisition module 601 includes a first acquisition unit, a second acquisition unit, and a generation unit, wherein the first acquisition unit is used to acquire the style sample voice, the timbre sample voice, and the input sample text; the second acquisition unit is used to acquire a reference style voice and a reference timbre voice, wherein the reference style voice is the same style as the style sample voice, the reference timbre voice is the same timbre as the timbre sample voice, and the text corresponding to the reference timbre voice is the input sample text; and the generation unit is used to generate the output sample voice based on the reference style voice and the reference timbre voice.
[0127] In one possible embodiment of the embodiments of the present disclosure, the first acquisition unit specifically acquires a style description sample text, the timbre sample audio, and the input sample text; acquires a style audio library which includes at least one style description text and a style audio corresponding to the style description text; queries the style audio library based on the style description sample text; acquires a first style description sample in the style audio library that matches the style description sample text; and uses the style audio corresponding to the first style description sample to determine the style sample audio.
[0128] As one possible implementation of an embodiment of the present disclosure, an autoregressive semantic feature extraction module is installed within the semantic coding network, and the training processing module 603 includes a decision unit, a concatenation unit and a training processing unit, wherein the decision unit is used to determine a timbre sample text corresponding to the timbre sample speech; the concatenation unit is used to concatenate the timbre sample text and the input sample text to obtain a concatenated sample text; and the training processing unit is used to perform training on the semantic coding network based on the style sample speech, the timbre sample speech, the concatenated sample text and the output sample speech.
[0129] In one possible implementation of the embodiments of this disclosure, the training processing unit specifically performs semantic feature extraction on the output sample audio to obtain an output semantic feature vector sequence, inputs the style sample audio, the timbre sample audio, and the concatenated sample text into the semantic coding network to obtain an intermediate semantic feature vector sequence corresponding to the concatenated sample text output by the semantic coding network, extracts a predicted semantic feature vector sequence corresponding to the input sample text from the intermediate semantic feature vector sequence, and performs parameter tuning on the semantic coding network based on the predicted semantic feature vector sequence and the output semantic feature vector sequence to obtain a trained semantic coding network.
[0130] As one possible implementation of an embodiment of the present disclosure, the semantic coding network includes a style coding module, a timbre coding module, a text coding module, and the autoregressive semantic feature extraction module, wherein the training processing unit is used to obtain a style representation vector by performing style coding on a style sample audio via the style coding module, to obtain a timbre representation vector by performing timbre coding on a timbre sample audio via the timbre coding module, to obtain a text representation vector sequence by performing text coding on the concatenated sample text via the text coding module, and to obtain the intermediate semantic feature vector sequence by performing semantic feature extraction on the input style representation vector, timbre representation vector, and text representation vector sequence via the autoregressive semantic feature extraction module.
[0131] As one possible implementation of the embodiments of the present disclosure, the method for processing the concatenated sample text by the text encoding module may include: performing character encoding on each character in the concatenated sample text to obtain a character vector sequence; performing character position encoding and / or paragraph position encoding on each character in the concatenated sample text to obtain a position vector sequence; and performing concatenation on the character vectors in the character vector sequence and the position vectors in the position vector sequence according to the characters to obtain a text display vector sequence output by the text encoding module.
[0132] In one possible implementation of an embodiment of the present disclosure, the autoregressive semantic feature extraction module includes sequentially connected sample context coding modules, a style-gated attention module, and a semantic output projector head module, wherein the training processing unit specifically obtains, for each of the current characters to be predicted in the concatenated sample text, the representation vector of the current character and the representation vectors of each character preceding the current character; performs context feature extraction on the representation vector of the current character and the representation vectors of each character via the sample context coding module to obtain a context representation vector corresponding to the current character; and also uses the style-gated attention module and the semantic output projector head module to perform semantic feature extraction on the context representation vector, the style representation vector, and the timbre representation vector to obtain a semantic feature vector of the current character.
[0133] In one possible implementation of an embodiment of the present disclosure, the style-gated attention module is equipped with sequentially connected historical semantic memory coding modules, attention adjustment modules, and style adjustment gated units to perform style tendency extraction processing based on an existing semantic feature vector sequence output by the autoregressive semantic feature extraction module, and the extracted style tendency is used to predict the semantic feature vector of the current character.
[0134] In one possible embodiment of the embodiments of the present disclosure, the training processing unit is also used to determine a numerical value for a sub-loss function of at least one of the following dimensions—style dimension, timbre dimension, semantic dimension, and style consistency dimension—based on the predicted semantic feature vector sequence and the output semantic feature vector sequence, and to perform parameter tuning on the semantic coding network based on the numerical value of the sub-loss function of at least one dimension to obtain a trained semantic coding network.
[0135] As one possible implementation of the embodiments of this disclosure, the numerical value of the sub-loss function for the style dimension is determined based on the style representation vector output by the style coding module and the predicted style representation vector, wherein the predicted style representation vector is obtained by performing style extraction on the predicted semantic feature vector sequence; the numerical value of the sub-loss function for the timbre dimension is determined based on the timbre representation vector corresponding to the timbre sample speech and the predicted timbre representation vector, wherein the predicted timbre representation vector is obtained by performing timbre extraction on the predicted semantic feature vector sequence; the numerical value of the sub-loss function for the semantic dimension is determined based on the predicted semantic feature vector sequence and the output semantic feature vector sequence; and the numerical value of the sub-loss function for the style consistency dimension is determined based on the predicted semantic feature vectors corresponding to adjacent characters in the predicted semantic feature vector sequence.
[0136] In one possible embodiment of the embodiments of the present disclosure, the training processing module 603 further comprises a third acquisition unit and a fourth acquisition unit, the third acquisition unit being used to input the style sample speech, the timbre sample speech and the input sample text into a trained semantic coding network to obtain a predicted semantic feature vector sequence output by the trained semantic coding network; the fourth acquisition unit being used to input the predicted semantic feature vector sequence into a semantic decoding network to obtain a predicted output speech output by the semantic decoding network; and the training processing unit being used to perform parameter tuning on the semantic decoding network based on the predicted output speech and the output sample speech to obtain a trained semantic decoding network.
[0137] In one possible embodiment of the embodiments of the present disclosure, the fourth acquisition unit is specifically used to acquire a style representation vector corresponding to the style sample speech, input the style representation vector and the predicted semantic feature vector sequence into the semantic decoding network, and acquire the output predicted speech output by the semantic decoding network.
[0138] In one possible embodiment of the embodiments of the present disclosure, the training processing unit is used to determine a numerical value for a sub-loss function of at least one of the following dimensions: speech dimension, speech rhythm dimension, and speech consistency dimension, based on the output predicted speech and the output sample speech, and to perform a parameter tuning process on the semantic decoding network based on the numerical value of the sub-loss function of at least one dimension, in order to obtain a trained semantic decoding network.
[0139] As one possible implementation of the embodiments of this disclosure, the numerical value of the sub-loss function for the speech dimension is determined based on the speech similarity between the predicted output speech and the output sample speech; the numerical value of the sub-loss function for the speech rhythm dimension is determined based on predicted rhythm data and sample rhythm data, the predicted rhythm data is obtained by rhythm extraction from the predicted output speech; the sample rhythm data is obtained by rhythm extraction from the output sample speech; and the numerical value of the sub-loss function for the speech consistency dimension is determined based on the differences between adjacent speech frames in the predicted output speech.
[0140] The speech synthesis model training apparatus of the embodiment of this disclosure acquires training data, the training samples in the training data include style sample speech, timbre sample speech, input sample text, and output sample speech, an initial speech synthesis model is acquired, the speech synthesis model includes a semantic coding network and a semantic decoding network, training is performed on the semantic coding network and the semantic decoding network based on the style sample speech, timbre sample speech, input sample text, and output sample speech, a trained semantic coding network and a trained semantic decoding network are acquired, and a trained speech synthesis model is acquired based on the trained semantic coding network and the trained semantic decoding network, where the style sample speech and timbre sample speech are expressive and can be combined with semantic feature extraction of the semantic coding network to extract more features and fuse them into the output speech obtained by synthesis, thereby avoiding training processing for each style and / or timbre, reducing the training cost of the speech synthesis model, and improving the generality of the trained speech synthesis model in new styles and / or timbres.
[0141] To realize the above embodiments, the present disclosure further provides a speech synthesis device. As shown in Figure 7, Figure 7 is a schematic diagram relating to a sixth embodiment of the present disclosure. The speech synthesis device 70 may include a first acquisition module 701, a second acquisition module 702, a third acquisition module 703, and a fourth acquisition module 704.
[0142] The first acquisition module 701 is used to acquire style speech, timbre speech, and input text to be processed; the second acquisition module 702 is used to acquire a speech synthesis model, the semantic coding network and semantic decoding network in the speech synthesis model being acquired by training them based on style sample speech, timbre sample speech, input sample text, and output sample speech, respectively; the third acquisition module 703 is used to input the style speech, timbre speech, and input text into the semantic coding network in the speech synthesis model and acquire the semantic feature vector sequence output by the semantic coding network; and the fourth acquisition module 704 is used to input the semantic feature vector sequence into the semantic decoding network and acquire the output speech corresponding to the input text output by the semantic decoding network.
[0143] The speech synthesis apparatus of the embodiment of this disclosure acquires style speech, timbre speech and input text to be processed, acquires a speech synthesis model, and the semantic coding network and semantic decoding network in the speech synthesis model are acquired by training them based on style sample speech, timbre sample speech, input sample text and output sample speech, respectively. The style speech, timbre speech and input text are input to the semantic coding network in the speech synthesis model to acquire a semantic feature vector sequence output by the semantic coding network, and the semantic feature vector sequence is input to the semantic decoding network to acquire output speech corresponding to the input text output by the semantic decoding network. Here, the semantic coding network and semantic decoding network are obtained based on training with style sample speech, timbre sample speech, input sample text and output sample speech, respectively. As a result, the semantic coding network can extract more features from the style speech and timbre speech and fuse them into the semantic feature vector sequence, thereby enabling decoding to obtain output speech, which can be applied to new styles and / or timbres, improving speech synthesis efficiency.
[0144] In the technical solutions disclosed herein, all processing of relevant user personal information, including collection, storage, use, processing, transmission, provision, and disclosure, is carried out with the user's consent, complies with the provisions of relevant laws and regulations, and does not violate public order and morals.
[0145] According to embodiments of the present disclosure, the present disclosure further provides electronic devices, readable storage media, and computer programs.
[0146] Figure 8 is an exemplary block diagram of an exemplary electronic device 800 for carrying out an embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, mobile phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the description herein and / or the implementation of the present disclosure as requested.
[0147] As shown in Figure 8, the electronic device 800 includes a computing unit 801 that performs various appropriate operations and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. The RAM 803 may also store various programs and data necessary for the operation of the electronic device 800. The computing unit 801, ROM 802, and RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0148] Multiple components of the electronic device 800 are connected to the I / O interface 805, which includes input units 806 such as a keyboard and mouse, output units 807 such as various types of displays and speakers, storage units 808 such as magnetic disks and optical disks, and communication units 809 such as a network card, modem, and wireless communication transceiver. The communication units 809 enable the device 800 to exchange information / data with other devices via computer networks such as the Internet and / or various types of telegraph networks.
[0149] The computing unit 801 may be a variety of general-purpose and / or dedicated processing components having processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that execute machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs each of the methods and processes described above, for example, a method for training a speech synthesis model or a method for synthesizing speech. For example, in some embodiments, a method for training a speech synthesis model or a method for synthesizing speech can be implemented as a computer software program tangibly contained in a machine-readable medium such as a memory unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed into the electronic device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the method for training a speech synthesis model or a method for synthesizing speech described above can be performed. Alternatively, in other embodiments, the computing unit 801 may be configured in any other suitable manner (e.g., via firmware) to perform a method for training a speech synthesis model or a method for speech synthesis.
[0150] Various embodiments of the systems and technologies described above in this specification can be implemented as digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include being implemented by one or more computer programs, which can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be an application-specific or general-purpose programmable processor, which can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, at least one input device, and at least one output device.
[0151] Program code for performing the methods of this disclosure can be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, a dedicated computer, or other programmable data processing device so that, when executed by the processor or controller, the functions / operations defined in the flowcharts and / or block diagrams are performed. The program code may run entirely on a machine, partially on a machine, as a standalone software package, partially on a machine, partially on a remote machine, or entirely on a remote machine or server.
[0152] In the context of this disclosure, a machine-readable medium may be a tangible medium that contains or can store a program used by or in combination with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any appropriate combination of the above. More specific examples of machine-readable storage media include one or more line-based electrical connections, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any appropriate combination of the above.
[0153] To provide user interaction, the systems and technologies described herein can be implemented on a computer, which includes a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) and a keyboard and pointing device (e.g., a mouse or trackball), and the user can provide input to the computer using the keyboard and pointing device. Other types of devices may also be used to provide user interaction, for example, the feedback provided to the user may be any form of sensing feedback (e.g., visual feedback, auditory feedback, or haptic feedback), and input from the user may be received in any form (including acoustic input, voice input, and haptic input).
[0154] The systems and technologies described herein can be run on computing systems including backend components (e.g., data servers), computing systems including middleware components (e.g., application servers), computing systems including frontend components (e.g., user computers having a graphical user interface or web browser, through which users can interact with embodiments of the systems and technologies described herein), or any combination of such backend components, middleware components, and frontend components. The components of the system can be interconnected by digital data communication (e.g., communication networks) in any form or medium. Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0155] A computer system can include clients and servers. Clients and servers are generally geographically separated and typically interact via a communication network. The client-server relationship is generated by computer programs running on corresponding computers that have a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server incorporating blockchain technology.
[0156] Furthermore, the steps can be rearranged, added, or deleted using the various forms of processes described above. For example, each step described herein may be performed in parallel, sequentially, or in a different order, as long as the desired results of the proposed technology disclosed herein can be achieved.
[0157] The specific embodiments described above do not limit the scope of protection of this disclosure. Those skilled in the art will understand that various modifications, combinations, subcombinations, and substitutions can be made depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for training speech synthesis models, A step of acquiring training data, wherein the training samples in the training data include style sample audio, timbre sample audio, input sample text, and output sample audio. A step of obtaining an initial speech synthesis model, wherein the speech synthesis model includes a semantic coding network and a semantic decoding network, The steps include: performing training on the semantic coding network and the semantic decoding network based on the style sample audio, the timbre sample audio, the input sample text, and the output sample audio, respectively, and obtaining the trained semantic coding network and the trained semantic decoding network; The process includes the step of obtaining a trained speech synthesis model based on the trained semantic coding network and the trained semantic decoding network. Training methods for speech synthesis models.
2. The step of obtaining training data is: The steps include obtaining the aforementioned style sample audio, the aforementioned tone sample audio, and the aforementioned input sample text, A step of obtaining a reference style voice and a reference timbre voice, wherein the reference style voice is the same style as the style sample voice, the reference timbre voice is the same timbre as the timbre sample voice, and the text corresponding to the reference timbre voice is the input sample text, The process includes the step of generating the output sample audio based on the reference style audio and the reference timbre audio, The method according to claim 1.
3. The steps of obtaining the aforementioned style sample audio, the aforementioned tone sample audio, and the aforementioned input sample text are as follows: The steps include obtaining a style description sample text, the timbre sample audio, and the input sample text, A step of obtaining a style voice library, wherein the style voice library includes at least one style description text and a style voice corresponding to the style description text. The steps include querying the style voice library based on the style description sample text and obtaining a first style description sample that matches the style description sample text within the style voice library, The step includes determining the style audio corresponding to the first style description sample as the style sample audio, The method according to claim 2.
4. An autoregressive semantic feature extraction module is installed within the aforementioned semantic coding network. The step of performing a training process on the semantic coding network based on the style sample audio, the timbre sample audio, the input sample text, and the output sample audio is: The steps include determining a tone sample text corresponding to the aforementioned tone sample audio, The steps include: concatenating the aforementioned tone sample text and the aforementioned input sample text to obtain the concatenated sample text; The process includes training the semantic coding network based on the style sample audio, the timbre sample audio, the concatenated sample text, and the output sample audio. The method according to claim 1.
5. The step of training the semantic coding network based on the style sample audio, the timbre sample audio, the concatenated sample text, and the output sample audio is: The steps include: performing semantic feature extraction on the output sample audio to obtain an output semantic feature vector sequence; The steps include inputting the style sample audio, the timbre sample audio, and the concatenated sample text into the semantic coding network to obtain an intermediate semantic feature vector sequence corresponding to the concatenated sample text output by the semantic coding network, The steps include: extracting a predictive semantic feature vector sequence corresponding to the input sample text from the intermediate semantic feature vector sequence; The process includes the step of performing parameter tuning on the semantic coding network based on the predicted semantic feature vector sequence and the output semantic feature vector sequence to obtain the trained semantic coding network. The method according to claim 4.
6. The semantic coding network includes a style coding module, a timbre coding module, a text coding module, and the autoregressive semantic feature extraction module. The step of inputting the style sample audio, the timbre sample audio, and the concatenated sample text into the semantic coding network to obtain an intermediate semantic feature vector sequence corresponding to the concatenated sample text output by the semantic coding network is: The steps include: performing style encoding processing on a style sample audio via the aforementioned style encoding module to obtain a style representation vector; The steps include: performing timbre encoding processing on the timbre sample audio via the timbre encoding module to obtain a timbre representation vector; The steps include: performing text encoding on the concatenated sample text via the text encoding module to obtain a text representation vector sequence; The process includes the step of performing semantic feature extraction on the input style representation vector, timbre representation vector, and text representation vector sequence via the autoregressive semantic feature extraction module to obtain the intermediate semantic feature vector sequence. The method according to claim 5.
7. The method for processing the concatenated sample text by the text encoding module is as follows: The steps include: performing character encoding on each character in the concatenated sample text to obtain a character vector sequence; The steps include: obtaining a position vector sequence by performing character position encoding and / or paragraph position encoding on each character in the concatenated sample text; The process includes the steps of concatenating the character vectors in the character vector sequence and the position vectors in the position vector sequence according to the characters, and obtaining the text display vector sequence output by the text encoding module. The method according to claim 6.
8. The autoregressive semantic feature extraction module includes sequentially connected sample context coding modules, style gated attention modules, and semantic output projector head modules. The step of performing semantic feature extraction processing on the input style representation vector, timbre representation vector, and text representation vector sequence via the autoregressive semantic feature extraction module to obtain the intermediate semantic feature vector sequence is as follows: For each of the current characters to be predicted in the concatenated sample text, the steps include obtaining the display vector of the current character and the display vectors of each character preceding the current character. The steps include: performing context feature extraction processing on the current character's representation vector and each of the character's representation vectors via the sample context coding module to obtain a context representation vector corresponding to the current character; The steps include: performing semantic feature extraction processing on the context display vector, the style display vector, and the timbre display vector via the style gated attention module and the semantic output projector head module to obtain the semantic feature vector of the current character; The method according to claim 6.
9. In the aforementioned style-gated attention module, a sequentially connected history semantic memory encoding module, attention adjustment module, and style adjustment gated unit are installed to perform style tendency extraction processing based on an existing semantic feature vector sequence output by the autoregressive semantic feature extraction module, and the extracted style tendency is used to predict the semantic feature vector of the current character. The method according to claim 8.
10. The step of performing parameter tuning on the semantic coding network based on the predicted semantic feature vector sequence and the output semantic feature vector sequence to obtain the trained semantic coding network is: A step of determining the numerical value of a sub-loss function for at least one of the following dimensions: style dimension, timbre dimension, semantic dimension, and style consistency dimension, based on the predicted semantic feature vector sequence and the output semantic feature vector sequence. The process includes the step of performing a parameter tuning process on the semantic coding network based on the numerical values of the sub-loss function of at least one dimension to obtain the trained semantic coding network. The method according to claim 5.
11. The numerical value of the sub-loss function for the style dimension is determined based on the style representation vector output by the style coding module and the predicted style representation vector, the predicted style representation vector is obtained by performing style extraction on the predicted semantic feature vector sequence. The numerical value of the sub-loss function for the timbre dimension is determined based on the timbre representation vector and the predicted timbre representation vector corresponding to the timbre sample audio, and the predicted timbre representation vector is obtained by performing timbre extraction on the predicted semantic feature vector sequence. The numerical value of the semantic dimension sub-loss function is determined based on the predicted semantic feature vector sequence and the output semantic feature vector sequence. The numerical value of the sub-loss function for the style consistency dimension is determined based on the predicted semantic feature vectors corresponding to adjacent characters in the predicted semantic feature vector sequence. The method according to claim 10.
12. The step of performing a training process on the semantic decoding network based on the style sample audio, the timbre sample audio, the input sample text, and the output sample audio is: The steps include inputting the style sample audio, the timbre sample audio, and the input sample text into a trained semantic coding network to obtain a predicted semantic feature vector sequence output by the trained semantic coding network, The steps include inputting the predicted semantic feature vector sequence into the semantic decoding network to obtain the predicted output speech output by the semantic decoding network, The process includes the step of performing parameter adjustment on the semantic decoding network based on the output prediction speech and the output sample speech to obtain the trained semantic decoding network. The method according to claim 1.
13. The step of inputting the predicted semantic feature vector sequence into the semantic decoding network and obtaining the predicted output speech output by the semantic decoding network is: The steps include obtaining a style display vector corresponding to the aforementioned style sample audio, The step includes inputting the style representation vector and the predicted semantic feature vector sequence into the semantic decoding network to obtain the predicted output speech output by the semantic decoding network, The method according to claim 12.
14. The step of performing parameter adjustment processing on the semantic decoding network based on the output prediction speech and the output sample speech to obtain the trained semantic decoding network is as follows: A step of determining the numerical value of a sub-loss function for at least one of the following dimensions: speech dimension, speech rhythm dimension, and speech consistency dimension, based on the output prediction speech and the output sample speech. The process includes the step of performing a parameter tuning process on the semantic decoding network based on the numerical values of the sub-loss function of at least one dimension to obtain the trained semantic decoding network. The method according to claim 12.
15. The numerical value of the sub-loss function for the speech dimension is determined based on the speech similarity between the predicted output speech and the output sample speech. The numerical value of the sub-loss function for the aforementioned speech rhythm dimension is determined based on predicted rhythm data and sample rhythm data, wherein the predicted rhythm data is obtained by rhythm extraction from the output predicted speech, and the sample rhythm data is obtained by rhythm extraction from the output sample speech. The numerical value of the sub-loss function for the speech consistency dimension is determined based on the differences between adjacent speech frames in the output prediction speech. The method according to claim 14.
16. A speech synthesis method, Steps include obtaining style voice, tone voice, and input text to be processed, A step of acquiring a speech synthesis model, wherein the semantic coding network and semantic decoding network in the speech synthesis model are acquired by training them based on style sample speech, timbre sample speech, input sample text, and output sample speech, respectively. The steps include inputting the aforementioned style voice, the aforementioned timbre voice, and the aforementioned input text into a semantic coding network in a speech synthesis model, and obtaining a semantic feature vector sequence output by the semantic coding network, The process includes the step of inputting the semantic feature vector sequence into the semantic decoding network and obtaining the output speech corresponding to the input text output by the semantic decoding network. Speech synthesis method.
17. A training device for speech synthesis models, A first acquisition module for acquiring training data, wherein the training samples in the training data include style sample audio, timbre sample audio, input sample text, and output sample audio. A second acquisition module for acquiring an initial speech synthesis model, wherein the speech synthesis model includes a semantic coding network and a semantic decoding network. A training processing module for performing training on the semantic coding network and the semantic decoding network based on the style sample audio, the timbre sample audio, the input sample text, and the output sample audio, respectively, and obtaining the trained semantic coding network and the trained semantic decoding network, A third acquisition module for acquiring a trained speech synthesis model based on the trained semantic coding network and the trained semantic decoding network, A training device for speech synthesis models.
18. A speech synthesis device, A first acquisition module for obtaining style voice, tone voice, and input text to be processed, A second acquisition module for acquiring a speech synthesis model, wherein the semantic coding network and semantic decoding network in the speech synthesis model are acquired by training them based on style sample speech, timbre sample speech, input sample text, and output sample speech, respectively. A third acquisition module for inputting the aforementioned style voice, the aforementioned timbre voice, and the aforementioned input text into a semantic coding network in a speech synthesis model, and for obtaining the semantic feature vector sequence output by the semantic coding network, A fourth acquisition module for inputting the semantic feature vector sequence into the semantic decoding network and acquiring output audio corresponding to the input text output by the semantic decoding network, is included. Speech synthesis device.
19. It is an electronic device, At least one process, The memory includes at least one processor and is communicably connected to it. The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the method according to any one of claims 1 to 15 or the method according to claim 16. Electronic devices.
20. A non-temporary computer-readable storage medium in which computer instructions are stored, The computer instruction is used to cause the computer to execute the method according to any one of claims 1 to 15 or the method according to claim 16. A non-temporary, computer-readable storage medium.
21. It is a computer program, When the computer program is executed by the processor, it performs the method according to any one of claims 1 to 15 or the method according to claim 16. Computer program.