Speech synthesis method and apparatus, electronic device, and storage medium
By generating and fusing the acoustic feature sequence of the target text into the speech synthesis model, the problems of pronunciation accuracy and clarity of the existing model are solved, and higher quality speech synthesis is achieved.
Patent Information
- Application Number
- PCT/CN2025/087108
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-08
- Filing Date
- 2025-04-03
- Publication Date
- 2025-10-16
AI Technical Summary
Existing speech synthesis models suffer from pronunciation errors and poor clarity when generating speech, especially when dealing with phoneme performance in different contexts and are unable to accurately capture acoustic features.
By generating an acoustic feature sequence corresponding to the target text and fusing it with the phoneme sequence extracted by the speech synthesis model to form a fused sequence, the acoustic feature information of the target text is supplemented, and the target synthesized speech is generated using the reference speech.
The pronunciation accuracy and clarity of the speech synthesis model are improved, pronunciation errors are reduced, and the quality and naturalness of speech synthesis are enhanced.
Smart Images

Figure CN2025087108_16102025_PF_FP_ABST
Abstract
Description
Speech synthesis method and device, electronic equipment and storage medium Cross-reference to related applications This application claims priority to the Chinese patent application No. 2024104169366, filed on April 8, 2024, and entitled "Speech synthesis method and device, electronic equipment and storage medium", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0001] The present application relates to, but is not limited to, the technical field of speech synthesis, and in particular to a speech synthesis method, device, electronic equipment and storage medium. BACKGROUND
[0002] Speech synthesis (TTS for short) is a technology that can convert any input text into corresponding speech. The synthesized speech can simulate the characteristics of human voice, such as tone, pitch and speed, allowing computers to interact with humans more intelligently. In daily life, speech synthesis technology can be widely used in voice assistants, automatic voice response machines, self-service terminals and other devices. Through this technology, devices can answer users' questions with personified voices. SUMMARY
[0003] The following is a summary of the subject matter described in detail herein. This summary is not intended to limit the scope of the claims.
[0004] The present application provides a speech synthesis method, device, electronic equipment and storage medium.
[0005] According to a first aspect of an embodiment of the present application, a speech synthesis method is provided, the method comprising: generating an acoustic feature sequence corresponding to a target text to be synthesized, the acoustic feature sequence being used to represent acoustic features of the target text; and fusing the acoustic feature sequence and a phoneme sequence of the target text extracted by a speech synthesis model into a fused sequence.
[0006] Obtaining a target synthesized speech generated by the speech synthesis model according to a reference speech and the fused sequence.
[0007] According to a second aspect of the embodiments of the present application, a speech synthesis device is provided, comprising an acoustic feature sequence generating unit, a fusion sequence generating unit and a target synthesized speech generating unit. The acoustic feature sequence generating unit is configured to generate an acoustic feature sequence corresponding to a target text to be synthesized, the acoustic feature sequence being configured to represent acoustic features of the target text. The fusion sequence generating unit is configured to fuse the acoustic feature sequence with a phoneme sequence of the target text extracted by a speech synthesis model into a fusion sequence. The target synthesized speech generating unit is configured to obtain a target synthesized speech generated by the speech synthesis model based on a reference speech and the fusion sequence.
[0008] According to a third aspect of the embodiments of the present application, an electronic device is provided, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method according to the first aspect when executing the program.
[0009] According to a fourth aspect of the embodiments of the present application, a computer readable storage medium is provided, having a computer program stored thereon, wherein the program is executable by a processor to implement the steps of the method according to the first aspect.
[0010] In the embodiments of the present application, an acoustic feature sequence corresponding to a target text to be synthesized is generated to represent acoustic features of the target text, and the acoustic feature sequence is fused with a phoneme sequence of the target text extracted by a speech synthesis model into a fusion sequence, so that the fusion sequence contains acoustic features representing the target text. Then, the speech synthesis model can generate a target synthesized speech based on the fusion sequence and a reference speech.
[0011] It can be seen that, in the process of generating a target synthesized speech based on a target text, the present application supplements the target text with corresponding acoustic feature information. Specifically, instead of directly inputting the target text and a reference speech into a speech synthesis model, an acoustic feature sequence representing acoustic features of the target text is first generated, and then the sequence is fused with a phoneme sequence of the target text extracted by the speech synthesis model into a fusion sequence, so that the speech synthesis model generates a corresponding target synthesized speech based on the fusion sequence and the reference speech. In this way, the speech synthesis model can more accurately convert the target text into a target synthesized speech, so as to reduce pronunciation errors of the target synthesized speech.
[0012] It should be understood that the general description above and the following detailed description are only exemplary and explanatory, and cannot limit the present application. Other aspects can be understood after reading and understanding the drawings and detailed description. BRIEF DESCRIPTION OF DRAWINGS
[0013] The accompanying drawings, which are incorporated herein and constitute part of the specification, illustrate embodiments consistent with the application and, together with the description, further serve to explain the principles of the application.
[0014] FIG. 1 is a structural schematic diagram of a speech synthesis model according to an example embodiment of the present application.
[0015] FIG. 2 is a flow schematic diagram of a speech synthesis method according to an example embodiment of the present application.
[0016] FIG. 3 is a model framework diagram of an acoustic feature predictor according to an example embodiment of the present application.
[0017] FIG. 4 is a flow schematic diagram of generating a fusion sequence according to an example embodiment of the present application.
[0018] FIG. 5 is a structural schematic diagram of a speech synthesis model in an inference stage according to an example embodiment of the present application.
[0019] FIG. 6 is a flow schematic diagram of a training method for an acoustic feature predictor according to an example embodiment of the present application.
[0020] FIG. 7 is a structural schematic diagram of a speech synthesis model in a training stage according to an example embodiment of the present application.
[0021] FIG. 8 is a schematic diagram of another training method for an acoustic feature predictor according to an example embodiment of the present application.
[0022] FIG. 9 is a flow schematic diagram of obtaining a training data set according to an example embodiment of the present application.
[0023] FIG. 10 is a flow schematic diagram of another method of obtaining a training data set according to an example embodiment of the present application.
[0024] FIG. 11 is a block diagram of an electronic device according to an example embodiment of the present application.
[0025] FIG. 12 is a block diagram of a speech synthesis apparatus according to an example embodiment of the present application. DETAILED DESCRIPTION
[0026] The example embodiments will be described in detail herein with reference to the accompanying drawings. In the following description, same numbers refer to same elements in all figures unless otherwise described. The following example embodiments described in the example embodiments do not represent all implementations consistent with the present application. Instead, they only describe examples of devices and methods consistent with some aspects of the present application, as detailed in the appended claims.
[0027] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting. As used in this application and the appended claims, the singular forms "a," "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0028] It should be understood that, although the terms first, second, third, etc. can be used herein to describe various information, the information should not be limited to these terms. These terms are only used to differentiate one piece of information from another piece of information. For example, a first information can also be called a second information without departing from the scope of the application, and similarly, a second information can also be called a first information. Depending on the context, the word "if' as used herein can be interpreted as "when" or "upon determination" or "in response to determining".
[0029] Next, the embodiments of the present application are described in detail.
[0030] The speech synthesis model is a specific algorithm or framework for implementing speech synthesis technology. For example, the VALL-E model can output synthesized speech expressing the text information by inputting the text information and the reference speech into the VALL-E model, and the synthesized speech has similar voice characteristics to the reference speech.
[0031] However, the current speech synthesis model still has problems of pronunciation errors, such as poor clarity of synthesized speech, poor distinction of homophonic words, etc.
[0032] As shown in FIG. 1, FIG. 1 shows a model architecture of a speech synthesis model 1 according to some embodiments of the present application. The speech synthesis model 1 can be a VALL-E model. Of course, it can also be other token-based speech synthesis models, such as Coqui TTS model or BASE TTS model. The model framework of the speech synthesis model 1 can include but is not limited to a phoneme conversion module 102, an audio codec encoder 105, an acoustic model 106, and an audio codec decoder 107. If the speech synthesis model is a VALL-E model, the acoustic model 106 is specifically a neural codec language model. Specifically, the phoneme conversion module 102 is used to convert the target text 101 to be synthesized into a phoneme sequence 103 and input into the acoustic model 106; the audio codec encoder 105 is used to convert the reference speech 104 into a first discrete acoustic sequence and input into the acoustic model 106; the second discrete acoustic sequence output by the acoustic model 106 is input into the audio codec decoder 107, and the target synthesized speech 108 is generated through the audio codec decoder 107.
[0033] It should be noted that in the speech synthesis model 1, specifically the processing steps in the dashed box 10: the phoneme conversion module 102 directly converts the target text 101 to be synthesized into a phoneme sequence 103 and inputs it into the acoustic model 106. As can be seen, since the above-mentioned phoneme sequence 103 input into the acoustic model 106 is directly converted from the target text 101 to be synthesized, it only contains the basic unit of the speech, lacking more fine-grained acoustic feature information about the target text 101 to be synthesized. However, in actual application, given that different phonemes have different acoustic manifestations in different contexts, the speech synthesis model 1 may not be able to capture this complex relationship when establishing the mapping of the acoustic features of the target text 101 to be synthesized to the target synthesized speech, resulting in inaccurate pronunciation of the target synthesized speech 108 finally generated.
[0034] As shown in FIG. 2, the present application provides a speech synthesis method. Specifically, it includes the following steps 201 to 203.
[0035] Step 201: generating an acoustic feature sequence corresponding to the target text to be synthesized.
[0036] In an embodiment, the target text to be synthesized is the text content that needs to generate the target synthesized speech. The target text to be synthesized can be any form of text, which can be a sentence, a paragraph, or even an article. For example, the target text to be synthesized input to the speech synthesis model can be a short sentence—“Tomorrow's weather forecast shows that it will rain in Hangzhou.”; or a longer article—“This is a research paper on artificial intelligence, which introduces the principles and applications of machine learning and deep learning in detail....”
[0037] In an embodiment, an acoustic feature sequence corresponding to the target text to be synthesized is generated, which is used to represent the acoustic features of the target text. The form of the “feature sequence” in this embodiment can be a feature vector or a feature matrix, etc. Specifically, the target text to be synthesized is converted into a corresponding phoneme sequence using a trained acoustic feature predictor, and the phoneme sequence is further converted into a corresponding acoustic feature sequence, which is used to represent the acoustic features of the target text. It should be noted that the acoustic feature predictor can refer to a deep learning model that has the ability to generate an acoustic feature sequence corresponding to the target text to be synthesized, and is not the name of a specific model. For example, the acoustic feature predictor can be a commonly used Transformer model, or a recurrent neural network (RNN) model. By training the Transformer model or the RNN model, it has the ability to generate an acoustic feature sequence corresponding to the target text to be synthesized. Of course, those skilled in the art can not use the model architecture of the existing deep learning model, but customize the model architecture of the acoustic feature predictor to improve the adaptability of the generated acoustic feature sequence.
[0038] In another embodiment, the model architecture of the acoustic feature predictor is customized in this embodiment:
[0039] As shown in FIG. 3, the acoustic feature predictor 3 can be composed of two parts, a phoneme conversion module 30 and an acoustic feature conversion module 31. The phoneme conversion module 30 can adopt the same model architecture as the phoneme conversion module 102 in the speech synthesis model 1 of FIG. 1, for converting the target text 101 to be synthesized into a phoneme sequence, and inputting the phoneme sequence to the acoustic feature conversion module 31, so as to generate an acoustic feature sequence 307 corresponding to the target text 101 to be synthesized by the acoustic feature conversion module 31. The acoustic feature conversion module 31 includes five conversion sub-modules connected in sequence, and the output of the previous sub-module is the input of the next sub-module. Among them, Conv1D represents one-dimensional convolution; RMSNorm stands for Root Mean Square Normalization; DP stands for Dropout; Linear represents linear transformation.
[0040] Step 202: Fuse the acoustic feature sequence and the phoneme sequence of the target text extracted by the speech synthesis model into a fusion sequence.
[0041] In an embodiment, as shown in FIG. 4, the target text 101 to be synthesized is input into the phoneme conversion module 102 and the acoustic feature predictor 3 respectively, and the phoneme sequence 103 output by the phoneme conversion module 102 and the acoustic feature sequence 307 output by the acoustic feature predictor 3 are processed into a fusion sequence 402 through fusion processing 401.
[0042] In an embodiment, the fusion processing 401 can be to splice the acoustic feature sequence 307 and the phoneme sequence 103 together to generate the fusion sequence. For example, if the acoustic feature sequence 307 is [1, 2, 1] and the phoneme sequence 103 is [1, 1, 1], the fusion sequence after splicing them together can be [1, 2, 1, 1, 1, 1] or [1, 1, 1, 1, 2, 1]. The application does not limit the splicing manner of the acoustic feature sequence and the phoneme sequence.
[0043] In another embodiment, the fusion processing 401 can also be to add the feature values at the corresponding positions in the acoustic feature sequence 307 and the phoneme sequence 103 to generate the fusion sequence. For example, if the acoustic feature sequence is [1, 2, 1] and the phoneme sequence is [1, 1, 1], the fusion sequence can be [2, 3, 2]. Of course, the feature values after addition can be further averaged, and then the fusion sequence becomes [2 / 2, 3 / 2, 2 / 2].
[0044] Through the above fusion processing, the features of the acoustic feature sequence and the corresponding phoneme sequence are combined together, so that the fusion sequence contains the features of the acoustic feature sequence and the phoneme sequence. In other words, the fusion sequence contains the acoustic feature information about the target text 101.
[0045] Step 203: obtaining a target synthesized speech generated by the speech synthesis model according to the reference speech and the fusion sequence.
[0046] In summary, in the process of generating a target synthesized speech based on a target text, the method supplements the target text with corresponding acoustic feature information. Specifically, instead of directly inputting the target text and the reference speech into the speech synthesis model, an acoustic feature sequence used to represent the acoustic features of the target text is first generated, and then the acoustic feature sequence and the phoneme sequence of the target text extracted by the speech synthesis model are fused into a fusion sequence, so that the speech synthesis model generates a corresponding target synthesized speech based on the fusion sequence and the reference speech. In this way, the speech synthesis model can more accurately convert the target text into a target synthesized speech, thereby reducing pronunciation errors of the target synthesized speech.
[0047] In an embodiment, the reference speech is a speech sample used as a reference in the speech synthesis model, which is usually artificially recorded and represents specific language, accent, speech style, etc. The reference speech can be used to guide the speech synthesis model to generate a synthesized speech of a specific style or tone. In other words, the speech synthesis model can generate similar speech output by learning the speech features in the reference speech.
[0048] In an embodiment, the target synthesized speech is a sound representation of the target text to be synthesized, used to convey the semantic information of the target text. For example, if the semantic information of the target text is "Hello, world!", the target synthesized speech is expressed using the speech features of the reference speech.
[0049] In an embodiment, as shown in FIG. 5, the speech synthesis model 2 generates a target synthesized speech 108 according to the reference speech 104 and the fusion sequence 402.
[0050] It should be noted that in the inference stage of the model, the difference between the speech synthesis model 2 of FIG. 5 (the speech synthesis model proposed in the present application, which includes the acoustic feature predictor 3) and the speech synthesis model 1 of FIG. 1 lies in the processing logic within the dashed box 10. The modules (including the audio encoder 105, the acoustic model 106, and the audio decoder 107) of the speech synthesis model 2 of FIG. 5, except for the dashed box 10, can adopt the model architecture of the corresponding modules in the speech synthesis model 1 of FIG. 1.
[0051] As shown in FIG. 6, FIG. 6 shows a training method of the acoustic feature predictor 3 of FIG. 5, and the training method of FIG. 6 will be explained in combination with FIG. 7. Specifically, the following steps 601 to 604 are included.
[0052] Step 601: obtaining a training data set, the training data set including text data and its corresponding speech data.
[0053] In an embodiment, the training dataset can be trained using a public corpus, such as the AISHELL dataset disclosed at present, which includes 178 hours of speech data. Of course, a dataset related to a specific task and scenario can also be collected and labeled, for example, using a recording device to record speech data and labeling the text data corresponding to the speech data. The present specification does not limit the source of the training dataset.
[0054] The training dataset can include text data and corresponding speech data. For example, the text data can be a sentence "Hello, world!", and the speech data can be a recording of reading the text data.
[0055] Step 602: inputting the text data into the acoustic feature predictor to be trained, and generating a first acoustic feature sequence corresponding to the text data through the acoustic feature predictor.
[0056] As shown in FIG. 7, the text data 701 is input into the acoustic feature predictor 3 to be trained, and a first acoustic feature sequence 702 corresponding to the text data 701 is generated through the acoustic feature predictor 3.
[0057] Step 603: inputting the speech data into the speech representation model to generate a speech representation vector corresponding to the speech data through the speech representation model; and inputting the speech representation vector and the text data into the acoustic feature aligner to convert the speech representation vector into a second acoustic feature sequence through the acoustic feature aligner.
[0058] In an embodiment, as shown in FIG. 7, the speech data 707 is input into the speech representation model 703 to generate a speech representation vector 704 corresponding to the speech data 707 through the speech representation model 703; and the speech representation vector 704 and the text data 701 are input into the acoustic feature aligner 705 to convert the speech representation vector 704 into a second acoustic feature sequence 706 through the acoustic feature aligner 705.
[0059] In an embodiment, the speech representation model 703 can specifically be a trained HuBERT model, and the speech representation model 703 can also be other trained deep learning models that have the ability to convert speech data into a speech representation vector 704. The present application does not make any limitation in this regard.
[0060] In an embodiment, the acoustic feature aligner 705 can employ a model architecture of a common Transform model, and the trained Transform model has the capability of converting the speech representation vector 704 into the second acoustic feature sequence 706. Of course, it can also be a custom model architecture that has the capability of converting the speech representation vector 704 into the second acoustic feature sequence 706. The present application does not impose any limitation on the model architecture of the acoustic feature aligner.
[0061] It should be noted that, in the training phase of the model, the embodiment adds two auxiliary training modules (including the speech representation model 703 and the acoustic feature aligner 705) to the speech synthesis model 2 of FIG. 5 for training the acoustic feature predictor 3 in the speech synthesis model 2 of FIG. 5, as shown in the processing logic of the dashed box 70 of FIG. 7. Either auxiliary training module can only appear in the training phase for the acoustic feature predictor 3. In the inference phase of the model, only the model architecture of the speech synthesis model 2 shown in FIG. 5 can be used. In other words, the actual inference process of the speech synthesis model 2 is shown in FIG. 5 and does not include the processing of the dashed box 70.
[0062] Step 604: updating the parameters of the acoustic feature predictor based on the similarity loss between the first acoustic feature sequence and the second acoustic feature sequence.
[0063] In an embodiment, the parameters of the acoustic feature predictor can be updated based on a loss function such as mean square error, absolute value error, or cross-entropy loss between the first acoustic feature sequence and the second acoustic feature sequence. It can be understood that the mean square error, absolute value error, or cross-entropy loss are all used to measure the similarity loss between the first acoustic feature sequence and the second acoustic feature sequence, and to measure the difference between them, so as to update the parameters of the acoustic feature predictor. By continuously updating the parameters of the acoustic feature predictor, the first acoustic feature sequence and the second acoustic feature sequence can be made more and more close.
[0064] The embodiment extracts the speech features of the speech data through the speech representation model, and further extracts the acoustic features in the speech features through the acoustic feature aligner. The similarity loss between the first acoustic feature sequence output by the acoustic feature predictor and the second acoustic feature sequence output by the acoustic feature aligner is used to reduce the difference between the second acoustic feature sequence extracted from the speech data and the first acoustic feature sequence extracted from the text data by the acoustic feature predictor, so as to train the acoustic feature predictor to have the capability of extracting acoustic features from text data.
[0065] In an embodiment, the training method for other modules of the speech synthesis model 2 (modules in FIG. 5 other than the acoustic feature predictor 3) can employ the loss function defined in FIG. 1 for other modules of the speech synthesis model 1. For example, if the speech synthesis model is a VALL-E model, the loss function defined by the VALL-E model for each module can be employed to train other modules. The present application does not make any limitation in this regard.
[0066] In order to overcome the problem that the speech representation model 703 in FIG. 7 generates a speech representation vector 704 containing too many interference factors related to the speech style, which further causes the acoustic feature sequence generated by the trained acoustic feature predictor 3 to also contain the interference factors, and finally interferes with the quality and naturalness of the synthesized speech generated by the speech synthesis model.
[0067] The present embodiment further processes the speech representation vector 704 output by the speech representation model 703 to eliminate the interference factors of the speech representation vector 704 before inputting it to the acoustic feature aligner 705.
[0068] Specifically, as shown in FIG. 8, all speech representation vectors 704 are divided into different categories by clustering processing 801, wherein the similarity between speech representation vectors 704 in the same category is higher than the similarity between speech representation vectors 704 in different categories. In this way, speech representation vectors with similar speech styles can be clustered into the same category.
[0069] Specifically, all speech representation vectors can be divided into different categories using a clustering method. The clustering method can be K-means or a density-based spatial clustering of applications with noise (DBSCAN) or the like. The present application does not make any limitation on the clustering method.
[0070] Taking the K-means method as an example, the specific division steps are as follows:
[0071] ① Predefine any 512 speech representation vectors as initial centroids.
[0072] ② For any speech representation vector, calculate the distance between it and each initial centroid, and assign it to the nearest cluster.
[0073] ③ Calculate the average value of the speech representation vectors in each cluster as a new centroid.
[0074] (4) Repeat steps (2) and (3) to iteratively reassign the speech representation vectors and update the centroids until the centroids do not change or a preset number of iterations is reached.
[0075] Finally, all speech representation vectors are divided into different categories.
[0076] For all speech representation vectors of the same category, each speech representation vector in the same category is converted into a numerical value identical number, and the number is converted into an embedding vector. For example, for the speech representation vectors [1, 2, 3], [2, 2, 3] and [1, 2, 2] in category 1, they are all converted into the number 1, where the number 1 can be the number of the category, and then the number 1 is converted into the embedding vectors [0.1, 0.2, 0.2], [0.1, 0.2, 0.2] and [0.1, 0.2, 0.2] respectively corresponding to each speech representation vector. It can be seen that through the above processing, the differences in speech style between speech representation vectors with similar speech styles can be eliminated, thereby eliminating the interference factors of speech representation vectors related to speech style.
[0077] Of course, in addition to the above manner, each speech representation vector in the same category can also be concatenated with the category embedding vector representing the category to form an embedding vector corresponding to the speech representation vector. For example, for the category embedding vector of category 1 [1, 1, 1], the speech representation vectors [1, 2, 3], [2, 2, 3] and [1, 2, 2] in category 1. By concatenating each speech representation vector with the category embedding vector, embedding vectors [1, 2, 3, 1, 1, 1], [2, 2, 3, 1, 1, 1] and [1, 2, 2, 1, 1, 1] are obtained. It can be seen that through the above processing, the differences in speech style between speech representation vectors with similar speech styles can be partially eliminated, and the interference phonemes of speech representation vectors related to speech style can be reduced.
[0078] The embedding vectors are input to the acoustic feature aligner, and each embedding vector is converted into a corresponding second acoustic feature sequence by the acoustic feature aligner.
[0079] In this embodiment, by clustering similar speech representation vectors together, it means that the speech representation vectors of the same category represent similar acoustic features. Then, the speech representation vectors of the same category are converted into the same or similar embedding vectors. Thus, the converted embedding vectors eliminate or reduce the interference factors related to speech style, helping the speech synthesis model to better capture and maintain the consistency of the speech, thereby improving the quality and naturalness of the synthesized speech.
[0080] Since it is a challenging task to obtain a large-scale speech dataset, it requires a lot of time, resources and cost, and when collecting the speech dataset, the data privacy and compliance issues also need to be paid attention to. Therefore, in order to ensure that there is enough speech data to fully train the speech synthesis model, the application provides a method of generating synthetic speech data to expand the training dataset in FIG. 9. Specifically, it includes the following steps 901 to 903.
[0081] Step 901: Obtain an initial dataset, which includes initial text data and its corresponding initial speech data.
[0082] The initial dataset can be a public corpus, such as the currently public AISHELL dataset.
[0083] Step 902: Generate synthetic speech data using at least part of the initial speech data in the initial dataset.
[0084] In an embodiment, part of the initial speech data in the initial dataset can be selected to generate synthetic speech data. For example, 10% of the initial speech data in the initial dataset can be selected to generate corresponding synthetic speech data, or all of the initial speech data in the initial dataset can be used to generate corresponding synthetic speech data.
[0085] In an embodiment, in view of the problem that it is difficult to collect speech data with a single timbre and long duration in the real world, synthetic speech data corresponding to this type of speech data can be generated to make up for the problem of too few of this type of speech data.
[0086] In an embodiment, for each initial speech data that needs to generate synthetic speech data, the initial speech data is input into a timbre transformation model to obtain a plurality of synthetic speech data output by the model, and any two of the plurality of synthetic speech data have timbre differences. For example, for the speech data of "Hello, world!", a plurality of synthetic speech data of "Hello, world!" read by different timbres can be obtained by the timbre transformation model.
[0087] In an embodiment, the timbre transformation model can be a trained UNet model. Of course, it can also be other trained models, and the other trained models have the ability of the UNet model.
[0088] Step 903: Add the synthetic speech data to the expanded dataset, and combine the initial dataset and the expanded dataset into the training dataset.
[0089] Through the embodiment, the synthesized speech data can be easily generated, so that the data scale of the training data set can be increased to improve the generalization ability of the model. Moreover, since the multiple synthesized speech data corresponding to the same initial speech data differ in vocal color, the diversity of the training data set in vocal color is increased, and thus the model can learn more vocal color patterns.
[0090] Although the current speech synthesis model can realize vocal color cloning, the inconsistency of the vocal color of the synthesized speech is still a problem that cannot be ignored in the speech synthesis task. For example, the first half of the synthesized speech is the vocal color feature of speaker A, and the second half is converted into the vocal color feature of speaker B, causing the inconsistency of the vocal color of the synthesized speech.
[0091] The present application generates a training data set containing vocal color perturbation through the method shown in FIG. 10 to enhance the consistency of the synthesized speech in vocal color generated by the trained speech synthesis model. Specifically, the following steps 1001 to 1003 are included.
[0092] Step 1001: Obtain an initial data set, wherein the initial data set includes initial text data and corresponding initial speech data.
[0093] Step 1002: Generate perturbed speech data using at least part of the initial speech data of the initial data set.
[0094] Step 1003: Add the perturbed speech data to the expanded data set, and combine the initial data set and the expanded data set into the training data set.
[0095] Step 1001 is the same as step 901, and specific embodiments can be seen in step 901, which will not be described here.
[0096] In an embodiment, part of the speech segment in any initial speech data is replaced by other speech segments containing the same semantics to obtain perturbed speech data corresponding to the any initial speech data.
[0097] In an embodiment, part of the speech segment in a certain initial speech data is replaced by a speech segment containing the same semantics of other initial speech data. For example, the first speech segment of initial speech data 1 is “hello, world”, the second speech segment of initial speech data 2 is “hello, world”, and initial speech data 1 and initial speech data 2 are not the same initial speech data; the first speech segment of initial speech data 1 can be replaced by the second speech segment of initial speech data 2, and the replaced initial speech data 1 is used as the perturbed speech data corresponding to initial speech data 1.
[0098] In an embodiment, any speech segment in the initial speech data is replaced by another speech segment in the initial speech data that contains the same semantics. For example, any speech segment in certain initial speech data is "Hello, world", and there is another speech segment "Hello, world" in another position in the initial speech data that has the same semantics as the any speech segment, the other speech segment is replaced by the any speech segment in the initial speech data to obtain the perturbed speech data corresponding to the initial speech data.
[0099] In the embodiment, the voice synthesis model trained based on the perturbed speech data with added voice timbre perturbation can have strong anti-interference ability to the voice timbre perturbation, thereby ensuring the voice timbre consistency of the synthesized speech.
[0100] Corresponding to the embodiments of the foregoing method, the present application also provides embodiments of a device and a terminal to which the device is applied.
[0101] As shown in FIG. 11, FIG. 11 is a structural schematic diagram of an electronic device according to an exemplary embodiment of the present application. At the hardware level, the device includes a processor 1102, an internal bus 1104, a network interface 1106, a memory 1108, and a non-volatile memory 1110, and of course can also include other hardware required by the business. One or more embodiments of the present application can be implemented in a software manner, such as reading a corresponding computer program from the non-volatile memory 1110 into the memory 1108 by the processor 1102 and then running. Of course, in addition to the software implementation, one or more embodiments of the present application do not exclude other implementation manners, such as a logic device or a combination of software and hardware, and the like, that is, the execution subject of the following processing flow is not limited to each logical module, but can also be hardware or a logic device.
[0102] As shown in FIG. 12, FIG. 12 is a block diagram of a voice synthesis device according to an exemplary embodiment of the present application. The voice synthesis device 12 can be applied to the electronic device as shown in FIG. 11 to implement the technical solutions of the present application. The voice synthesis device 12 includes an acoustic feature sequence generation unit 1204, a fusion sequence generation unit 1206, and a target synthesized speech generation unit 1208.
[0103] The acoustic feature sequence generation unit 1204 is configured to generate an acoustic feature sequence corresponding to a target text to be synthesized, the acoustic feature sequence being used to represent acoustic features of the target text.
[0104] The fusion sequence generation unit 1206 is configured to fuse the acoustic feature sequence and a phoneme sequence of the target text into a fusion sequence.
[0105] The target synthesized speech generation unit 1208 is configured to obtain a target synthesized speech generated by the speech synthesis model according to the reference speech and the fusion sequence.
[0106] Optionally, the speech synthesis model comprises an acoustic feature predictor, and the acoustic feature sequence generation unit 1204 is specifically configured to input the target text to be synthesized into the trained acoustic feature predictor, so as to generate an acoustic feature sequence corresponding to the target text by the acoustic feature predictor.
[0107] Optionally, the speech synthesis model comprises a speech representation model and an acoustic feature aligner, and the acoustic feature predictor is trained by the following operations: obtaining a training data set comprising text data and corresponding speech data; inputting the text data into the acoustic feature predictor to be trained, so as to generate a first acoustic feature sequence corresponding to the text data by the acoustic feature predictor to be trained; inputting the speech data into the speech representation model, so as to generate a speech feature vector corresponding to the speech data by the speech representation model; and inputting the speech feature vector and the text data into the acoustic feature aligner, so as to convert the speech feature vector into a second acoustic feature sequence by the acoustic feature aligner; and updating the parameters of the acoustic feature predictor based on the similarity loss between the first acoustic feature sequence and the second acoustic feature sequence.
[0108] Optionally, the apparatus 12 further comprises a clustering unit 1202. The clustering unit 1202 is configured to divide the speech feature vectors into different categories, wherein the similarity between the speech feature vectors in the same category is higher than the similarity between the speech feature vectors in different categories. The clustering unit 1202 is further configured to determine an embedding vector corresponding to each category after division, wherein the embedding vector corresponding to each category is respectively converted from a plurality of speech feature vectors belonging to the category. The inputting the speech feature vector and the text data into the acoustic feature aligner and converting the speech feature vector into a second acoustic feature sequence by the acoustic feature aligner comprises: inputting the embedding vector into the acoustic feature aligner, and converting each embedding vector in the embedding vector into a corresponding second acoustic feature sequence by the acoustic feature aligner.
[0109] Optionally, the clustering unit 1202 is specifically configured to divide the speech feature vectors into different categories by a clustering algorithm.
[0110] Optionally, the training data set comprises an initial data set and an extended data set, and the obtaining the training data set comprises: obtaining the initial data set, the initial data set comprising initial text data and corresponding initial speech data; generating synthetic speech data by using at least part of the initial speech data in the initial data set, wherein any initial speech data is input into a timbre transformation model to obtain a plurality of synthetic speech data output by the model, and any two of the plurality of synthetic speech data have timbre differences; adding the plurality of synthetic speech data to the extended data set, and combining the initial data set and the extended data set as the training data set.
[0111] Optionally, the training data set comprises an initial data set and an extended data set, and the obtaining the training data set comprises: obtaining the initial data set, the initial data set comprising initial text data and corresponding initial speech data; generating perturbed speech data by using at least part of the initial speech data in the initial data set, wherein part of the speech segments in the at least part of the initial speech data are replaced by other speech segments containing the same semantics to obtain perturbed speech data corresponding to the at least part of the initial speech data; adding the perturbed speech data to the extended data set, and combining the initial data set and the extended data set as the training data set.
[0112] The implementation process of the functions and roles of the modules in the above device is specifically described in the implementation process of the corresponding steps in the above method, which will not be repeated here.
[0113] For the device embodiment, since it basically corresponds to the method embodiment, the related parts can be referred to the part of the method embodiment. The device embodiments described above are only illustrative, and the modules described as separate components can or can not be physically separated, and the components displayed as modules can or can not be physical modules, i.e. they can be located in one place or distributed on multiple network modules. According to actual needs, part or all of the modules can be selected to achieve the purpose of the present application. Those skilled in the art can understand and implement without creative labor.
[0114] Those skilled in the art can understand that all or part of the steps of the foregoing method can be instructed by a program to relevant hardware (for example, a processor), and the program can be stored in a computer readable storage medium, such as a read-only memory, a magnetic disk or an optical disk. Alternatively, all or part of the steps of the foregoing embodiments can also be implemented using one or more integrated circuits. Accordingly, each module / unit in the foregoing embodiments can be implemented in the form of hardware, for example, by an integrated circuit to implement its corresponding function, or in the form of a software function module, for example, by a processor executing a program / instruction stored in a memory to implement its corresponding function. The present application is not limited to any specific form of combination of hardware and software.
[0115] The present application also provides a computer readable storage medium, which stores a computer program, and the program is executed by a processor to implement the steps of any of the foregoing voice synthesis methods provided by the present application.
[0116] In particular, computer readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.
[0117] The foregoing described aspects and implementations of the present application have been described in particular detail. Other implementations can be apparent to those of ordinary skill in the art from the description and drawings. The scope of the application should not be limited to the described implementations but should be given the widest scope of the appended claims and the following claims.
[0118] Other embodiments of the present application will be apparent to those skilled in the art from consideration of the specification and practice of the application disclosed herein. It is intended that the present application cover any and all variations of the application that come within the scope of the following claims and their equivalents. It is intended that the application be construed as including all such variations as fall within the scope of the appended claims.
[0119] It should be understood that the application is not limited to the precise construction that has been described above and illustrated in the accompanying drawings, and that various modifications and changes can be made by those skilled in the art without departing from the scope of the application. The scope of the application should be limited only by the appended claims.
[0120] The above merely provides the optional embodiments of the present application, and is not used to limit the present application, and any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of protection of the present application.
Claims
1. A speech synthesis method, comprising: generating an acoustic feature sequence corresponding to a target text to be synthesized, wherein the acoustic feature sequence is used to characterize acoustic features of the target text; fusing the acoustic feature sequence with the phoneme sequence of the target text extracted by the speech synthesis model into a fused sequence; A target synthesized speech generated by the speech synthesis model according to the reference speech and the fusion sequence is obtained.
2. The method according to claim 1, characterized in that The speech synthesis model includes an acoustic feature predictor, and generating an acoustic feature sequence corresponding to a target text to be synthesized includes: The target text to be synthesized is input into the trained acoustic feature predictor, so that the acoustic feature predictor generates an acoustic feature sequence corresponding to the target text.
3. The method according to claim 2, characterized in that The speech synthesis model includes a speech representation model and an acoustic feature aligner, and the acoustic feature predictor is trained by the following operations, including: Acquire a training data set, wherein the training data set includes text data and its corresponding voice data; Inputting the text data into the acoustic feature predictor to be trained, and generating a first acoustic feature sequence corresponding to the text data by the acoustic feature predictor to be trained; Inputting the speech data into a speech representation model, and generating a speech representation vector corresponding to the speech data through the speech representation model; and inputting the speech representation vector and the text data into an acoustic feature aligner, and converting the speech representation vector into a second acoustic feature sequence through the acoustic feature aligner; Based on the similarity loss between the first acoustic feature sequence and the second acoustic feature sequence, parameters of the acoustic feature predictor are updated.
4. The method according to claim 3, characterized in that The method further comprises: Classifying the speech representation vectors into different categories, wherein the similarity between speech representation vectors of the same category is higher than the similarity between speech representation vectors of different categories; Determine an embedding vector corresponding to each of the divided categories, wherein the embedding vector corresponding to each category is obtained by converting multiple speech representation vectors belonging to the category; Inputting the speech representation vector and the text data into the acoustic feature aligner, and converting the speech representation vector into a second acoustic feature sequence by the acoustic feature aligner, comprises: The embedding vectors are input into the acoustic feature aligner, and the acoustic feature aligner converts each embedding vector in the embedding vector into a corresponding second acoustic feature sequence.
5. The method according to claim 4, characterized in that The dividing the speech representation vectors into different categories includes: The speech representation vectors are divided into different categories by a clustering algorithm.
6. The method according to claim 3, characterized in that The training data set includes an initial data set and an extended data set, and obtaining the training data set includes: Acquire an initial data set, the initial data set including initial text data and its corresponding initial voice data; Inputting at least part of the initial speech data into a timbre conversion model to obtain a plurality of synthesized speech data output by the timbre conversion model, wherein there is a timbre difference between any two of the plurality of synthesized speech data; The plurality of synthesized speech data are added to the extended dataset, and the initial dataset and the extended dataset are merged into the training dataset.
7. The method according to claim 3, characterized in that The training data set includes an initial data set and an extended data set, and obtaining the training data set includes: Acquire an initial data set, the initial data set including initial text data and its corresponding initial voice data; Replacing part of the speech segments in at least part of the initial speech data with other speech segments containing the same semantics to obtain disturbed speech data corresponding to the at least part of the initial speech data; The disturbed speech data is added to the extended dataset, and the initial dataset and the extended dataset are merged into the training dataset.
8. A speech synthesis device, comprising: an acoustic feature sequence generating unit, configured to generate an acoustic feature sequence corresponding to a target text to be synthesized, wherein the acoustic feature sequence is used to characterize acoustic features of the target text; a fusion sequence generating unit, configured to fuse the acoustic feature sequence with the phoneme sequence of the target text extracted by the speech synthesis model into a fusion sequence; The target synthesized speech generation unit is used to obtain the target synthesized speech generated by the speech synthesis model according to the reference speech and the fusion sequence.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Voice synthesis training data generation method and related equipment
CN112037754A
Speech synthesis method and device based on artificial intelligence, computer equipment and medium
CN112837673A
Voice style migration method and device, readable medium and electronic equipment
CN112927674A
Speech synthesis method and speech synthesis model training method and device
CN113112987A
Speech synthesis method and device, equipment, storage medium and program product
CN114242033A