A training method for speech synthesis model and speech synthesis method for Cantonese
By using an initial network based on multiple language types to iteratively train the pronunciation synthesis model of local languages such as Cantonese, the problem of poor training effect in the existing technology is solved, and efficient and natural pronunciation synthesis effect is achieved.
Patent Information
- Application Number
- CN202210322437.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-29
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-03-29
AI Technical Summary
For local languages such as Cantonese, it is difficult for the existing technology to effectively train the pronunciation synthesis model, resulting in poor synthesis effect.
By obtaining the training sample set of the target language type and the initial network associated with it, iterative training is performed to obtain the speech synthesis model. This initial network is trained based on training sample sets of multiple language types and has stronger learning ability.
This greatly reduces the model training time, improves training efficiency, and can train a speech synthesis model with better performance on a small number of training samples, reducing the possibility of overfitting the model, thereby improving the naturalness of speech synthesis.
Smart Images

Figure CN114944144B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of speech processing technology, and more specifically, to a training method for a speech synthesis model for Cantonese and a speech synthesis method. Background Art
[0002] Speech synthesis, also known as text to speech (TTS), is a technology that can convert text information generated by a computer or input from the outside into understandable and fluent corresponding speech. It is an important research branch in the field of natural language processing.
[0003] In the related art, a large number of training samples are generally used for model training to obtain a speech synthesis model for synthesizing speech. However, for some local languages, such as Cantonese, Minnan dialect or languages used by some ethnic minorities, the number of training samples obtained is very limited, which will lead to poor synthesis effect of the speech synthesis model trained for the local language. Summary of the invention
[0004] In view of this, the present application proposes a training method for a speech synthesis model and a speech synthesis method for Cantonese.
[0005] In a first aspect, an embodiment of the present application provides a method for training a speech synthesis model for Cantonese, the method comprising: obtaining a first training sample set corresponding to a target language type, the first training sample set comprising a first text sample set and a first voice sample set, the first voice sample in the first voice sample set corresponding one-to-one to the first text sample in the first text sample set; obtaining a first initial network associated with the target language type as an initial model, the first initial network being trained based on a second training sample set corresponding to a plurality of language types, the plurality of language types being associated with the target language type; inputting the first text sample into the initial model to obtain a synthesized speech corresponding to the first text sample; iteratively training the initial model based on the synthesized speech corresponding to the first text sample and the first voice sample of the target language type corresponding to the first text sample until a first preset condition is met, thereby obtaining a trained speech synthesis model, the speech synthesis model being used to synthesize speech of the target language type corresponding to the text to be processed.
[0006] In a second aspect, an embodiment of the present application provides a speech synthesis method for Cantonese, the method comprising: obtaining a text to be processed; inputting the text to be processed into a pre-trained speech synthesis model to obtain a speech of a target language type corresponding to the text to be processed, the pre-trained speech synthesis model being obtained by iteratively training an initial model using a first training sample set corresponding to the target language type until a first preset condition is met, the first training sample set comprising a first text sample set and a first speech sample set, the first speech sample in the first speech sample set corresponding one-to-one to the first text sample in the first text sample set, the initial model being a first initial network associated with the target language type, the first initial network being trained based on a second training sample set corresponding to a plurality of language types, the plurality of language types being associated with the target language type.
[0007] In the solution provided by the present application, a first training sample set corresponding to a target language type is obtained, the first training sample set includes a first text sample set and a first voice sample set, and the first voice sample in the first voice sample set corresponds one-to-one to the first text sample in the first text sample set; a first initial network associated with the target language type is obtained as an initial model, the first initial network is trained based on a second training sample set corresponding to a plurality of language types, and the plurality of language types are associated with the target language type; a first text sample is input into the initial model to obtain a synthesized speech corresponding to the first text sample; based on the synthesized speech corresponding to the first text sample and the first voice sample of the target language type corresponding to the first text sample, the initial model is iteratively trained until a first preset condition is met, and a trained speech synthesis model is obtained, and the speech synthesis model is used to synthesize the speech of the target language type corresponding to the text to be processed. In this way, since the first initial network trained based on the second training sample set corresponding to multiple language types has stronger learning ability, using the first initial network as the initial model to train the speech synthesis model greatly reduces the training time of the model and improves the efficiency of model training; and it is possible to train a speech synthesis model with better performance on a smaller number of first training samples, and optimize the model on a small number of first training samples, which greatly reduces the possibility of overfitting of the model, thereby improving the naturalness of the speech synthesized by the speech synthesis model finally obtained by training. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0009] Figure 1 A schematic diagram showing an application scenario of a method for training a speech synthesis model for Cantonese provided in an embodiment of the present application.
[0010] Figure 2 A flowchart of a method for training a speech synthesis model for Cantonese provided in one embodiment of the present application is shown.
[0011] Figure 3 Shows Figure 2 FIG. 1 is a schematic diagram of a sub-step flow chart of step S240 in one implementation manner.
[0012] Figure 4 A flowchart of a method for training a speech synthesis model for Cantonese provided in another embodiment of the present application is shown.
[0013] Figure 5 Shows Figure 4 FIG. 1 is a schematic diagram of a sub-step flow chart of step S340 in one implementation manner.
[0014] Figure 6 Shows Figure 5 A schematic flowchart of the sub-steps of step S343 in one implementation manner.
[0015] Figure 7 A flowchart of a method for training a speech synthesis model for Cantonese provided in yet another embodiment of the present application is shown.
[0016] Figure 8 A flow chart of a speech synthesis method for Cantonese provided in one embodiment of the present application is shown.
[0017] Fig. 9 It is a block diagram of a training device for a speech synthesis model for Cantonese provided according to an embodiment of the present application.
[0018] Fig.10 It is a block diagram of a speech synthesis device for Cantonese provided according to an embodiment of the present application.
[0019] Fig.11 It is a block diagram of a computer device according to an embodiment of the present application for executing a training method for a speech synthesis model for Cantonese according to an embodiment of the present application.
[0020] Fig.12 It is a storage unit of an embodiment of the present application for storing or carrying program codes for implementing a training method for a speech synthesis model for Cantonese according to an embodiment of the present application. DETAILED DESCRIPTION
[0021] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.
[0022] Speech synthesis, also known as text to speech (TTS), is a technology that can convert text information generated by a computer or input from the outside into understandable and fluent corresponding speech. It is an important research branch in the field of natural language processing.
[0023] In the related art, a large number of training samples are generally used for model training to obtain a speech synthesis model for synthesizing speech. However, for some local languages, such as Cantonese, Minnan dialect or languages used by some ethnic minorities, the number of training samples obtained is very limited, which will lead to poor synthesis effect of the speech synthesis model trained for the local language.
[0024] In response to the above problems, the inventors proposed a training method and a speech synthesis method for a Cantonese speech synthesis model, in which a first text sample is input into a first initial network associated with a target language type and trained based on a second training sample set corresponding to multiple language types, and iterative training is performed in order to obtain a speech synthesis model of the target language type.
[0025] See also Figure 1 , Figure 1A schematic diagram of an application scenario of a method for training a speech synthesis model for Cantonese provided in an embodiment of the present application, the application scenario includes a training system 10 for a speech synthesis model for Cantonese. The training system 10 for a speech synthesis model for Cantonese includes a computer device 110 and a first training sample set 120. The computer device 110 may be an electronic terminal with a data processing function, including but not limited to a smart phone, a tablet computer, and a laptop computer; of course, the computer device 110 may also be a server, which may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, network acceleration services (Content Delivery Network, CDN), and big data and artificial intelligence platforms. The first training sample set 120 includes a first text sample set 121 and a first voice sample set 122. The first voice sample in the first voice sample set 122 corresponds one-to-one to the first text sample in the first text sample set 121. It can be understood that the first training sample set 120 can be a pre-stored training sample directly obtained by the computer device 110 from a local database, or the required training sample can be downloaded from a network database through a wired or wireless network. Of course, other methods of obtaining training samples are also within the protection scope of this application and are no longer specifically limited here.
[0026] In some embodiments, the computer device 110 obtains a first training sample set 120 corresponding to the target language type; obtains a first initial network associated with the target language type as an initial model, the first initial network is trained based on a second training sample set corresponding to a plurality of language types, and the plurality of language types are associated with the target language type; the computer device 110 then inputs a first text sample 121 into the initial model to obtain a synthesized speech corresponding to the first text sample 121; based on the synthesized speech corresponding to the first text sample 121 and a first speech sample 122 of the target language type corresponding to the first text sample 121, the initial model is iteratively trained until a first preset condition is met, thereby obtaining a trained speech synthesis model, the speech synthesis model being used to synthesize a speech of the target language type corresponding to the text to be processed.
[0027] Please refer to Figure 2 , Figure 2 A flowchart of a method for training a speech synthesis model for Cantonese is provided in accordance with an embodiment of the present application. Figure 2The training method for the speech synthesis model for Cantonese provided in the embodiment of the present application is described in detail. The training method for the speech synthesis model for Cantonese may include the following steps:
[0028] Step S210: Acquire a first training sample set corresponding to the target language type, wherein the first training sample set includes a first text sample set and a first voice sample set, and the first voice sample in the first voice sample set corresponds one-to-one to the first text sample in the first text sample set.
[0029] In this embodiment, the target language type is Cantonese, and of course it can be any language of any country, including but not limited to Mandarin Chinese, Hakka dialect, Minnan dialect, Shanghai dialect, Sichuan dialect, Northeast dialect, British English, American English, Japanese, Korean, etc. It can be understood that the target language type can be determined according to actual needs. For example, if it is desired to train a speech synthesis model for synthesizing Cantonese, Cantonese is used as the above target language type; if it is desired to train a speech synthesis model for synthesizing Sichuan dialect, Sichuan dialect is used as the above target language type. This embodiment does not limit this.
[0030] Optionally, after determining the target language type, a first training sample set corresponding to the target language type is further obtained, the first training sample set including a first text sample set and a first voice sample set, and the first voice sample in the first voice sample set corresponds to the first text sample in the first text sample set. For example, if the target language type is Cantonese, if the first voice sample set includes a first Cantonese voice sample of "hello", the first text sample set also includes a first text sample of "hello" corresponding thereto.
[0031] Step S220: obtaining a first initial network associated with the target language type as an initial model, wherein the first initial network is trained based on a second training sample set corresponding to a plurality of language types, and the plurality of language types are associated with the target language type.
[0032] In this embodiment, the first initial network may be a neural network, including but not limited to deep neural networks (DNN), feedforward neural networks (FF), recurrent neural networks (RNN), long short-term memory networks (LSTM), deep residual networks (DRN) and other neural networks. It can be understood that if an untrained neural network is directly used as the initial model for training the speech synthesis model, since the initial model parameters of the untrained neural network are artificially set or randomly assigned, it is uncertain whether the initial model parameters are appropriate. When the artificially set or randomly assigned initial model parameters are far from the ideal model parameters, it will lead to a long time to train the initial model using the first training sample set, which will waste a lot of computing resources and be time-consuming and labor-intensive; at the same time, a large number of first training samples will be required to train and adjust the parameters of the initial model. If it is not possible to obtain enough first training samples of the target language type, the speech synthesis of the trained speech synthesis model will be unnatural, that is, the speech synthesis effect is very poor.
[0033] Based on this, a neural network with parameters initialized in advance can be used as the initial model of the speech synthesis model for training. It can be understood that a neural network with parameters initialized in advance is equivalent to being endowed with more prior knowledge. Analogously, if a person who has not graduated from elementary school listens to advanced mathematics, he obviously cannot understand it; while if a high school graduate who has scored full marks in college entrance examination mathematics listens to it, he may learn much more easily. The reason is that the amount of knowledge accumulated by the two is different. The former has accumulated less knowledge, while the latter has accumulated more knowledge. Therefore, the latter can learn knowledge that has not been learned faster. Similarly, a neural network with parameters initialized in advance has a stronger learning ability than a neural network that is artificially set or randomly assigned. In this way, the time for training a speech synthesis model with a neural network with parameters initialized in advance is shorter, and the number of first training samples used is also smaller, that is, after training on fewer first training samples, a speech synthesis model with good performance can be obtained.
[0034] Therefore, the first initial network trained based on the second training sample set corresponding to multiple language types can be obtained as the initial model. It can be seen that the first initial network is a network with better model parameters trained by other tasks first, that is, the first initial network is a network that can synthesize speech of any language type in multiple language types through training first, and the network has better model parameters for the speech synthesis function. Based on this, the first initial network is selected as the initial model to train the speech synthesis model of the target language type in a shorter time, and the speech synthesis effect of the trained speech synthesis model is also better. In addition, the aforementioned multiple language types are language types associated with the target language type. For example, the target language type is "Cantonese", and the multiple language types can include "Cantonese", "Dialects in Guangxi", "Dialects in Hong Kong", etc., that is, language types with similar pronunciation to the speech of the target language type. In this way, the first initial network can have model parameters that are closer to the speech for synthesizing the target language type, further improving the subsequent training speed of the speech synthesis model of the target language type based on the first initial network; at the same time, fewer first training samples of the target language type can be used for training, and it can be ensured that the trained speech synthesis model still has a good speech synthesis effect.
[0035] In some embodiments, a computer device may pre-train multiple first initial networks and store the multiple pre-trained first initial networks. Different first initial networks may be used as initial models for speech synthesis models of different language types. When the computer device needs to train a speech synthesis model of a target language type, the pre-trained first initial network associated with the target language type is obtained from the multiple pre-trained first initial networks as the initial model. In this way, the training efficiency of the speech synthesis model of the target language type can be improved, and the reuse of multiple first initial networks can be achieved, thereby avoiding the need to train the first initial network associated with the language type every time a speech synthesis model of a different language type needs to be trained, thereby saving computing resources of the computer device.
[0036] In other embodiments, the computer device may obtain a second training sample set of multiple language types associated with the target language type when determining that the speech synthesis model of the target language type needs to be trained, and then train the first initial network based on the second training sample set of multiple language types; and use the trained first initial network as the initial model. It can be understood that the first initial network is associated with each language type in the multiple language types in addition to being associated with the target language type. Therefore, the trained first initial network can be stored, and can be stored in the computer device, or it can be uploaded to the cloud server for storage to save storage resources in the computer device. In this way, the trained first initial network is stored, so that when the speech synthesis model of the language type in the multiple language types is subsequently trained, the stored first initial network can be directly called without retraining, which greatly improves the training efficiency of the speech synthesis model, while also saving the computing resources of the computer device and reducing the computing cost.
[0037] Step S230: input the first text sample into the initial model to obtain synthesized speech corresponding to the first text sample.
[0038] After the initial model is obtained, the first text sample is input into the initial model, the initial model text is converted into a phoneme sequence, and the start and end time, frequency change and other information of each phoneme are marked, and then the synthesized speech corresponding to the first text sample is generated according to the phoneme sequence (and the marked start and end time, frequency change and other information). Among them, the method of generating speech includes but is not limited to the splicing method, the parameter method and the vocal tract simulation method.
[0039] Step S240: Based on the synthesized speech corresponding to the first text sample and the first speech sample of the target language type corresponding to the first text sample, the initial model is iteratively trained until a first preset condition is met, thereby obtaining a trained speech synthesis model, which is used to synthesize the speech of the target language type corresponding to the text to be processed.
[0040] In some embodiments, see Figure 3 , step S240 may include the following steps:
[0041] Step S241: determining a second loss value according to a difference between the synthesized speech corresponding to the first text sample and the first speech sample of the target language type corresponding to the first text sample.
[0042] In this embodiment, the difference between the synthesized speech and the first speech sample of the target language type corresponding to the first text sample can be calculated by a loss function to obtain a second loss value. Among them, the loss function can be a mean square error loss function, and of course it can also be other loss functions, which is not limited in this embodiment. Taking the mean square error loss function as an example, since the mean square error loss function is a measure used to calculate the degree of difference between the estimator and the estimated quantity, the mean square error loss function can be used to calculate the degree of difference between the speech features of the synthesized speech and the speech features of the first speech sample of the target language type corresponding to the first text sample, and obtain the second loss value. It can be understood that the smaller the second loss value is, the smaller the difference between the synthesized speech and the corresponding first speech sample is, that is, the closer the synthesized speech is to the real speech of the target language type, and the higher the naturalness of the speech synthesis.
[0043] Step S242: According to the second loss value, the initial model is iteratively trained until the first preset condition is met, thereby obtaining a trained speech synthesis model.
[0044] Based on this, after obtaining the second loss value used to characterize the difference between the synthesized speech and the corresponding first speech sample, the initial model can be iteratively trained according to the second loss value until the first preset condition is met, and the trained initial model is obtained as the above-mentioned speech synthesis model.
[0045] Among them, the first preset condition can be: the second loss value is less than the preset value, the second loss value no longer changes, or the number of training times reaches the preset number of times, etc. It can be understood that after the initial model is iteratively trained for multiple training cycles according to the first training sample set, each training cycle includes multiple iterative trainings, and the parameters in the initial model are continuously optimized, so that the above-mentioned second loss value becomes smaller and smaller, and finally becomes a fixed value, or is less than the preset value. At this time, it means that the initial model has converged; of course, it can also be determined that the initial model has converged after the number of training times reaches the preset number of times. At this time, the initial model can be used as the above-mentioned speech synthesis model. Among them, the preset value and the preset number of times are both pre-set, and their values can also be adjusted according to different application scenarios. The optimization of the parameters in the initial model can be optimized by gradient descent, for example, batch gradient descent method, stochastic gradient descent method or small batch gradient descent method. Of course, the parameters in the initial model can also be optimized by using optimization algorithms such as Newton method, quasi-Newton method, DFP (Davidon-Fletcher-Powell algorithm) algorithm, or improved iterative scaling method, etc., which are not limited in this embodiment.
[0046] Furthermore, the trained speech synthesis model can be used to synthesize speech of the target language type corresponding to the text to be processed, wherein the speech synthesis model can be applied to a variety of application scenarios. For example, in the field of reading and listening to books, the text information in the book can be used as the aforementioned text to be processed, and then the speech synthesis model is used to synthesize the text to be processed into the speech of the corresponding target language type, which can provide users with personalized reading functions, free up the user's hands and eyes, and provide a more extreme reading experience. For another example, in the order broadcast scenario (taxi, restaurant call number or queue call number), the text information of the order broadcast is used as the aforementioned text to be processed, and the speech synthesis model is used to synthesize the text to be processed into the speech of the corresponding target language type to realize voice order broadcast, so that users can more conveniently obtain the broadcast notification information.
[0047] In this embodiment, a first initial network associated with the target language type and pre-initialized with parameters is obtained as an initial model, and then the initial model is trained using the first training sample set to obtain a trained speech synthesis model. In this way, since the first initial network with parameters pre-initialized has a stronger learning ability, the training time of the model is greatly reduced and the efficiency of model training is improved by using the first initial network as the initial model for training the speech synthesis model; moreover, a speech synthesis model with better performance can be trained on a smaller number of first training samples, and the model is optimized on a small number of first training samples, which greatly reduces the possibility of overfitting of the model, thereby improving the naturalness of the speech synthesized by the speech synthesis model obtained by training.
[0048] Please refer to Figure 4 , Figure 4 A flowchart of a method for training a speech synthesis model for Cantonese is provided in another embodiment of the present application. Figure 4 The training method for the speech synthesis model for Cantonese provided in the embodiment of the present application is described in detail. The training method for the speech synthesis model for Cantonese may include the following steps:
[0049] Step S310: obtaining a first training sample set corresponding to the target language type, wherein the first training sample set includes a first text sample set and a first voice sample set, and the first voice sample in the first voice sample set corresponds one-to-one to the first text sample in the first text sample set.
[0050] In this embodiment, the specific implementation of step S310 can refer to the content of the above-mentioned embodiment, which will not be repeated here.
[0051] Step S320: Obtain a second training sample set corresponding to each language type in the multiple language types, wherein the second training sample set includes a second text training sample set and a second speech training sample, and the second speech training samples in the second speech training sample set correspond one-to-one to the second text training samples in the second text training sample set.
[0052] In this embodiment, the multiple language types are language types associated with the target language type, that is, language types with similar speech pronunciation to the target language type. After determining that the speech synthesis model of the target language type needs to be trained, a second training sample set corresponding to each language type in the multiple language types associated with it can be obtained. In this way, the computer device uses the second training sample set corresponding to each language type to initialize the network in advance, thereby improving the generalization ability of the network when facing different tasks, that is, improving the network's speech synthesis ability for synthesizing different language types.
[0053] Step S330: inputting the second text training samples of each language type into the second initial network corresponding to each language type respectively, to obtain the synthesized speech of each language type corresponding to the second text training samples.
[0054] In this embodiment, the specific implementation method of generating the synthesized speech of each language type corresponding to the second text training sample can refer to the content of the above embodiments, which will not be repeated here.
[0055] Among them, the second initial network can be an untrained neural network. The second text training samples of each language type are respectively input into the second initial network corresponding to each language type to obtain the synthesized speech of each language type corresponding to the second text training samples. In this way, the computer device can optimize the parameters of the second initial network of each language type based on the synthesized speech of each language type.
[0056] For example, if the number of the multiple language types is 3, including language type 1, language type 2, and language type 3, when the second text training sample is input, the second initial network corresponding to language type 1 can output the synthesized speech of language type 1 corresponding to the second text training sample, the second initial network corresponding to language type 2 can output the synthesized speech of language type 2 corresponding to the second text training sample, and the second initial network corresponding to language type 3 can output the synthesized speech of language type 3 corresponding to the second text training sample. In this way, the parameters of the second initial network of each language type can be optimized according to the synthesized speech of the three language types.
[0057] Step S340: Based on the synthesized speech of each language type corresponding to the second text training sample and the second speech training sample of each language type corresponding to the second text training sample, iteratively train the second initial network corresponding to each language type until a second preset condition is met, and obtain the trained second initial network corresponding to each language type as the first initial network corresponding to each language type.
[0058] In some embodiments, the second training sample set further includes a second text test sample set and a second speech test sample set, and the second speech test samples in the second speech test sample set correspond to the second text test samples in the second text test sample set in one-to-one correspondence. Figure 5 , step S340 may include the following steps:
[0059] Step S341: Based on the synthesized speech of each language type corresponding to the second text training sample and the second speech training sample of each language type corresponding to the second text training sample, iteratively train the second initial network corresponding to each language type until a third preset condition is met, and obtain the trained second initial network corresponding to each language type as the third initial network corresponding to each language type.
[0060] In some embodiments, a third loss value of the second initial network corresponding to each language type is determined based on the difference between the synthesized speech of each language type corresponding to the second text training sample and the second speech training sample of each language type corresponding to the second text training sample; and based on the third loss value of the second initial network corresponding to each language type, the second initial network corresponding to each language type is iteratively trained until a third preset condition is met, so as to obtain the trained second initial network corresponding to each language type as the third initial network corresponding to each language type.
[0061] Among them, the third preset condition can be: the third loss value is less than the preset value, the third loss value no longer changes, or the number of training times reaches the preset number of times, etc. It can be understood that after the second initial network corresponding to each language type is iteratively trained for multiple training cycles according to the second text training samples corresponding to each language type, wherein each training cycle includes multiple iterative trainings, the parameters in the initial model are continuously optimized, so that the above-mentioned third loss value becomes smaller and smaller, and finally becomes a fixed value, or is less than the preset value. At this time, it means that the initial model has converged; of course, it can also be determined that the second initial network has converged after the number of training times reaches the preset number of times. At this time, the second initial network can be used as the third initial network. Among them, the preset value and the preset number of times are both pre-set, and their values can also be adjusted according to different application scenarios. The parameters in the second initial network corresponding to each language type can be optimized by gradient descent, for example, batch gradient descent, stochastic gradient descent or mini-batch gradient descent. Of course, Newton's method, quasi-Newton's method, DFP algorithm, or improved iterative scaling method and other optimization algorithms can also be used to optimize the parameters in the second initial network corresponding to each language type. This embodiment does not limit this.
[0062] Exemplarily, still taking multiple language types including language type 1, language type 2 and language type 3 as an example, based on the synthesized speech of language type 1 corresponding to the second text training sample and the second speech training sample of language type 1 corresponding to the second text training sample, the second initial network corresponding to language type 1 is iteratively trained until the third preset condition is met, and the trained second initial network corresponding to language type 1 is obtained as the third initial network corresponding to language type 1. Similarly, the third initial network corresponding to language type 2 and the third initial network corresponding to language type 3 can be trained to obtain. It can be understood that at this time, the third initial network corresponding to language type 1 has a better effect on synthesizing text into speech of language type 1, the third initial network corresponding to language type 2 has a better effect on synthesizing text into speech of language type 2, and the third initial network corresponding to language type 3 has a better effect on synthesizing text into speech of language type 3. It can be seen that the optimization of the parameters of the second initial network facing different tasks is achieved.
[0063] Step S342: inputting the second text test sample of each language type into the third initial network corresponding to each language type respectively, to obtain the synthesized speech of each language type corresponding to the second text test sample.
[0064] Step S343: Based on the synthesized speech of each language type corresponding to the second text test sample and the second speech test sample of each language type corresponding to the second text test sample, iteratively train the third initial network corresponding to each language type until the second preset condition is met, and obtain the trained third initial network corresponding to each language type as the first initial network corresponding to each language type.
[0065] It is understandable that after training to obtain a third initial network with good speech synthesis effect for each language type, the third initial networks of different language types may not be good at synthesizing speech of other language types in multiple language types. Therefore, the second text test sample set and the second speech test sample set can be used to optimize the parameters of the third initial network of each language type at the same time to obtain parameters that are good for each language type in multiple language types, that is, the network terms with these good parameters have a good effect in synthesizing speech of each language type.
[0066] Specifically, the second text test samples of each language type are respectively input into the third initial network corresponding to each language type to obtain the synthesized speech of each language type corresponding to the second text test samples; then based on the synthesized speech of each language type corresponding to the second text test samples and the second speech test samples of each language type corresponding to the second text test samples, the third initial network corresponding to each language type is iteratively trained until the second preset condition is met, and the trained third initial network corresponding to each language type is obtained as the first initial network corresponding to each language type.
[0067] In some embodiments, see Figure 6 , step S343 may include the following steps:
[0068] Step S3431: Obtain a total loss value based on the synthesized speech of each language type corresponding to the second text test sample and the second speech test sample of each language type corresponding to the second text test sample.
[0069] Specifically, according to the difference between the synthesized speech of each language type corresponding to the second text test sample and the second speech test sample of each language type corresponding to the second text test sample, the loss value corresponding to the third initial network of each language type is determined to obtain multiple first loss values; and the sum of the multiple first loss values is obtained as the above-mentioned total loss value.
[0070] Step S3432: According to the total loss value, iteratively train the third initial network corresponding to each language type until the second preset condition is met, and obtain the trained third initial network corresponding to each language type as the first initial network corresponding to each language type.
[0071] Among them, the second preset condition can be: the total loss value is less than the preset value, the total loss value no longer changes, or the number of training times reaches the preset number of times, etc. It can be understood that after the third initial network corresponding to each language type is iteratively trained for multiple training cycles according to the second text training sample corresponding to each language type, wherein each training cycle includes multiple iterative trainings, the parameters in the third initial network corresponding to each language type are continuously optimized, so that the above-mentioned total loss value becomes smaller and smaller, and finally becomes a fixed value, or less than the preset value. At this time, it means that the third initial network corresponding to each language type has converged; of course, it can also be determined that the third initial network corresponding to each language type has converged after the number of training times reaches the preset number of times. At this time, the trained third initial network corresponding to each language type can be used as the first initial network corresponding to each language type. Among them, the preset value and the preset number of times are both pre-set, and their values can also be adjusted according to different application scenarios. The parameters in the third initial network corresponding to each language type can be optimized by gradient descent, for example, batch gradient descent, stochastic gradient descent or mini-batch gradient descent. Of course, Newton's method, quasi-Newton's method, DFP algorithm, or improved iterative scaling method and other optimization algorithms can also be used to optimize the parameters in the third initial network corresponding to each language type. This embodiment does not limit this.
[0072] It can be understood that the contents of step S320 to step S340 are equivalent to the learning process of meta-learning. Specifically, with task as the basic unit, each task has its own independent loss function. During training, the training set (Support set) is first used to train the model of each task, and the loss value is calculated using its own independent loss function. The corresponding model parameters of each task are independently iteratively optimized for the loss value of each task to obtain a model for each task; then the query set (Query set) is used to perform performance testing on the optimized model of each task, that is, secondary model training. According to the sum of the loss values of all tasks, the model parameters of all tasks are uniformly optimized to obtain good model parameters for each task, and finally a model with good model parameters is obtained, which has good execution capabilities for each task. In this way, the overall learning ability of the model is improved, rather than the ability to solve a specific problem. During training, different tasks are constantly switched to achieve the purpose of optimizing network parameters. The final model can learn faster when facing new tasks. That is, the first initial network can complete the training at a faster speed when facing the training task of the speech synthesis model of any new language type.
[0073] Step S350: using the first initial network corresponding to any language type among the multiple language types as the initial model.
[0074] Based on this, after obtaining the first initial network corresponding to each language type, since the network structure and parameters of the first initial network corresponding to each language type are the same at this time, the first initial network corresponding to any language type in multiple language types can be used as the initial model.
[0075] Step S360: input the first text sample into the initial model to obtain synthesized speech corresponding to the first text sample.
[0076] Step S370: Based on the synthesized speech corresponding to the first text sample and the first speech sample of the target language type corresponding to the first text sample, the initial model is iteratively trained until a first preset condition is met, so as to obtain a trained speech synthesis model, which is used to synthesize the speech of the target language type corresponding to the text to be processed.
[0077] In this embodiment, the specific implementation of step S360 to step S370 can refer to the content of the above-mentioned embodiment, which will not be repeated here.
[0078] In this embodiment, firstly, the network of each language type is independently optimized through the second text training sample set and the second voice training sample set, and the optimized second initial network corresponding to each language type is obtained as the third initial network corresponding to each language type; then, based on the second text test sample set and the second voice test sample set, all the third initial networks are jointly optimized to obtain the optimized third initial network as the initial model. In this way, the initial model can use previous knowledge and experience to guide the learning of new tasks, so that the initial network has a stronger learning ability, so that it can have good performance after training on a small amount of sample data, that is, the speech synthesis model obtained by the initial model training has a good speech synthesis effect, and the training time of the model is greatly reduced, and the efficiency of model training is improved; and it can be realized that a speech synthesis model with good performance can be trained on a small number of first training samples, and the model optimization is performed on a small number of first training samples, which greatly reduces the possibility of overfitting of the model, thereby improving the naturalness of the speech synthesized by the speech synthesis model obtained by training.
[0079] Please refer to Figure 7 , Figure 7 A flowchart of a method for training a speech synthesis model for Cantonese is provided in accordance with another embodiment of the present application. Figure 7 The training method for the speech synthesis model for Cantonese provided in the embodiment of the present application is described in detail. The training method for the speech synthesis model for Cantonese may include the following steps:
[0080] Step S401: obtaining a first training sample set corresponding to a target language type, wherein the first training sample set includes a first text sample set and a first voice sample set, wherein a first voice sample in the first voice sample set corresponds one-to-one to a first text sample in the first text sample set.
[0081] In this embodiment, the specific implementation of step S401 can refer to the content of the above-mentioned embodiment, which will not be repeated here.
[0082] Step S402: obtaining a preset voice of a preset text in each preset language type among a plurality of preset language types.
[0083] Step S403: extracting speech features of the preset speech in each preset language type to obtain preset speech features corresponding to each preset language type.
[0084] In this embodiment, the computer device may pre-store multiple preset language types, preset texts, and preset voices of the preset texts in each preset language type. Based on this, after determining the target language type, the computer device may filter out multiple preset language types associated with the target language type from the multiple preset language types.
[0085] In some implementations, a preset speech in each preset language type may be acquired, and speech features of the preset speech in each preset language type may be extracted to obtain preset speech features corresponding to each preset language type.
[0086] Among them, the speech features can be extracted by feature extraction algorithms, and the feature extraction algorithm treasure house includes but is not limited to Mel Frequency Cepstrum Coefficient (MFCC) algorithm, filter bank (FBANK) algorithm, constant-Q transform (CQT) algorithm, Linear Predictive Cepstral Coefficient (LPCC) algorithm, Perceptual Linear Predictive Coefficient (PLP) algorithm, Linear Prediction Coefficients (LPC) algorithm, etc., and the corresponding speech features can be MFCC features, FBANK features, CQT features, LPCC features, PLP features or LPC features. Of course, the speech features of the preset speech can also be extracted by a pre-trained feature extraction network, which is not limited in this embodiment.
[0087] Step S404: Obtain a target voice of the preset text in the target language type.
[0088] Step S405: extracting speech features of the target speech as target speech features corresponding to the target language type.
[0089] Furthermore, the computer device can obtain the target speech of the preset text in the target language type, and extract the speech features of the target speech as the target speech features corresponding to the target language type. In this way, the target speech features and the preset speech features can be used as the screening basis to screen out multiple preset language types associated with the target language type from multiple preset language types.
[0090] Step S406: obtaining preset language types corresponding to preset speech features whose similarity with the target speech features reaches a preset similarity threshold, as the multiple language types.
[0091] In some implementations, the similarity between the target speech feature and the preset speech feature corresponding to each preset language type can be obtained to obtain multiple target similarities; the preset language types corresponding to the preset speech features whose target similarities reach the preset similarity threshold are obtained as multiple language types. In this way, using the speech features as the screening basis, multiple language types that are more similar in pronunciation to the target language type can be obtained, thereby making the first initial network obtained by training have model parameters that are closer to the speech used to synthesize the target language type, that is, it has the ability to better learn the speech synthesis of the target language type.
[0092] Step S407: Obtain a second training sample set corresponding to each language type in the multiple language types, wherein the second training sample set includes a second text training sample set and a second speech training sample, and the second speech training samples in the second speech training sample set correspond one-to-one to the second text training samples in the second text training sample set.
[0093] Step S408: inputting the second text training samples of each language type into the second initial network corresponding to each language type respectively, to obtain the synthesized speech of each language type corresponding to the second text training samples.
[0094] Step S409: Based on the synthesized speech of each language type corresponding to the second text training sample and the second speech training sample of each language type corresponding to the second text training sample, iteratively train the second initial network corresponding to each language type until a second preset condition is met, and obtain the trained second initial network corresponding to each language type as the first initial network corresponding to each language type.
[0095] Step S410: using the first initial network corresponding to any language type among the multiple language types as the initial model.
[0096] Step S411: input the first text sample into the initial model to obtain synthesized speech corresponding to the first text sample.
[0097] Step S412: Based on the synthesized speech corresponding to the first text sample and the first speech sample of the target language type corresponding to the first text sample, the initial model is iteratively trained until a first preset condition is met, so as to obtain a trained speech synthesis model, which is used to synthesize the speech of the target language type corresponding to the text to be processed.
[0098] In this embodiment, the specific implementation of steps S407 to S412 can refer to the contents of the above embodiments and will not be repeated here.
[0099] In this embodiment, by taking speech features as the screening basis, a plurality of language types that are more similar in pronunciation to the target language type can be obtained, so that the trained first initial network has model parameters that are closer to the speech for synthesizing the target language type, that is, it has the ability to better learn the speech synthesis of the target language type; the subsequent training speed of the speech synthesis model of the target language type based on the first initial network is further improved; at the same time, fewer first training samples of the target language type can also be used for training, and it can be ensured that the trained speech synthesis model still has a good speech synthesis effect.
[0100] Please refer to Figure 8 , Figure 8 A flowchart of a speech synthesis method for Cantonese provided in an embodiment of the present application is provided. Figure 8 The speech synthesis method for Cantonese provided in the embodiment of the present application is described in detail. The speech synthesis method for Cantonese may include the following steps:
[0101] Step S510: Obtain the text to be processed.
[0102] In this embodiment, the text to be processed refers to data presented in text form and containing content information. The text to be processed may be text that a user inputs into a computer device and hopes to perform speech synthesis on, or may be a broadcast text that the computer device obtains from a pre-stored broadcast text and needs to perform speech synthesis on, or may be a reply text generated by the computer device for the text input by the user, or may be a text obtained by the computer device by performing text conversion on a speech to be converted that is different from the target language type, and this embodiment does not limit this.
[0103] Step S520: input the text to be processed into a pre-trained speech synthesis model to obtain a synthesized speech of a target language type corresponding to the text to be processed, wherein the pre-trained speech synthesis model is obtained by iteratively training an initial model using a first training sample set corresponding to the target language type until a first preset condition is met, wherein the first training sample set includes a first text sample set and a first speech sample set, wherein a first speech sample in the first speech sample set corresponds one-to-one to a first text sample in the first text sample set, and the initial model is a first initial network associated with the target language type, wherein the first initial network is trained based on a second training sample set corresponding to a plurality of language types, and the plurality of language types are associated with the target language type.
[0104] Based on this, after obtaining the text to be processed, the text to be processed can be converted into a phoneme sequence, and then the phoneme sequence can be input into a pre-trained speech synthesis model to obtain a synthesized speech of the target language type corresponding to the text to be processed. Among them, the target language type is Cantonese, of course, it can also be any language of any country, including but not limited to Mandarin Chinese, Hakka dialect, Minnan dialect, Shanghai dialect, Sichuan dialect, Northeast dialect, British English, American English, Japanese, Korean, etc.
[0105] Exemplarily, the text to be processed is "Hello, the order has been completed, please pick up the food". After querying the phoneme table, the computer device can convert the text to be processed into 10 phoneme sequences: {n,i,n}, {h,a,o}, {d,i,n,g}, {d,a,n}, {y,i}, {w,a,n}, {c,h,e,n,g}, {q,i,n,g}, {q,u}, {c,a,n}. Each word in the text to be processed corresponds to a phoneme sequence. By inputting the above 10 phoneme sequences into the speech synthesis model, the synthesized speech of the target language type corresponding to the text to be processed can be synthesized.
[0106] In some embodiments, the speech synthesis model performs speech synthesis processing on the phoneme sequence to obtain a Mel spectrum corresponding to the text to be processed, and the Mel spectrum contains the audio features corresponding to each word in the text to be processed, and then performs Fourier transform processing on the Mel spectrum to obtain a synthesized speech of the target language type corresponding to the text to be processed. Optionally, in order to make the final synthesized speech more realistic, background noise data can be calculated based on a preset signal-to-noise ratio, and the background noise data can be added to the synthesized speech corresponding to the text to be processed, so that a speech with a real background environment can be obtained, and the speech also includes breathing sounds, which can better reflect the realism of the speech.
[0107] In this embodiment, since the first initial network with pre-initialized parameters has stronger learning ability, a speech synthesis model with better performance can be trained on a smaller number of first training samples, thereby greatly improving the naturalness of the synthesized speech based on the trained speech synthesis model.
[0108] Please refer to Fig. 9 , which shows a structural block diagram of a speech synthesis model training device 600 for Cantonese provided by an embodiment of the present application. The device 600 may include: a sample acquisition module 610, an initial model acquisition module 620, a speech synthesis module 630 and a model training module 640.
[0109] The sample acquisition module 610 is used to acquire a first training sample set corresponding to the target language type, wherein the first training sample set includes a first text sample set and a first voice sample set, wherein the first voice sample in the first voice sample set corresponds one-to-one to the first text sample in the first text sample set.
[0110] The initial model acquisition module 620 is used to acquire a first initial network associated with the target language type as an initial model, wherein the first initial network is trained based on a second training sample set corresponding to a plurality of language types, and the plurality of language types are associated with the target language type.
[0111] The speech synthesis module 630 is used to input the first text sample into the initial model to obtain the synthesized speech corresponding to the first text sample.
[0112] The model training module 640 is used to iteratively train the initial model based on the synthesized speech corresponding to the first text sample and the first speech sample of the target language type corresponding to the first text sample until a first preset condition is met, thereby obtaining a trained speech synthesis model, which is used to synthesize the speech of the target language type corresponding to the text to be processed.
[0113] In some embodiments, the initial model acquisition module 620 may include: a second sample acquisition unit, a second speech synthesis unit, a first initial network acquisition unit, and an initial model acquisition unit. The second sample acquisition unit may be used to acquire a second training sample set corresponding to each of the multiple language types, the second training sample set including a second text training sample set and a second speech training sample, the second speech training sample in the second speech training sample set corresponding to the second text training sample set. The second speech synthesis unit may be used to input the second text training sample of each language type into the second initial network corresponding to each language type, respectively, to obtain the synthesized speech of each language type corresponding to the second text training sample. The first initial network acquisition unit may be used to iteratively train the second initial network corresponding to each language type based on the synthesized speech of each language type corresponding to the second text training sample and the second speech training sample of each language type corresponding to the second text training sample, until the second preset condition is met, and obtain the trained second initial network corresponding to each language type as the first initial network corresponding to each language type. The initial model acquisition unit may be used to use the first initial network corresponding to any of the multiple language types as the initial model.
[0114] In this manner, the initial model acquisition module 620 may further include: a preset speech acquisition unit, a preset speech feature acquisition unit, a target speech acquisition unit, a target speech feature acquisition unit, and a language type determination unit. The preset speech acquisition unit may be used to acquire the preset speech of the preset text in each preset language type in the plurality of preset language types before acquiring the second training sample set corresponding to each language type in the plurality of language types. The preset speech feature acquisition unit may be used to extract the speech features of the preset speech in each preset language type to obtain the preset speech features corresponding to each preset language type. The target speech acquisition unit may be used to acquire the target speech of the preset text in the target language type. The target speech feature acquisition unit may be used to extract the speech features of the target speech as the target speech features corresponding to the target language type. The language type determination unit may be used to acquire the preset language type corresponding to the preset speech features whose similarity with the target speech features reaches a preset similarity threshold as the plurality of language types.
[0115] In some embodiments, the second training sample set further includes a second text test sample set and a second speech test sample set, and the second speech test sample in the second speech test sample set corresponds to the second text test sample in the second text test sample set one by one; the first initial network acquisition unit may include: a third network acquisition subunit, a third speech synthesis subunit, and a first network acquisition subunit. The third network acquisition subunit may be used to iteratively train the second initial network corresponding to each language type based on the synthesized speech of each language type corresponding to the second text training sample and the second speech training sample of each language type corresponding to the second text training sample, until the third preset condition is met, and the trained second initial network corresponding to each language type is obtained as the third initial network corresponding to each language type. The third speech synthesis subunit may be used to input the second text test sample of each language type into the third initial network corresponding to each language type, respectively, to obtain the synthesized speech of each language type corresponding to the second text test sample. The first network acquisition subunit can be used to iteratively train the third initial network corresponding to each language type based on the synthesized speech of each language type corresponding to the second text test sample and the second speech test sample of each language type corresponding to the second text test sample, until the second preset condition is met, so as to obtain the trained third initial network corresponding to each language type as the first initial network corresponding to each language type.
[0116] In some embodiments, the first network acquisition subunit can be specifically used to: obtain a total loss value based on the synthesized speech of each language type corresponding to the second text test sample and the second speech test sample of each language type corresponding to the second text test sample; and iteratively train the third initial network corresponding to each language type according to the total loss value until the second preset condition is met, so as to obtain the trained third initial network corresponding to each language type as the first initial network corresponding to each language type.
[0117] In this manner, the first network acquisition subunit can also be specifically used to: determine the loss value corresponding to the third initial network of each language type based on the difference between the synthesized speech of each language type corresponding to the second text test sample and the second speech test sample of each language type corresponding to the second text test sample, to obtain multiple first loss values; and obtain the sum of the multiple first loss values as the total loss value.
[0118] In some embodiments, the model training module 640 may include: a loss value determination unit and a model training unit. The loss value determination unit may be used to determine a second loss value based on the difference between the synthesized speech corresponding to the first text sample and the first speech sample of the target language type corresponding to the first text sample. The model training unit may be used to iteratively train the initial model according to the second loss value until the first preset condition is met to obtain a trained speech synthesis model.
[0119] Please refer to Fig.10 , which shows a structural block diagram of a speech synthesis device 700 for Cantonese provided in an embodiment of the present application. The device 700 may include: a text acquisition module 710 and a speech synthesis module 720.
[0120] The text acquisition module 710 is used to acquire the text to be processed.
[0121] The speech synthesis module 720 is used to input the text to be processed into a pre-trained speech synthesis model to obtain the speech of the target language type corresponding to the text to be processed, wherein the pre-trained speech synthesis model is obtained by iteratively training the initial model using a first training sample set corresponding to the target language type until a first preset condition is met, wherein the first training sample set includes a first text sample set and a first speech sample set, wherein the first speech sample in the first speech sample set corresponds one-to-one to the first text sample in the first text sample set, and the initial model is a first initial network associated with the target language type, wherein the first initial network is trained based on a second training sample set corresponding to a plurality of language types, and the plurality of language types are associated with the target language type.
[0122] In some embodiments, the speech synthesis device 700 for Cantonese may include: a text conversion module. The text conversion module may be used to convert the text to be processed into a phoneme sequence. The speech synthesis module 720 may be specifically used to input the phoneme sequence into the pre-trained speech synthesis model to obtain a synthesized speech of the target language type corresponding to the text to be processed.
[0123] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here.
[0124] In several embodiments provided in the present application, the coupling between modules may be electrical, mechanical or other forms of coupling.
[0125] In addition, each functional module in each embodiment of the present application can be integrated into a processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The above integrated modules can be implemented in the form of hardware or software functional modules.
[0126] In summary, in the solution provided by the embodiment of the present application, a first training sample set corresponding to the target language type is obtained, the first training sample set includes a first text sample set and a first voice sample set, and the first voice sample in the first voice sample set corresponds one-to-one to the first text sample in the first text sample set; a first initial network associated with the target language type is obtained as an initial model, the first initial network is trained based on a second training sample set corresponding to a plurality of language types, and the plurality of language types are associated with the target language type; the first text sample is input into the initial model to obtain a synthesized speech corresponding to the first text sample; based on the synthesized speech corresponding to the first text sample and the first voice sample of the target language type corresponding to the first text sample, the initial model is iteratively trained until a first preset condition is met, and a trained speech synthesis model is obtained, and the speech synthesis model is used to synthesize the speech of the target language type corresponding to the text to be processed. In this way, since the first initial network trained based on the second training sample set corresponding to multiple language types has stronger learning ability, using the first initial network as the initial model to train the speech synthesis model greatly reduces the training time of the model and improves the efficiency of model training; and it is possible to train a speech synthesis model with better performance on a smaller number of first training samples, and optimize the model on a small number of first training samples, which greatly reduces the possibility of overfitting of the model, thereby improving the naturalness of the speech synthesized by the speech synthesis model finally obtained by training.
[0127] The following will be combined Fig.11 A computer device provided by the present application is described.
[0128] Reference Fig.11 , Fig.11 The structural block diagram of a computer device 800 provided in an embodiment of the present application is shown, and the above method provided in an embodiment of the present application can be executed by the computer device 800. The computer device 800 can be a device capable of running an application, such as a smart phone, a tablet computer, a smart watch, a laptop computer, a desktop computer, a server, a voice recorder, etc.
[0129] The computer device 800 in the embodiment of the present application may include one or more of the following components: a processor 801, a memory 802, and one or more applications, wherein the one or more applications may be stored in the memory 802 and configured to be executed by one or more processors 801, and the one or more programs are configured to execute the method as described in the aforementioned method embodiment.
[0130] The processor 801 may include one or more processing cores. The processor 801 uses various interfaces and lines to connect various parts of the entire computer device 800, and executes various functions and processes data of the computer device 800 by running or executing instructions, programs, code sets or instruction sets stored in the memory 802, and calling data stored in the memory 802. Optionally, the processor 801 can be implemented in at least one hardware form of digital signal processing (Digital Signal Processing, DSP), field programmable gate array (Field-Programmable Gate Array, FPGA), and programmable logic array (Programmable Logic Array, PLA). The processor 801 can integrate one or a combination of a central processing unit (Central Processing Unit, CPU), a graphics processing unit (Graphics Processing Unit, GPU) and a modem. Among them, the CPU mainly processes the operating system, user interface and application programs; the GPU is responsible for rendering and drawing display content; and the modem is used to process wireless communications. It can be understood that the above-mentioned modem can also be integrated into the processor 801 and implemented separately through a communication chip.
[0131] The memory 802 may include a random access memory (RAM) or a read-only memory (ROM). The memory 802 may be used to store instructions, programs, codes, code sets or instruction sets. The memory 802 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the following various method embodiments, etc. The data storage area may also store data (such as the various corresponding relationships described above) created by the computer device 800 during use.
[0132] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here.
[0133] In several embodiments provided in the present application, the coupling or direct coupling or communication connection between the modules shown or discussed may be an indirect coupling or communication connection through some interfaces, devices or modules, which may be electrical, mechanical or other forms.
[0134] In addition, each functional module in each embodiment of the present application can be integrated into a processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The above integrated modules can be implemented in the form of hardware or software functional modules.
[0135] Please refer to Fig.12 , which shows a structural block diagram of a computer-readable storage medium provided in an embodiment of the present application. The computer-readable medium 900 stores program codes, which can be called by a processor to execute the method described in the above method embodiment.
[0136] The computer readable storage medium 900 may be an electronic memory such as a flash memory, an EEPROM (electrically erasable programmable read-only memory), an EPROM, a hard disk, or a ROM. Optionally, the computer readable storage medium 900 includes a non-transitory computer-readable storage medium. The computer readable storage medium 900 has storage space for program code 910 that performs any method step of the above method. These program codes can be read from or written to one or more computer program products. The program code 910 can be compressed, for example, in an appropriate form.
[0137] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for training a speech synthesis model for Cantonese, characterized in that: The method comprises: Acquire a first training sample set corresponding to the target language type, wherein the first training sample set includes a first text sample set and a first voice sample set, and a first voice sample in the first voice sample set corresponds one-to-one to a first text sample in the first text sample set; Acquire a first initial network associated with the target language type as an initial model, wherein the first initial network is trained based on a second training sample set corresponding to a plurality of language types, wherein the plurality of language types are associated with the target language type; the plurality of language types being associated with the target language type means that the plurality of language types have similar pronunciation to the target language type; Inputting the first text sample into the initial model to obtain a synthesized speech corresponding to the first text sample; Based on the synthesized speech corresponding to the first text sample and the first speech sample of the target language type corresponding to the first text sample, the initial model is iteratively trained until a first preset condition is met, so as to obtain a trained speech synthesis model, which is used to synthesize the speech of the target language type corresponding to the text to be processed.
2. The method according to claim 1, characterized in that Obtaining a first initial network corresponding to the target language type as an initial model includes: Acquire a second training sample set corresponding to each language type in the multiple language types, the second training sample set comprising a second text training sample set and a second speech training sample, the second speech training sample in the second speech training sample set corresponding one-to-one to the second text training sample in the second text training sample set; Inputting the second text training samples of each language type into the second initial network corresponding to each language type respectively, to obtain the synthesized speech of each language type corresponding to the second text training samples; Based on the synthesized speech of each language type corresponding to the second text training sample and the second speech training sample of each language type corresponding to the second text training sample, iteratively train the second initial network corresponding to each language type until a second preset condition is met, and obtain the trained second initial network corresponding to each language type as the first initial network corresponding to each language type; The first initial network corresponding to any language type among the multiple language types is used as the initial model.
3. The method according to claim 2, characterized in that Before acquiring the second training sample set corresponding to each language type in the multiple language types, the method further includes: Obtain the preset voice of the preset text in each preset language type among multiple preset language types; Extracting speech features of the preset speech under each preset language type to obtain preset speech features corresponding to each preset language type; Obtaining a target voice of the preset text in the target language type; Extracting speech features of the target speech as target speech features corresponding to the target language type; The preset language types corresponding to the preset voice features whose similarity with the target voice features reaches a preset similarity threshold are obtained as the multiple language types.
4. The method according to claim 2, characterized in that: The second training sample set further includes a second text test sample set and a second voice test sample set, and the second voice test samples in the second voice test sample set correspond one-to-one to the second text test samples in the second text test sample set; The iterative training of the second initial network corresponding to each language type based on the synthesized speech of each language type corresponding to the second text training sample and the second speech training sample of each language type corresponding to the second text training sample until a second preset condition is met, to obtain the trained second initial network corresponding to each language type as the first initial network corresponding to each language type, includes: Based on the synthesized speech of each language type corresponding to the second text training sample and the second speech training sample of each language type corresponding to the second text training sample, iteratively train the second initial network corresponding to each language type until a third preset condition is met, and obtain the trained second initial network corresponding to each language type as the third initial network corresponding to each language type; Inputting the second text test sample of each language type into the third initial network corresponding to each language type respectively, to obtain the synthesized speech of each language type corresponding to the second text test sample; Based on the synthesized speech of each language type corresponding to the second text test sample and the second speech test sample of each language type corresponding to the second text test sample, the third initial network corresponding to each language type is iteratively trained until the second preset condition is met, so as to obtain the trained third initial network corresponding to each language type as the first initial network corresponding to each language type.
5. The method according to claim 4, characterized in that The iterative training of the third initial network corresponding to each language type based on the synthesized speech of each language type corresponding to the second text test sample and the second speech test sample of each language type corresponding to the second text test sample until the second preset condition is met, to obtain the trained third initial network corresponding to each language type as the first initial network corresponding to each language type, includes: Acquire a total loss value based on the synthesized speech of each language type corresponding to the second text test sample and the second speech test sample of each language type corresponding to the second text test sample; According to the total loss value, the third initial network corresponding to each language type is iteratively trained until the second preset condition is met, so as to obtain the trained third initial network corresponding to each language type as the first initial network corresponding to each language type.
6. The method according to claim 5, characterized in that The obtaining of a total loss value based on the synthesized speech of each language type corresponding to the second text test sample and the second speech test sample of each language type corresponding to the second text test sample comprises: Determine a loss value corresponding to a third initial network of each language type according to a difference between the synthesized speech of each language type corresponding to the second text test sample and the second speech test sample of each language type corresponding to the second text test sample, to obtain a plurality of first loss values; The sum of the multiple first loss values is obtained as the total loss value.
7. The method according to claim 1, characterized in that The step of obtaining a first initial network associated with the target language type as an initial model includes: From a plurality of pre-trained first initial networks, a pre-trained first initial network associated with the target language type is obtained as the initial model.
8. The method according to claim 1, characterized in that The iterative training of the initial model based on the synthesized speech corresponding to the first text sample and the first speech sample of the target language type corresponding to the first text sample until a first preset condition is met to obtain a trained speech synthesis model includes: Determine a second loss value according to a difference between the synthesized speech corresponding to the first text sample and the first speech sample of the target language type corresponding to the first text sample; According to the second loss value, the initial model is iteratively trained until the first preset condition is met, thereby obtaining a trained speech synthesis model.
9. A speech synthesis method for Cantonese, characterized in that: The method comprises: Get the text to be processed; The text to be processed is input into a pre-trained speech synthesis model to obtain a speech of a target language type corresponding to the text to be processed, wherein the pre-trained speech synthesis model is obtained by iteratively training an initial model using a first training sample set corresponding to the target language type until a first preset condition is met, wherein the first training sample set includes a first text sample set and a first speech sample set, wherein a first speech sample in the first speech sample set corresponds one-to-one to a first text sample in the first text sample set, and the initial model is a first initial network associated with the target language type, wherein the first initial network is trained based on a second training sample set corresponding to a plurality of language types, wherein the plurality of language types are associated with the target language type; wherein the plurality of language types are associated with the target language type means that the plurality of language types have similar speech pronunciations to the target language type.
10. The method according to claim 9, characterized in that After obtaining the text to be processed, the method further includes: Converting the text to be processed into a phoneme sequence; The step of inputting the text to be processed into a pre-trained speech synthesis model to obtain a synthesized speech of a target language type corresponding to the text to be processed includes: The phoneme sequence is input into the pre-trained speech synthesis model to obtain synthesized speech of the target language type corresponding to the text to be processed.
Citation Information
Patent Citations
Voice data processing method and device, equipment and medium
CN111383627A
Voice synthesis model training method and device, storage medium and electronic equipment
CN112309365A