Training Method, Device, Equipment, Storage Medium and Product of Speech Synthesis Model
By training the tone learning module in the speech synthesis model and performing transfer learning, the time-consuming and labor-intensive update of the specified tone in the prior art is solved, and the voice data of the target tone is efficiently synthesized, and the deployment efficiency and speech synthesis quality are improved.
Patent Information
- Application Number
- CN202210146068.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-17
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-02-17
AI Technical Summary
The existing voice synthesis technology needs to update all parameters when updating the voice data of the specified tone, which is time-consuming and labor-intensive and unfavorable to deployment.
By obtaining the first training data to train the tone learning module of the basic model, distinguishing the differences between different tones, and using the second training data for transfer learning, only modifying the model parameters of the tone learning module to generate the target model.
It realizes that the voice data of the target tone can be synthesized without updating all parameters, improving deployment efficiency and speech synthesis quality.
Smart Images

Figure CN114464163B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technologies, and particularly to a method, apparatus, device, storage medium, and product for training a voice synthesis model. Background Art
[0002] With the development of computer technologies, voice synthesis technologies have emerged.
[0003] Voice synthesis technologies can enable a trained model to synthesize voice data corresponding to a specified voice color, such as synthesizing voice data corresponding to the voice color of Zhang San or Li Si. If it is necessary to synthesize voice data corresponding to a newly added specified voice color, it is necessary to collect a large amount of training voice data corresponding to the newly added specified voice color to update all parameters of the previously trained model.
[0004] Updating the specified voice color of a trained model requires updating all parameters, which is time-consuming and laborious and is not conducive to deployment. Summary of the Invention
[0005] Based on this, in view of the above technical problems, it is necessary to provide a method, apparatus, device, storage medium, and product for training a voice synthesis model that does not require updating all parameters.
[0006] In a first aspect, the present application provides a method for training a voice synthesis model. The method includes: obtaining first training data; training the first training data to obtain a basic model, where the basic model includes a voice color learning module, and the voice color learning module differentiates the differences between different voice colors during the training of the first training data and obtains model parameters corresponding to the different voice colors; obtaining second training data; and performing transfer learning on the trained basic model according to the second training data to obtain a target model, where only the model parameters of the voice color learning module are modified during the transfer learning.
[0007] In one embodiment, training the first training data to obtain a basic model includes: preprocessing the first training data; encoding the preprocessed first training data to obtain corresponding feature encoding vectors, and processing the feature encoding vectors through the voice color learning module to obtain voice color features; decoding according to the feature encoding vectors and the voice color features to obtain corresponding target spectrograms; generating an objective function according to the first training data, the voice color features, and the target spectrograms; and performing iterative training according to the objective function to obtain the basic model.
[0008] In one embodiment, preprocessing the first training data includes: performing word segmentation on the text information in the first training data; and converting the segmented text information into corresponding standard phonemes.
[0009] In one embodiment, the processing of the feature encoding vector by the timbre learning module to obtain a timbre feature includes: obtaining a corresponding feature encoding vector obtained by each semantic encoding layer encoding the preprocessed first training data; generating a first timbre basis function according to each feature encoding vector, and generating a first timbre adjustment coefficient corresponding to each feature encoding vector; generating a first timbre feature according to each first timbre basis function and the first timbre adjustment coefficient.
[0010] In one embodiment, after the processing of the feature encoding vector by the timbre learning module to obtain a timbre feature, it further includes: performing length adjustment on each feature encoding vector to obtain a corresponding target encoding feature vector; the performing length adjustment on each feature encoding vector to obtain a corresponding target encoding feature vector includes: obtaining corresponding length information according to the first timbre feature; performing corresponding length adjustment on the feature encoding vector according to the length information to obtain a corresponding target encoding vector.
[0011] In one embodiment, the processing of the feature encoding vector by the timbre learning module to obtain a timbre feature includes: obtaining a corresponding target decoding vector obtained by each semantic decoding layer decoding the target encoding vector; generating a second timbre basis function according to each target decoding vector, and generating a second timbre adjustment coefficient corresponding to each target decoding vector; generating a second timbre feature according to each second timbre basis function and the second timbre adjustment coefficient.
[0012] In one embodiment, the decoding according to the feature encoding vector and the timbre feature to obtain a corresponding target spectrum includes: decoding according to the target feature decoding vector and the second timbre feature to obtain a corresponding target spectrum.
[0013] In one embodiment, the generating of the objective function according to the first training data, the timbre feature, and the target spectrum includes: obtaining the real loss value according to the difference between the real spectrum and the target spectrum, where the real spectrum is the spectrum data corresponding to the speech data; obtaining the length loss value according to the difference between the length information and the target length information, where the length information is obtained according to each first timbre feature, and the target length information is the length data corresponding to the speech data; obtaining the timbre loss value according to the first timbre feature and the second timbre feature; generating the objective function according to the real loss value, the length loss value, and the timbre loss value.
[0014] In one embodiment, obtaining a target model by performing transfer learning on the trained base model according to the second training data includes: modifying the expression of the timbre adjustment coefficient of the timbre learning module; and training according to the second training data to train the timbre adjustment coefficient with the modified expression of the timbre learning module to obtain the target model.
[0015] In a second aspect, the present application further provides a speech synthesis method, including: obtaining text information; and inputting the text information into the target model obtained by transfer learning to obtain target speech corresponding to the text information.
[0016] In a third aspect, the present application further provides a training device for a speech synthesis model. The device includes: a first acquisition module, configured to acquire first training data; a training module, configured to train the first training data to obtain a base model, where the base model includes a timbre learning module, and the timbre learning module differentiates differences in different timbres during the training of the first training data and obtains model parameters corresponding to the different timbres; a second acquisition module, configured to acquire second training data; and a second training module, configured to perform transfer learning on the trained base model according to the second training data to obtain a target model, where only the model parameters of the timbre learning module are modified during the transfer learning.
[0017] In a fourth aspect, the present application further provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method according to any one of claims 1 to 9 or 10 are implemented.
[0018] In a fifth aspect, the present application further provides a computer-readable storage medium. On the computer-readable storage medium, a computer program is stored, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 9 or 10 are implemented.
[0019] In a sixth aspect, the present application further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 9 or 10 are implemented.
[0020] The training method, device, equipment, storage medium and product of the above voice synthesis model are trained to obtain a basic model according to the acquired first training data. The timbre learning module of the basic model differentiates the differences of different timbres during the training of the first training data and obtains the model parameters corresponding to different timbres, so that the basic model can synthesize the voice data corresponding to different timbres. When it is necessary to synthesize the voice data of the target timbre, transfer learning is performed on the trained basic model according to the acquired second training data to obtain the target model. Among them, only the model parameters of the timbre learning module are modified in transfer learning, so that the target model can synthesize the voice data corresponding to the timbre in the second training data, which is beneficial to deployment. Description of the Drawings
[0021] Figure 1 It is an application environment diagram of the training method of the voice synthesis model in an embodiment;
[0022] Figure 2 It is a schematic flowchart of the training method of the voice synthesis model in an embodiment;
[0023] Figure 3 It is a schematic diagram of the architecture of the voice synthesis model in an embodiment;
[0024] Figure 4 It is a schematic diagram of the voice feature transformer in an embodiment;
[0025] Figure 5 It is a structural block diagram of the training device of the voice synthesis model in an embodiment;
[0026] Figure 6 It is a structural block diagram of the voice synthesis device in an embodiment;
[0027] Figure 7 It is an internal structure diagram of a computer device in an embodiment. Detailed Embodiments
[0028] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0029] The training method of the voice synthesis model provided by the embodiments of the present application can be applied to such as Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or placed on the cloud or other network servers. The first training data and / or the second training data are stored in the data storage system and / or the terminal 102. The server 104 obtains the first training data from the database storage system or the terminal 102 to obtain the basic model. The server 104 obtains the second training data from the database storage system or the terminal 102, and performs transfer learning on the basic model through the second training data to obtain the target model. It should be noted that the first training data and / or the second training data stored in the terminal 102 are uploaded by the user or downloaded by the terminal 102 through the network. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers.
[0030] To make it clearly understood by those skilled in the art, taking map voice synthesis as an example, the server 104 obtains the first training data from the terminal 102 or the database storage system, and trains the first training data to obtain the basic model. The server 104 obtains the second training data from the terminal 102 or the database storage system, and the second training data is data of a specific timbre; transfer learning is performed on the trained basic model according to the second training data to obtain the target model. The server 104 receives the target voice conversion instruction sent by the terminal 102, inputs the text part of the map navigation in the database storage system into the target model, and the target model converts the text part of the map navigation into the voice data corresponding to the specified timbre included in the target model. The timbre included in the above target model is the timbre included in the first training data and the second training data.
[0031] In one embodiment, as Figure 2 shown, a method for training a voice synthesis model is provided, taking the method applied to Figure 1 the terminal 102 in
[0032] Step 202, obtain the first training data.
[0033] Among them, the first training data is speech synthesis training data of at least two types, and the speech synthesis data of one type has the same timbre. The speech synthesis training data includes two parts: speech data and text information. Among them, the speech data and the text information correspond one by one. The speech data refers to the sound of the text information, and the text data is the specific text expression. For example, the speech synthesis training data of "Hello" can include the sound data of "Hello" and the text "Hello". Optionally, the speech synthesis training data of each person includes multiple speech data, and the length of each speech data meets the length condition, for example, less than 10s. Preferably, the total duration of all the speech data in the speech synthesis training data of each person meets the time condition, for example, more than 10h. It should be noted that the above length condition and time condition can be set as needed and will not be specifically limited here.
[0034] Specifically, the terminal can read at least two types of speech synthesis training data from the database. In one embodiment, the terminal can read the speech synthesis training data of different users from the database respectively to meet the requirements.
[0035] Step 204: Train the first training data to obtain a basic model. The basic model includes a timbre learning module. The timbre learning module distinguishes the differences of different timbres during the training of the first training data and obtains the model parameters corresponding to different timbres.
[0036] Among them, the basic model is a model that can obtain speech data corresponding to different timbres according to text information by training the first training data. That is, by inputting text information, speech data of different timbres can be obtained. The basic model can be the base speech synthesis system in the fune-tune (transfer learning) speech synthesis scheme. Specifically, the basic model can include a phoneme generation network, a feature extraction network, a length adjustment network, a spectrum generation network, and a sound generation network. The phoneme generation network is used to convert the text information in the first training data into corresponding standard phonemes. The feature extraction network is used to extract the timbre features of the speech data in the first training data. The length adjustment network is used to convert the length of the standard phonemes into the phoneme lengths corresponding to the timbres of the speech data in the first training data. The spectrum generation network is used to convert the adjusted standard phonemes into corresponding spectrum data. The sound generation network is used to convert the spectrum data into sound data. In this way, the basic model can convert text information into sound data of the corresponding timbre. The basic model can synthesize the speech data corresponding to the timbre in the first training data.
[0037] Specifically, the terminal trains on multiple types of speech synthesis training data to obtain a basic model that can synthesize speech data corresponding to different timbres in the first training data. During the training process, the timbre learning module of the basic model is trained based on multiple types of speech synthesis training data to distinguish the differences between the timbres corresponding to different types of speech synthesis data, and obtain the model parameters of the timbres corresponding to different types of speech synthesis data.
[0038] Step 206, obtain the second training data.
[0039] Among them, the second training data is the speech synthesis training data of the target type. The timbre of the speech synthesis training data of the target type is different from the timbres in the first training data, and the timbre of the speech synthesis training data of the target type is the same timbre. Optionally, the speech synthesis training data of the target type includes multiple speech data, and the length of each speech data satisfies the second length condition, for example, less than 10s. Preferably, the total number of all speech data meets the quantity condition, for example, at least 20. It should be noted that the speech data in the speech synthesis training data of the target type also corresponds one-to-one with the text information. It should be noted that the above second length condition and quantity condition can be set as needed and are not specifically limited here.
[0040] Specifically, the terminal can read the speech synthesis training data of the target type from the database. In one embodiment, the terminal can read the speech synthesis training data of the target user from the database to meet the requirements.
[0041] Obtain the speech synthesis training data of the target type, that is, the second training data.
[0042] Step 208, perform transfer learning on the trained basic model according to the second training data to obtain a target model, where only the model parameters of the timbre learning module are modified during the transfer learning.
[0043] Among them, the target model is a model that can obtain speech data of the target timbre according to the text information by training on the second training data, and the speech data of the target timbre can be obtained by inputting the text information. The target model can be a model after transfer learning in the fine-tune speech synthesis scheme. Specifically, the target model only modifies the feature extraction network of the above basic model, and the modified feature extraction network is used to extract the target timbre features of the speech data in the second training data, so that the target model can convert the text information into the corresponding sound data of the target timbre. The target model can synthesize the speech data corresponding to the target timbre in the second training data.
[0044] It can synthesize the speech data corresponding to the target timbre corresponding to the second training data.
[0045] Specifically, the terminal performs transfer learning on the speech synthesis training data of the target to obtain a target model that can synthesize the speech data corresponding to the target timbre in the speech synthesis training data of the target. During the transfer learning process, the terminal only learns the target model parameters of the timbre learning module in the target model based on the speech synthesis training data of the target, and modifies the model parameters of the timbre learning module of the target model with the target model parameters.
[0046] In the above training method of the speech synthesis model, a basic model is trained according to the obtained first training data. The timbre learning module of the basic model differentiates the differences of different timbres during the training of the first training data, and obtains the model parameters corresponding to different timbres, so that the basic model can synthesize the speech data corresponding to different timbres. When it is necessary to synthesize the speech data of the target timbre, transfer learning is performed on the trained basic model according to the obtained second training data to obtain a target model. Among them, transfer learning only modifies the model parameters of the timbre learning module, so that the target model can synthesize the speech data corresponding to the timbre in the second training data, which is beneficial to deployment.
[0047] In one embodiment, training the first training data to obtain a basic model includes: preprocessing the first training data. Encoding the preprocessed first training data to obtain a corresponding feature encoding vector, and processing the feature encoding vector through a timbre learning module to obtain a timbre feature. Decoding according to the feature encoding vector and the timbre feature to obtain a corresponding target spectrum. Generating an objective function according to the first training data, the timbre feature and the target spectrum. Specifically, the terminal preprocesses the sentences in the text information of the first training data to obtain corresponding phonemes. For example, the sentence "I love my motherland" in the first training data is converted into phonemes as "w", "ou", "ai", "zu", "g", "uo".
[0048] Among them, different people have different timbre characteristics.
[0049] Specifically, the terminal encodes the phonemes obtained according to the preprocessing to obtain a corresponding feature encoding vector, and this feature encoding vector is used to represent the same features of the phonemes obtained by the preprocessing; and processes the feature encoding vector through a phoneme learning module to obtain different features of the speech data corresponding to the phonemes obtained by the preprocessing, that is, timbre features.
[0050] Specifically, the terminal decodes the feature encoding vector to obtain the same features of each phoneme, and decodes the feature encoding vector according to the timbre feature to obtain the different features of the speech data corresponding to each phoneme, that is, the timbre feature, so as to obtain the encoding vector of the speech data corresponding to each phoneme, and decode the encoding vector to obtain the corresponding target spectrum. Optionally, the target spectrum is a Mel spectrum. This embodiment does not limit the type of the target spectrum, as long as it can be used to represent the features of the sound.
[0051] Among them, the objective function is a function used to optimize the basic model.
[0052] Specifically, the terminal generates an objective function according to the first training data, the timbre feature, and the target spectrum.
[0053] According to the objective function, iterative training is performed to obtain the basic model. In the above training method of the speech synthesis model, by preprocessing the first training data, the granularity of the synthesis modeling is reduced, and the quality of the speech synthesis is improved; the timbre learning module is used to learn the different timbre features of each timbre in the first training data; an objective function is generated according to the first training data, the timbre feature, and the target spectrum, so as to perform iterative training to obtain the basic model.
[0054] In one embodiment, preprocessing the first training data includes: performing word segmentation on the text information in the first training data; converting the segmented text information into corresponding standard phonemes.
[0055] Among them, the first training data includes speech data and text information, and the speech data and the text information are in one-to-one correspondence. The standard phoneme is a phoneme with the same timbre, and the phonemes corresponding to the same word have the same timbre.
[0056] Specifically, the terminal performs word segmentation on the sentences in the text information of the first training data to obtain the corresponding words, and converts the obtained words after word segmentation into standard phonemes. For example, the sentence in the first training data is "I love my country", the corresponding words obtained after word segmentation are "I", "love", "my country", and further converted into phonemes, namely "w", "ou", "ai", "zu", "g", "uo".
[0057] In the above training method of the speech synthesis model, the granularity of the speech synthesis model modeling is reduced, the problem of uneven text data is reduced, and the quality of the speech synthesis is improved.
[0058] In one embodiment, the timbre learning module processes the feature encoding vector to obtain the timbre feature, including: obtaining the corresponding feature encoding vector by encoding the preprocessed first training data through each semantic encoding layer; generating the first timbre basis function according to each feature encoding vector, and generating the first timbre adjustment coefficient corresponding to each feature encoding vector; generating the first timbre feature according to each first timbre basis function and the first timbre adjustment coefficient.
[0059] Specifically, the terminal encodes the standard phoneme through each semantic encoding layer to obtain the corresponding feature encoding vector, that is, extracts the same features of the standard phoneme through each semantic encoding layer to obtain the corresponding feature encoding vector, which is x in formula (1). The terminal generates the first timbre basis function β i (x) and the corresponding first timbre adjustment coefficient
[0060] In the specific implementation process, the timbre learning module processes the feature encoding vector x through formula (1) to obtain the timbre feature, and formula (1) is:
[0061]
[0062] Among them, l in formula (1) represents the id of the network layer, and x is the feature encoding vector corresponding to each semantic encoding layer. t represents the id of different timbres, that is, different types of speech synthesis training data in the first training data correspond to different timbres and have different ids, and each different timbre has corresponding adaptation parameters. β i (x) is the adaptive basis function of the above adaptation parameters, a total of M, and different β i (x) functions are orthogonal. represents the adaptive coefficient of the β i (x) function corresponding to the t-th person, is a learnable coefficient. Through the voice data in different types of speech synthesis training data in the first training data, the different coefficients of each layer of each different person can be determined, so as to ensure that each person has different adaptive coefficients for different layers.
[0063] In the above training method of the speech synthesis model, through the extraction of the feature encoding vector, the same features corresponding to the above standard phoneme are obtained. The first timbre basis function and the corresponding first timbre adjustment coefficient are generated through each feature encoding vector; the first timbre feature is generated according to each first timbre basis function and the first timbre adjustment coefficient to obtain the different features of different timbres in each first training data corresponding to the above standard phoneme.
[0064] In one embodiment, after the timbre learning module processes the feature encoding vectors to obtain timbre features, it further includes: adjusting the lengths of the feature encoding vectors to obtain corresponding target encoding feature vectors; adjusting the lengths of the feature encoding vectors to obtain corresponding target encoding feature vectors, including: obtaining corresponding length information according to the first timbre feature; performing corresponding length adjustment on the feature encoding vectors according to the length information to obtain corresponding target encoding vectors.
[0065] Specifically, the terminal adjusts the lengths of the feature encoding vectors according to the lengths of the feature vectors corresponding to different types of timbres to obtain target encoding feature vectors.
[0066] Specifically, the terminal obtains the length information of the feature vectors corresponding to different types of timbres according to the first timbre feature, and performs corresponding length adjustment on the feature encoding vectors according to the length information to obtain corresponding target encoding vectors. For example, if the terminal receives 4 feature encoding vectors and parses the lengths of the above feature encoding vectors as "1", "2", "3", and "4" according to the first timbre feature, then the terminal copies the corresponding number of copies of the 4 feature encoding vectors according to the above length information to obtain corresponding target encoding vectors.
[0067] In the above method for training a speech synthesis model, length information is obtained through the first timbre feature; corresponding parallel length adjustment is performed on the feature encoding vectors according to the length information to obtain corresponding target encoding vectors, which greatly increases the speed of speech synthesis.
[0068] In one embodiment, processing the feature encoding vectors by the timbre learning module to obtain timbre features includes: obtaining, by each semantic decoding layer, a corresponding target decoding vector by decoding the target encoding vector; generating a second timbre basis function according to each target decoding vector, and generating a second timbre adjustment coefficient corresponding to each target decoding vector; generating a second timbre feature according to each second timbre basis function and the second timbre adjustment coefficient.
[0069] Specifically, the terminal obtains a corresponding target decoding vector by decoding the target encoding vector through each semantic decoding layer, that is, extracts the same features of the target encoding vector through each semantic decoding layer to obtain the corresponding target decoding vector, which is x in formula (1). The terminal generates a second timbre basis function β i (x) and the corresponding second timbre adjustment coefficient
[0070] In a specific implementation process, the timbre learning module processes the target decoding vector x through formula (1) to obtain a second timbre feature.
[0071] Among them, l in formula (1) represents the id of the network layer, x is the target decoding vector corresponding to each semantic decoding layer. t represents the id of different timbres, that is, different types of speech synthesis training data in the first training data correspond to different timbres and have different ids, and each different timbre has corresponding adaptation parameters. β i (x) is the adaptive basis function of the above adaptation parameters, a total of M, and different β i (x) functions are orthogonal. represents the adaptive coefficient of the function corresponding to the t-th person, which is a learnable coefficient. Through the speech data in different types of speech synthesis training data in the first training data, different coefficients of each layer of each different person can be determined, so as to ensure that each person has different adaptive coefficients for different layers.
[0072] It should be noted that x in formula (1) is the feature encoding vector and the target encoding vector, and the feature encoding vector, the target encoding vector, and the target decoding vector are of the same type of data. The number of M is the sum of the encoding layers and the decoding layers.
[0073] In the above training method of the speech synthesis model, through the extraction of the target decoding vector, the same features corresponding to the above target encoding vector are obtained. A second timbre basis function and a corresponding second timbre adjustment coefficient are generated through each target decoding vector; according to each second timbre basis function and the second timbre adjustment coefficient, a second timbre feature is generated to obtain different features of different timbres in each of the above first training data corresponding to the above target decoding vector.
[0074] In one embodiment, decoding according to the feature encoding vector and the timbre feature to obtain a corresponding target spectrum includes: decoding according to the target feature decoding vector and the second timbre feature to obtain a corresponding target spectrum.
[0075] Specifically, the terminal decodes according to the target feature decoding vector to obtain the Mel spectrum corresponding to the same feature, and decodes according to the second timbre feature to obtain the Mel spectrum corresponding to different features of each timbre. The sum of the Mel spectra corresponding to the same feature and different features is the target Mel spectrum.
[0076] The above training method of the speech synthesis model obtains the target spectrum data required in the speech synthesis step.
[0077] In one embodiment, generating an objective function based on first training data, timbre features, and a target spectrum includes: obtaining a true loss value based on the difference between a true spectrum and the target spectrum, where the true spectrum is the spectral data corresponding to the speech data. Obtaining a length loss value based on the difference between length information and target length information, where the length information is obtained from each feature encoding vector, and the target length information is the length data corresponding to the speech data. Obtaining a timbre loss value based on the first timbre feature and the second timbre feature. Generating an objective function based on the true loss value, the length loss value, and the timbre loss value.
[0078] Among them, the true spectrum is the spectral data corresponding to the speech data in the first training data.
[0079] Specifically, obtaining a true loss value based on the difference between the terminal true spectrum and the target spectrum, where the true spectrum is the spectral data corresponding to the speech data in the first training data, and the target spectrum is the spectral data obtained according to the training steps in the above embodiments corresponding to the above speech data. The above true loss value L1 can be obtained according to formula (2).
[0080] L1 = ||R - O|| (2)
[0081] Among them, R in formula (2) is the true spectrum matrix corresponding to the true spectrum, and O is the target spectrum matrix corresponding to the target spectrum. Optionally, if the true spectrum is a Mel spectrum, the types of the true spectrum matrix and the target spectrum matrix are Mel spectrum matrices.
[0082] Specifically, obtaining a length loss value based on the difference between the length information and the target length information, where the length information is obtained from each first timbre feature, the target length information is the length data corresponding to the speech data, and each first timbre feature is learned from the speech data and corresponds to each other. The above length loss value L2 can be obtained according to formula (3).
[0083]
[0084] Among them, p in formula (3) is the length of the speech phoneme, r i is the length of the speech Mel spectrum of the i-th phoneme in the speech data, o i is the length of the speech Mel spectrum of the i-th phoneme obtained from the feature encoding vector. Optionally, the length of the speech Mel spectrum of each phoneme in the speech data can be obtained by obtaining alignment information through an MFA tool.
[0085] Specifically, the calculation formula of the timbre loss value l3 is as shown in formula (4).
[0086]
[0087] Among them, L in formula (4) is the network layer id, al is the adaptive matrix of the l-th layer, and l3 is the adaptive loss function for measuring different speakers, which reflects that the adaptive modules of speakers with the same timbre are similar, while those of speakers with different timbres are different. Among them, the first timbre feature and the second timbre feature respectively represent the adaptive matrices under different network layers.
[0088] Specifically, the objective function L is obtained through formula (5).
[0089] L = L1 + L2 + L3 (5)
[0090] In the above training method of the voice synthesis model, the loss between the target spectrum and the corresponding real spectrogram is considered through the real loss value, the loss between the predicted length information and the corresponding real length information is considered through the length loss value, and the adaptive loss function of different speakers is considered through the timbre loss value.
[0091] In one embodiment, performing transfer learning on the trained basic model according to the second training data to obtain the target model includes: modifying the expression of the timbre adjustment coefficient of the timbre learning module. Training according to the second training data to train the modified expression of the timbre adjustment coefficient of the timbre learning module to obtain the target model.
[0092] Specifically, the terminal modifies the in formula (1) to obtain the target model.
[0093] Specifically, the terminal trains according to the voice data in the voice synthesis training data of the target category to obtain different features of the timbre corresponding to the voice data of the target category, that is, modifies the of the basic model to the corresponding to the target timbre, that is, the target model is obtained.
[0094] In the above training method of the voice synthesis model, training the modified expression of the timbre adjustment coefficient of the timbre learning module with a small amount of second training data, and modifying the expression of the timbre adjustment coefficient of the timbre learning module in the basic model can obtain the target model. Only modifying one parameter of the basic model can obtain the target model that can synthesize voice data of the target timbre. The number of modified parameters is small, which is convenient for deployment and reuse. And the second training data for training the target model is small and easy to obtain.
[0095] In one embodiment, a voice synthesis method includes: obtaining text information; inputting the text information into the target model obtained by transfer learning to obtain the target voice corresponding to the text information.
[0096] In a specific embodiment, a training method of a voice synthesis model includes: a terminal, a server, and a data storage system.
[0097] The first training data is speech synthesis training data of at least two different persons. Among them, the speech synthesis training data of each person includes a plurality of speech data and text information, the speech data and the text information are in one-to-one correspondence, the length of each speech data is less than 10s, and the sum of the speech data of each person is at least 10h. Optionally, the speech data may be a recording file corresponding to the text information.
[0098] The second training data is speech synthesis training data of a customized tone. The speech synthesis training data of the customized tone also includes a plurality of speech data and text information, the speech data and the text information are in one-to-one correspondence, the length of each speech data is less than 10s, and at least 20 pieces of speech data of the customized tone are required. It should be noted that the tones of the speech data in the speech synthesis training data of the customized tone are the same, which is the same tone. Optionally, the speech data of the customized tone may be downloaded offline or uploaded by the user.
[0099] The training process of the speech synthesis model is as Figure 3 shown. The server obtains the first training data and inputs the text information in the first training data as a text sequence into the training model.
[0100] The terminal performs word segmentation processing on the text sequence through the word preprocessing module of the training model, and performs standard phoneme conversion on each word after the word segmentation processing. For example: the text sequence is I love my motherland -> I love my motherland -> w ou ai zu g uo. Among them, the standard phoneme can be understood as the printed Chinese character. The shapes of the same printed Chinese character are the same, that is, the pronunciation of the same standard phoneme is always the same and does not change with the target tone. The word preprocessing module reduces the modeling granularity of the speech synthesis model, reduces the problem of uneven text data, and can improve the quality of speech synthesis.
[0101] The terminal uses the output standard phoneme of the word preprocessing module as the input of the semantic information feature encoding module. Each encoding layer of the semantic feature encoding module is used to extract the same feature x corresponding to each standard phoneme, and uses the same feature x of each standard phoneme extracted by the above encoding layers as the input of the adaptive activation module. The adaptive activation module learns the different features of each tone in the first training data according to formula (1), and uses the above different features as the output of the adaptive activation module. The above adaptive network module exists in each layer of the neural network of each semantic information encoding module. Among them, the same feature x is the feature encoding vector in the above embodiment, and the different features of each tone and the first tone feature in the above embodiment. Optionally, the semantic information feature encoding module is implemented by a transformer (a neural network) network or a Bi-LSTM (bidirectional long short-term memory neural network) network.
[0102] The terminal uses the same feature x output by the voice information feature encoding module and the different features of each timbre output by the adaptive activation module as the input of the voice feature transformer module. The voice feature transformer module includes a length predictor and a length adjuster. The terminal uses the different features of each timbre output by the adaptive activation module as the input of the length predictor to obtain the length of the same feature x output by the voice information feature encoding module. The length adjuster adjusts the length of the same feature x according to the output of the length predictor to obtain the target encoding vector. As Figure 4 shown, the input is the encoded feature vectors of four phonemes, and the predictor expects their lengths to be 1, 2, 3, and 4 respectively. Then the length adjuster copies the encoded feature vectors 1, 2, 3, and 4 times respectively. Through the voice feature transformer module, parallel operation of voice synthesis can be carried out, greatly increasing the speed of voice synthesis.
[0103] The terminal uses the output target encoding vector of the voice feature transformer module as the input of the semantic information feature decoding module. Each decoding layer of the semantic feature decoding module is used to extract the same feature x corresponding to each target encoding vector, and uses the same feature x of each target encoding vector extracted by the above-mentioned decoding layers as the input of the adaptive activation module. The adaptive activation module learns the different features of each timbre in the first training data according to formula (1), and uses the above different features as the output of the adaptive activation module. The terminal synthesizes the corresponding Mel spectrogram according to the same feature x corresponding to each target encoding vector and the different features of each timbre. The above adaptive network module exists in each layer of the neural network of each semantic information decoding module. Among them, the same feature x is the target decoding vector in the above embodiment, the different features of each timbre and the second timbre feature in the above embodiment. Optionally, the semantic information feature encoding module is implemented by a transformer (a neural network) network or a Bi-LSTM (bidirectional long short-term memory neural network) network.
[0104] The terminal uses the output Mel spectrogram of the voice information decoding module as the input of the voice vocoder module, and converts the above Mel spectrogram into a target voice waveform diagram perceptible to the human ear through the voice vocoder module. Optionally, the voice vocoder module is a Hi-Fi GAN (high-fidelity vocoder) vocoder or a MelGAN (generative adversarial network for fast audio generation) vocoder.
[0105] The terminal uses formula (5) as the text information for inputting the first training data into the training model, and outputs the loss function of the training process of the Mel spectrogram. This loss function includes three parts, namely (2), (3), and (4). Among them, formula (2) is the loss between the predicted Mel spectrogram and the real Mel spectrogram, which can be designed as L1-loss. R represents the real Mel spectrogram matrix, and O is the network-predicted Mel spectrogram matrix. Formula (3) is the loss function of the length predictor, where p is the length of the voice phoneme, ri is the length of the speech Mel spectrogram of the i-th phoneme, o i is the length of the speech Mel spectrogram of the i-th phoneme predicted by the phoneme length predictor. Formula (4) is an adaptive loss function for measuring different speakers, which reflects that the adaptive modules of speakers with the same timbre are similar, and the adaptive modules of speakers with different timbres are different, where L is the network layer id, a l is the adaptive matrix of the l-th layer. Optionally, the terminal can pre-emphasize, frame, window, perform fast Fourier transform, and Mel spectrogram transform on the original waveform of the speech data in the first training data to obtain the real Mel spectrogram matrix of R. Optionally, the terminal obtains r i alignment information of
[0106] After the model training for generating the Mel spectrogram is completed, the terminal uses it to generate a predicted Mel spectrogram, and trains the vocoder module according to the one-to-one correspondence between the predicted Mel spectrogram and the real speech waveform diagram. The corresponding basic model is obtained according to the model for generating the Mel spectrogram after training and the vocoder module after training.
[0107] Based on the basic model after training, the adaptive coefficient of formula (1) for training the basic model is updated using the speech data of the target timbre in the second training data to obtain formula (6).
[0108] Updating formula (1) in the basic model to formula (6) is the target model.
[0109]
[0110] Among them, the loss function Y corresponding to formula (6) is formula (7).
[0111] Y = L1 + L2 (7)
[0112] It should be noted that the first timbre feature and the second timbre feature, formula (1) and formula (6) all refer to model parameters.
[0113] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the indications of the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0114] Based on the same inventive concept, an embodiment of the present application further provides a training device for a speech synthesis model for implementing the training method of the speech synthesis model involved above. The solution provided by this device for solving problems is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the training device for the speech synthesis model provided below can refer to the limitations on the training method of the speech synthesis model in the above text, and will not be repeated here.
[0115] In one embodiment, as Figure 5 shown, a training device for a speech synthesis model is provided, including: a first acquisition module 100, a training module 200, a second acquisition module 300, and a second training module 400, where:
[0116] The first acquisition module 100 is configured to acquire first training data.
[0117] The training module 200 is configured to train the first training data to obtain a basic model. The basic model includes a timbre learning module. The timbre learning module differentiates the differences of different timbres during the training of the first training data and obtains the model parameters corresponding to different timbres.
[0118] The second acquisition module 300 is configured to acquire second training data.
[0119] The second training module 400 is configured to perform transfer learning on the trained basic model according to the second training data to obtain a target model, where only the model parameters of the timbre learning module are modified during the transfer learning.
[0120] In one embodiment, the training module includes: a preprocessing module for preprocessing the first training data; an encoding module for encoding the preprocessed first training data to obtain corresponding feature encoding vectors, and processing the feature encoding vectors through a timbre learning module to obtain timbre features; a decoding module for decoding according to the feature encoding vectors and the timbre features to obtain corresponding target spectra; a target function generation module for generating a target function according to the first training data, the timbre features, and the target spectra; and a training module for performing iterative training according to the target function to obtain a basic model.
[0121] In one embodiment, the preprocessing module includes: a word segmentation module for performing word segmentation on the text information in the first training data; and a phoneme conversion module for converting the segmented text information into corresponding standard phonemes.
[0122] In one embodiment, the encoding module includes: a first encoding module for obtaining, by each semantic encoding layer, corresponding feature encoding vectors by encoding the preprocessed first training data; a first generation module for generating a first timbre basis function according to each feature encoding vector and generating a first timbre adjustment coefficient corresponding to each feature encoding vector; and a first feature generation module for generating a first timbre feature according to each first timbre basis function and the first timbre adjustment coefficient.
[0123] In one embodiment, it further includes: a length adjustment module for adjusting the lengths of the feature encoding vectors to obtain corresponding target encoding feature vectors. Adjusting the lengths of the feature encoding vectors to obtain corresponding target encoding feature vectors includes: a length obtaining module for obtaining corresponding length information according to the first timbre feature; and a target encoding vector generation module for performing corresponding length adjustment on the feature encoding vectors according to the length information to obtain corresponding target encoding vectors.
[0124] In one embodiment, the encoding module includes: a first decoding module for obtaining, by each semantic decoding layer, corresponding target decoding vectors by decoding the target encoding vectors; a second generation module for generating a second timbre basis function according to each target decoding vector and generating a second timbre adjustment coefficient corresponding to each target decoding vector; and a second feature generation module for generating a second timbre feature according to each second timbre basis function and the second timbre adjustment coefficient.
[0125] In one embodiment, the decoding module includes: a target spectrum generation module for decoding according to the target feature decoding vectors and the second timbre features to obtain corresponding target spectra.
[0126] In one embodiment, the objective function generation module includes: a true loss value generation module for obtaining a true loss value based on the difference between the true spectrum and the target spectrum, where the true spectrum is the spectral data corresponding to the voice data; a length loss value generation module for obtaining a length loss value based on the difference between the length information and the target length information, where the length information is obtained from each first timbre feature and the target length information is the length data corresponding to the voice data; a timbre loss value generation module for obtaining a timbre loss value based on the first timbre feature and the second timbre feature; and a second objective function generation module for generating an objective function based on the true loss value, the length loss value, and the timbre loss value.
[0127] In one embodiment, the second training module includes: a modification module for modifying the expression of the timbre adjustment coefficient of the timbre learning module; and a target model generation module for training based on the second training data to train the timbre adjustment coefficient with the modified expression of the timbre learning module to obtain a target model.
[0128] In one embodiment, as Figure 6 shown, a voice synthesis device is provided, including: an information acquisition module 500 and a target voice generation module 600, where
[0129] The information acquisition module 500 is configured to acquire text information.
[0130] The target voice generation module 600 is configured to input the text information into the target model obtained by transfer learning to obtain the target voice corresponding to the text information.
[0131] Each module in the above voice synthesis model training device and the voice synthesis device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in hardware form or be independent of it, or can be stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above respective modules.
[0132] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 7As shown in the figure. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a method for training a voice synthesis model and a voice synthesis method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0133] Those skilled in the art can understand that Figure 7 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0134] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented: obtaining first training data; training the first training data to obtain a basic model. The basic model includes a timbre learning module, and the timbre learning module differentiates the differences of different timbres during the training of the first training data and obtains the model parameters corresponding to different timbres; obtaining second training data; performing transfer learning on the trained basic model according to the second training data to obtain a target model, where only the model parameters of the timbre learning module are modified during the transfer learning.
[0135] In one embodiment, the training of the first training data to obtain a basic model implemented when the processor executes the computer program includes: preprocessing the first training data; encoding the preprocessed first training data to obtain a corresponding feature encoding vector, and processing the feature encoding vector through the timbre learning module to obtain timbre features; decoding according to the feature encoding vector and the timbre features to obtain a corresponding target spectrum; generating an objective function according to the first training data, the timbre features, and the target spectrum; and performing iterative training according to the objective function to obtain a basic model.
[0136] In one embodiment, the preprocessing of the first training data implemented when the processor executes a computer program includes: performing word segmentation on the text information in the first training data; converting the word-segmented text information into corresponding standard phonemes.
[0137] In one embodiment, the processing of the feature encoding vectors by the timbre learning module to obtain timbre features implemented when the processor executes a computer program includes: obtaining the corresponding feature encoding vectors obtained by each semantic encoding layer encoding the preprocessed first training data; generating a first timbre basis function according to each feature encoding vector, and generating a first timbre adjustment coefficient corresponding to each feature encoding vector; generating a first timbre feature according to each first timbre basis function and the first timbre adjustment coefficient.
[0138] In one embodiment, after the processing of the feature encoding vectors by the timbre learning module to obtain timbre features implemented when the processor executes a computer program, it further includes: performing length adjustment on each feature encoding vector to obtain the corresponding target encoding feature vector; performing length adjustment on each feature encoding vector to obtain the corresponding target encoding feature vector, including: obtaining the corresponding length information according to the first timbre feature; performing corresponding length adjustment on the feature encoding vector according to the length information to obtain the corresponding target encoding vector.
[0139] In one embodiment, the processing of the feature encoding vectors by the timbre learning module to obtain timbre features implemented when the processor executes a computer program includes: obtaining the corresponding target decoding vectors obtained by each semantic decoding layer decoding the target encoding vector; generating a second timbre basis function according to each target decoding vector, and generating a second timbre adjustment coefficient corresponding to each target decoding vector; generating a second timbre feature according to each second timbre basis function and the second timbre adjustment coefficient.
[0140] In one embodiment, decoding according to the feature encoding vector and the timbre feature to obtain the corresponding target spectrum implemented when the processor executes a computer program includes: decoding according to the target feature decoding vector and the second timbre feature to obtain the corresponding target spectrum.
[0141] In one embodiment, generating an objective function according to the first training data, the timbre feature, and the target spectrum implemented when the processor executes a computer program includes: obtaining a real loss value according to the difference between the real spectrum and the target spectrum, where the real spectrum is the spectrum data corresponding to the voice data; obtaining a length loss value according to the difference between the length information and the target length information, where the length information is obtained according to each first timbre feature, and the target length information is the length data corresponding to the voice data; obtaining a timbre loss value according to the first timbre feature and the second timbre feature; generating an objective function according to the real loss value, the length loss value, and the timbre loss value.
[0142] In one embodiment, when the processor executes a computer program, performing transfer learning on a trained base model according to second training data to obtain a target model includes: modifying the expression of the timbre adjustment coefficient of the timbre learning module; and training according to the second training data to train the modified expression of the timbre adjustment coefficient of the timbre learning module to obtain the target model. In one embodiment, a speech synthesis method implemented when the processor executes a computer program includes: obtaining text information; and inputting the text information into the target model obtained by transfer learning to obtain target speech corresponding to the text information.
[0143] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: obtaining first training data; training the first training data to obtain a base model, the base model including a timbre learning module, the timbre learning module differentiating the differences of different timbres during the training of the first training data and obtaining model parameters corresponding to different timbres; obtaining second training data; and performing transfer learning on the trained base model according to the second training data to obtain a target model, wherein only the model parameters of the timbre learning module are modified during the transfer learning.
[0144] In one embodiment, training the first training data to obtain a base model implemented when the computer program is executed by a processor includes: preprocessing the first training data; encoding the preprocessed first training data to obtain corresponding feature encoding vectors, and processing the feature encoding vectors through a timbre learning module to obtain timbre features; decoding according to the feature encoding vectors and the timbre features to obtain corresponding target spectrograms; generating an objective function according to the first training data, the timbre features, and the target spectrograms; and performing iterative training according to the objective function to obtain the base model.
[0145] In one embodiment, preprocessing the first training data implemented when the computer program is executed by a processor includes: performing word segmentation on the text information in the first training data; and converting the segmented text information into corresponding standard phonemes.
[0146] In one embodiment, processing the feature encoding vectors through a timbre learning module to obtain timbre features implemented when the computer program is executed by a processor includes: obtaining corresponding feature encoding vectors encoded by each semantic encoding layer for the preprocessed first training data; generating first timbre basis functions according to each feature encoding vector, and generating first timbre adjustment coefficients corresponding to each feature encoding vector; and generating first timbre features according to each first timbre basis function and the first timbre adjustment coefficients.
[0147] In one embodiment, after the tone learning module processes the feature encoding vector to obtain the tone feature when the computer program is executed by the processor, it further includes: adjusting the length of each feature encoding vector to obtain the corresponding target encoding feature vector; adjusting the length of each feature encoding vector to obtain the corresponding target encoding feature vector, including: obtaining the corresponding length information according to the first tone feature; performing corresponding length adjustment on the feature encoding vector according to the length information to obtain the corresponding target encoding vector.
[0148] In one embodiment, when the computer program is executed by the processor, the process of obtaining the tone feature by processing the feature encoding vector through the tone learning module includes: obtaining each target decoding vector obtained by decoding the target encoding vector by each semantic decoding layer; generating the second tone basis function according to each target decoding vector, and generating the second tone adjustment coefficient corresponding to each target decoding vector; generating the second tone feature according to each second tone basis function and the second tone adjustment coefficient.
[0149] In one embodiment, when the computer program is executed by the processor, decoding according to the feature encoding vector and the tone feature to obtain the corresponding target spectrum includes: decoding according to the target feature decoding vector and the second tone feature to obtain the corresponding target spectrum.
[0150] In one embodiment, when the computer program is executed by the processor, generating the objective function according to the first training data, the tone feature and the target spectrum includes: obtaining the real loss value according to the difference between the real spectrum and the target spectrum, where the real spectrum is the spectrum data corresponding to the speech data; obtaining the length loss value according to the difference between the length information and the target length information, where the length information is obtained according to each first tone feature and the target length information is the length data corresponding to the speech data; obtaining the tone loss value according to the first tone feature and the second tone feature; generating the objective function according to the real loss value, the length loss value and the tone loss value.
[0151] In one embodiment, when the computer program is executed by the processor, performing transfer learning on the trained basic model according to the second training data to obtain the target model includes: modifying the expression of the tone adjustment coefficient of the tone learning module; training according to the second training data to train the modified expression of the tone adjustment coefficient of the tone learning module to obtain the target model.
[0152] In one embodiment, a speech synthesis method implemented when the computer program is executed by the processor includes: obtaining text information; inputting the text information into the target model obtained by transfer learning to obtain the target speech corresponding to the text information.
[0153] In one embodiment, a computer program product is provided, including a computer program, which when executed by a processor implements the following steps: obtaining first training data; training the first training data to obtain a basic model, the basic model including a timbre learning module, the timbre learning module differentiating the differences of different timbres during the training of the first training data and obtaining model parameters corresponding to different timbres; obtaining second training data; performing transfer learning on the trained basic model according to the second training data to obtain a target model, wherein only the model parameters of the timbre learning module are modified during the transfer learning.
[0154] In one embodiment, the training of the first training data to obtain a basic model when the computer program is executed by a processor includes: preprocessing the first training data; encoding the preprocessed first training data to obtain corresponding feature encoding vectors, and processing the feature encoding vectors through a timbre learning module to obtain timbre features; decoding according to the feature encoding vectors and the timbre features to obtain corresponding target spectra; generating an objective function according to the first training data, the timbre features and the target spectra; and performing iterative training according to the objective function to obtain a basic model.
[0155] In one embodiment, the preprocessing of the first training data when the computer program is executed by a processor includes: performing word segmentation on the text information in the first training data; and converting the segmented text information into corresponding standard phonemes.
[0156] In one embodiment, the processing of the feature encoding vectors through a timbre learning module to obtain timbre features when the computer program is executed by a processor includes: obtaining corresponding feature encoding vectors obtained by each semantic encoding layer encoding the preprocessed first training data; generating a first timbre basis function according to each feature encoding vector and generating a first timbre adjustment coefficient corresponding to each feature encoding vector; and generating first timbre features according to each first timbre basis function and the first timbre adjustment coefficient.
[0157] In one embodiment, after the processing of the feature encoding vectors through a timbre learning module to obtain timbre features when the computer program is executed by a processor, it further includes: adjusting the lengths of the feature encoding vectors to obtain corresponding target encoding feature vectors; adjusting the lengths of the feature encoding vectors to obtain corresponding target encoding feature vectors, including: obtaining corresponding length information according to the first timbre features; and performing corresponding length adjustment on the feature encoding vectors according to the length information to obtain corresponding target encoding vectors.
[0158] In one embodiment, when the computer program is executed by a processor, processing the feature encoding vector through a timbre learning module to obtain a timbre feature includes: obtaining a corresponding target decoding vector by decoding the target encoding vector for each semantic decoding layer; generating a second timbre basis function according to each target decoding vector, and generating a second timbre adjustment coefficient corresponding to each target decoding vector; generating a second timbre feature according to each second timbre basis function and the second timbre adjustment coefficient.
[0159] In one embodiment, when the computer program is executed by a processor, decoding according to the feature encoding vector and the timbre feature to obtain a corresponding target spectrum includes: decoding according to the target feature decoding vector and the second timbre feature to obtain a corresponding target spectrum.
[0160] In one embodiment, when the computer program is executed by a processor, generating an objective function according to the first training data, the timbre feature, and the target spectrum includes: obtaining a real loss value according to the difference between the real spectrum and the target spectrum, where the real spectrum is the spectrum data corresponding to the speech data; obtaining a length loss value according to the difference between the length information and the target length information, where the length information is obtained according to each first timbre feature, and the target length information is the length data corresponding to the speech data; obtaining a timbre loss value according to the first timbre feature and the second timbre feature; generating an objective function according to the real loss value, the length loss value, and the timbre loss value.
[0161] In one embodiment, when the computer program is executed by a processor, performing transfer learning on the trained basic model according to the second training data to obtain a target model includes: modifying the expression of the timbre adjustment coefficient of the timbre learning module; performing training according to the second training data to train the modified expression of the timbre adjustment coefficient of the timbre learning module to obtain a target model.
[0162] In one embodiment, a speech synthesis method implemented when the computer program is executed by a processor includes: obtaining text information; inputting the text information into the target model obtained by transfer learning to obtain a target speech corresponding to the text information.
[0163] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0164] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0165] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0166] The above-described embodiments merely represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A training method for a speech synthesis model, characterized in that, The method includes: Obtaining first training data; Preprocessing the first training data; Encoding the preprocessed first training data to obtain corresponding feature encoding vectors, and processing the feature encoding vectors through the timbre learning module to obtain timbre features; Decoding according to the feature encoding vectors and the timbre features to obtain corresponding target spectrograms; Generating an objective function according to the first training data, the timbre features and the target spectrograms; Performing iterative training according to the objective function to obtain a basic model, the basic model includes a timbre learning module, and the timbre learning module differentiates the differences of different timbres during the training process of the first training data and obtains the model parameters corresponding to the different timbres; Obtaining second training data; Performing transfer learning on the trained basic model according to the second training data to obtain a target model, wherein only the model parameters of the timbre learning module are modified during the transfer learning; The processing of the feature encoding vectors through the timbre learning module to obtain timbre features includes: Obtaining the corresponding feature encoding vectors encoded by each semantic encoding layer for the preprocessed first training data; Generating a first timbre basis function according to each feature encoding vector, and generating a first timbre adjustment coefficient corresponding to each feature encoding vector; Generating a first timbre feature according to each first timbre basis function and the first timbre adjustment coefficient.
2. The method according to claim 1, wherein The preprocessing of the first training data includes: Performing word segmentation on the text information in the first training data; Converting the segmented text information into corresponding standard phonemes.
3. The method according to claim 2, wherein After the processing of the feature encoding vectors through the timbre learning module to obtain timbre features, it further includes: Adjusting the lengths of the feature encoding vectors to obtain corresponding target encoding feature vectors; The adjusting the lengths of the feature encoding vectors to obtain corresponding target encoding feature vectors includes: Obtaining corresponding length information according to the first timbre feature; Performing corresponding length adjustment on the feature encoding vectors according to the length information to obtain corresponding target encoding vectors.
4. The method according to claim 3, characterized in that, The processing of the feature encoding vectors through the timbre learning module to obtain timbre features includes: Obtaining the corresponding target decoding vectors decoded by each semantic decoding layer for the target encoding vectors; Generating a second timbre basis function according to each target decoding vector, and generating a second timbre adjustment coefficient corresponding to each target decoding vector; Generating a second timbre feature according to each second timbre basis function and the second timbre adjustment coefficient.
5. The method according to claim 4, wherein The decoding according to the feature encoding vectors and the timbre features to obtain corresponding target spectrograms includes: Decoding according to the target decoding vectors and the second timbre features to obtain corresponding target spectrograms.
6. The method according to claim 5, characterized in that, The generating an objective function according to the first training data, the timbre features and the target spectrograms includes: Obtaining a real loss value according to the difference between the real spectrogram and the target spectrogram, and the real spectrogram is the spectrogram data corresponding to the speech data; A length loss value is obtained according to the difference between the length information and the target length information, where the length information is obtained based on each of the first timbre features, and the target length information is the length data corresponding to the voice data; A timbre loss value is obtained according to the first timbre feature and the second timbre feature; The objective function is generated according to the true loss value, the length loss value, and the timbre loss value.
7. The method according to claim 6, wherein The obtaining the target model by performing transfer learning on the trained basic model according to the second training data includes: Modifying the expression of the timbre adjustment coefficient of the timbre learning module; Training according to the second training data to train the timbre adjustment coefficient of the modified expression of the timbre learning module to obtain the target model.
8. A voice synthesis method, characterized in that, Including: Obtaining text information; Inputting the text information into a target model trained by using the training method of the voice synthesis model according to any one of claims 1-7 to obtain the target voice corresponding to the text information.
9. A training device for a speech synthesis model, characterized in that, The device includes: A first acquisition module, configured to acquire first training data; A training module, configured to preprocess the first training data; encode the preprocessed first training data to obtain corresponding feature encoding vectors, and acquire the corresponding feature encoding vectors obtained by each semantic encoding layer encoding the preprocessed first training data; generate first timbre basis functions according to each of the feature encoding vectors, and generate first timbre adjustment coefficients corresponding to each of the feature encoding vectors; generate first timbre features according to each of the first timbre basis functions and the first timbre adjustment coefficients; decode according to the feature encoding vectors and the timbre features to obtain corresponding target spectra; generate an objective function according to the first training data, the timbre features, and the target spectra; perform iterative training according to the objective function to obtain a basic model, where the basic model includes a timbre learning module, and the timbre learning module differentiates the differences of different timbres during the training process of the first training data and obtains the model parameters corresponding to the different timbres; A second acquisition module, configured to acquire second training data; A second training module, configured to perform transfer learning on the trained basic model according to the second training data to obtain a target model, where only the model parameters of the timbre learning module are modified during the transfer learning.
10. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 or 8 are implemented.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 7 or 8 are implemented.
12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 7 or 8 are implemented.
Citation Information
Patent Citations
Voice synthesis model training method and device
CN111508470A