Training Method, Device, Electronic Device and Storage Medium for Speech Synthesis Model

Through the VAE and GAN network architecture speech synthesis model, phoneme, pitch, duration, linguistic and vibrato information is input, and the model parameters are trained, which solves the problems of unnatural and inaccurate speech synthesis, and achieves high-quality natural speech synthesis.

CN115953995BActive Publication Date: 2025-08-05SHANGHAI MOBVOI INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211636225.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-19
Publication Date
2025-08-05
Estimated Expiration
2042-12-19

AI Technical Summary

Technical Problem

In the existing speech synthesis technology, the synthesized speech is unnatural, inaccurate pitch, and lack of adjustment of vibrato and lycophony, resulting in differences in the synthesis results from actual needs.

Method used

The speech synthesis model is adopted with VAE network and GAN network architecture, and phoneme, pitch, duration, lingot and vibrato information is input, and components such as the duration extraction module, linear transformation module and decoder are trained to adjust the model parameters to achieve natural speech synthesis.

Benefits of technology

It realizes high-quality speech synthesis, supports prediction and adjustment of vibrato and lycophony, and the synthesis is more natural and the pitch is more accurate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115953995B_ABST
    Figure CN115953995B_ABST
Patent Text Reader

Abstract

The present disclosure provides a training method, device, electronic device and storage medium for a speech synthesis model, the method comprising: inputting music information corresponding to a first speech sample into a duration extraction module to obtain a score sample embedding value; inputting the score sample embedding value and the pitch sample embedding value corresponding to the score sample embedding value into a linear transformation module to perform dimensionality reduction; using the output of the linear transformation module as the input of a framework network module to obtain a first predicted sample feature corresponding to the music information; obtaining a latent feature corresponding to the first speech sample; inputting the latent feature into a decoder to obtain a predicted speech sample corresponding to the latent feature; adjusting the parameters of the decoder based on the first speech sample and the predicted speech sample; adjusting the parameters of the linear transformation module and the framework network module based on the first predicted sample feature and the latent feature; and adjusting the parameters of the pitch extraction module based on the pitch sample embedding value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of speech synthesis technology, and in particular to a training method, device, electronic device, and storage medium for a speech synthesis model. Background Art

[0002] In related technologies, speech synthesis is usually achieved using a two-stage speech synthesis model, which requires separate training of the acoustic model and vocoder. In addition, the information considered during speech synthesis is not comprehensive, resulting in a difference between the final synthesis result and the actual needs, which cannot meet user needs. Summary of the Invention

[0003] The present disclosure provides a training method, device, electronic device and storage medium for a speech synthesis model to at least solve the above technical problems existing in the prior art.

[0004] According to a first aspect of the present disclosure, a method for training a speech synthesis model is provided, comprising:

[0005] Input the music information corresponding to the first speech sample into the duration extraction module to obtain the embedded value of the music score sample;

[0006] Inputting the music score sample embedding value and the pitch sample embedding value corresponding to the music score sample embedding value into a linear transformation module for dimensionality reduction;

[0007] Using the output of the linear transformation module as the input of the framework network module to obtain the first predicted sample feature corresponding to the music information;

[0008] Obtaining latent features corresponding to the first speech sample;

[0009] Inputting the latent features into a decoder to obtain a predicted speech sample corresponding to the latent features;

[0010] Based on the first speech sample and the predicted speech sample, the parameters of the decoder are adjusted; based on the first predicted sample features and the latent features, the parameters of the linear transformation module and the framework network module are adjusted; based on the pitch sample embedding value, the parameters of the pitch extraction module are adjusted.

[0011] According to a second aspect of the present disclosure, a speech synthesis method is provided, which is implemented based on a speech synthesis model obtained by the speech synthesis model training method provided in the first aspect above, and the method comprises:

[0012] Input the music information to be synthesized into the duration extraction module included in the speech synthesis model to obtain the music score embedding value;

[0013] Inputting the music score embedding value and the pitch sample embedding value corresponding to the music score embedding value into a linear transformation module included in a speech synthesis model for dimensionality reduction;

[0014] Using the output of the linear transformation module as the input of the framework network module included in the speech synthesis model to obtain the first feature corresponding to the music information to be synthesized;

[0015] The first feature is input into the decoder included in the speech synthesis model to obtain speech information corresponding to the music information to be synthesized.

[0016] According to a third aspect of the present disclosure, a device for training a speech synthesis model is provided, the device comprising:

[0017] A first embedding value acquisition unit is used to input the music information corresponding to the first voice sample into the duration extraction module to obtain the embedding value of the music score sample;

[0018] A first linear transformation unit is configured to input the music score sample embedding value and the pitch sample embedding value corresponding to the music score sample embedding value into a linear transformation module for dimensionality reduction;

[0019] a first acquisition unit, configured to use the output of the linear transformation module as the input of the framework network module to acquire a first predicted sample feature corresponding to the music information;

[0020] A second acquiring unit, configured to acquire latent features corresponding to the first speech sample;

[0021] A first decoding unit, configured to input the latent feature into a decoder to obtain a predicted speech sample corresponding to the latent feature;

[0022] An adjustment unit is used to adjust the parameters of the decoder based on the first speech sample and the predicted speech sample; adjust the parameters of the linear transformation module and the framework network module based on the first predicted sample feature and the latent feature; and adjust the parameters of the pitch extraction module based on the pitch sample embedding value.

[0023] According to a fourth aspect of the present disclosure, a speech synthesis device is provided, which is implemented based on a speech synthesis model obtained by the speech synthesis model training method provided in the first aspect above, and the device includes:

[0024] A second embedding value acquisition unit is used to input the music information to be synthesized into the duration extraction module included in the speech synthesis model to obtain the music score embedding value;

[0025] a second linear transformation unit, configured to input the music score embedding value and the pitch sample embedding value corresponding to the music score embedding value into a linear transformation module included in a speech synthesis model for dimensionality reduction;

[0026] a third acquisition unit, configured to use the output of the linear transformation module as an input to a framework network module included in a speech synthesis model, and acquire a first feature corresponding to the music information to be synthesized;

[0027] The second decoding unit is used to input the first feature into the decoder included in the speech synthesis model to obtain the speech information corresponding to the music information to be synthesized.

[0028] According to a fifth aspect of the present disclosure, there is provided an electronic device, including:

[0029] at least one processor; and

[0030] a memory communicatively connected to the at least one processor; wherein,

[0031] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method described in the present disclosure.

[0032] According to a sixth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method described in the present disclosure.

[0033] The training method of the speech synthesis model disclosed in the present invention comprises the following steps: inputting the music information corresponding to the first speech sample into the duration extraction module to obtain the score sample embedding value; inputting the score sample embedding value and the pitch sample embedding value corresponding to the score sample embedding value into the linear transformation module for dimensionality reduction; using the output of the linear transformation module as the input of the framework network module to obtain the first predicted sample feature corresponding to the music information; obtaining the latent feature corresponding to the first speech sample; inputting the latent feature into the decoder to obtain the predicted speech sample corresponding to the latent feature; adjusting the parameters of the decoder based on the first speech sample and the predicted speech sample; adjusting the parameters of the linear transformation module and the framework network module based on the first predicted sample feature and the latent feature; and adjusting the parameters of the pitch extraction module based on the pitch sample embedding value. In this way, the problem of unnatural synthesized speech and inaccurate pitch in the current singing synthesis process can be effectively solved.

[0034] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood by reading the detailed description below with reference to the accompanying drawings, in which several embodiments of the present disclosure are shown by way of example and not limitation, wherein:

[0036] In the drawings, the same or corresponding reference numerals denote the same or corresponding parts.

[0037] Figure 1 An optional flow chart of a method for training a speech synthesis model provided by an embodiment of the present disclosure is shown;

[0038] Figure 2 A schematic diagram of a speech synthesis model in the training phase provided by an embodiment of the present disclosure is shown;

[0039] Figure 3 A schematic diagram showing an optional flow chart of a speech synthesis method provided by an embodiment of the present disclosure is shown;

[0040] Figure 4 A schematic diagram of a speech synthesis model in the inference stage provided by an embodiment of the present disclosure is shown;

[0041] Figure 5 An optional structural diagram of a training device for a speech synthesis model provided in an embodiment of the present disclosure is shown;

[0042] Figure 6 An optional structural diagram of a speech synthesis device provided by an embodiment of the present disclosure is shown;

[0043] Figure 7 A schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0044] To make the purposes, features, and advantages of the present disclosure more apparent and understandable, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present disclosure without creative work shall fall within the scope of protection of the present disclosure.

[0045] First, the English abbreviations involved in the embodiments of the present disclosure are explained:

[0046] 1)VAE is a speech synthesis acoustic model structure.

[0047] 2) HIFiGan Vocoder is a vocoder based on Gan network (generative adversarial network).

[0048] 3) Gan is a generative adversarial network.

[0049] Related technologies in the field of speech synthesis, particularly those synthesizing audio (singing) from text, suffer from unnatural synthesis effects, the need to separately train acoustic models and vocoders, and a lack of support for note adjustments like vibrato and legato. Current existing technologies do not explicitly introduce these adjustments during synthesis.

[0050] To address the shortcomings of related technologies, the present invention provides a speech synthesis model training method and speech synthesis method based on a VAE and GAN network architecture. The acoustic model is primarily a VAE model, and the vocoder is a GAN. The model inputs include phonemes, pitch information (pitch nodes, including the fundamental frequency mean converted from musicxml or midinode), duration information, legato information, and vibrato information, addressing all or some of the aforementioned technical issues.

[0051] Figure 1 An optional flow chart of a method for training a speech synthesis model provided by an embodiment of the present disclosure is shown; Figure 2 A schematic diagram of a speech synthesis model in the training phase provided by an embodiment of the present disclosure is shown.

[0052] like Figure 2 As shown, the speech synthesis model can be represented by an acoustic model and a vocoder, wherein the posterior encoder part includes a feature extraction module and a second encoding module; the acoustic model part includes a first prediction module, a first encoding module, a second prediction module, a pitch extraction module, a pitch prediction module, a duration extraction module, a linear transformation module and a frame network module; the vocoder includes a decoder.

[0053] Step S101: input the music information corresponding to the first speech sample into a duration extraction module to obtain an embedding value of the music score sample.

[0054] In some embodiments, the music information includes phonemes, duration information, pitch information, legato information, and vibrato information corresponding to the first voice sample.

[0055] The phonemes are converted from the text corresponding to the first speech sample. If the first speech sample is a Chinese song, the phonemes are initials and finals. The duration information can include the duration of each character, or the duration of each character's initial and final. The pitch information can include pitch features in musicxml. The legato information is the legato feature in the first speech sample. The vibrato information is the vibrato feature in the first speech sample. In other words, in the disclosed embodiment, the speech synthesis model is trained based on the first speech sample and the music information corresponding to the first speech sample.

[0056] In some embodiments, before inputting the music information corresponding to the first speech sample into the duration extraction module to obtain the score sample embedding value, the training device of the speech synthesis model (hereinafter referred to as the first device) also needs to obtain the phoneme sample embedding value of the phoneme, the duration sample embedding value of the duration information, and the first sample embedding value corresponding to the legato information and the vibrato information; wherein the first sample embedding value includes the legato embedding value corresponding to the legato information and the vibrato embedding value corresponding to the vibrato information.

[0057] When implementing it specifically, Figure 2 As shown, the first device inputs the phoneme into the first encoding module to obtain the phoneme sample embedding value (Text embedding) corresponding to the phoneme; inputs the duration information into the second prediction module to obtain the duration sample embedding value corresponding to the duration information; inputs the legato information and the vibrato information into the first prediction module to obtain the first sample embedding value corresponding to the legato information and the vibrato information; it should be noted that the legato information and the vibrato information can be input into the first prediction module separately to obtain the corresponding embedding values respectively, or can be input into the first prediction module at the same time.

[0058] Specifically, the first device inputs the phoneme sample embedding value, the duration sample embedding value, the first sample embedding value, and the pitch information into the duration extraction module (LR) to obtain the music score sample embedding value corresponding to the first speech sample. In some optional embodiments, the first device may also input a user identifier (user ID) together with the phoneme sample embedding value, the duration sample embedding value, the first sample embedding value, and the pitch information into the LR to obtain the music score sample embedding value.

[0059] The duration extraction module performs upsampling processing on the phoneme sample embedding value, the duration sample embedding value, the first sample embedding value and the pitch information, so that the phoneme sample embedding value, the duration sample embedding value, the first sample embedding value and the pitch information reach the frame level.

[0060] Step S102: input the music score sample embedding value and the pitch sample embedding value corresponding to the music score sample embedding value into a linear transformation module for dimensionality reduction.

[0061] In some embodiments, as Figure 2 As shown, the first device inputs the music score sample embedding value and the pitch sample embedding value corresponding to the music score sample embedding value into the linear transformation module. Before performing dimensionality reduction, the pitch sample embedding value corresponding to the music score sample embedding value can also be obtained.

[0062] During specific implementation, the first device inputs the pitch sample embedding value into the pitch extraction module to obtain the pitch sample embedding value.

[0063] In some embodiments, the first device inputs the music score sample embedding value and the pitch sample embedding value into the linear transformation module, and performs dimensionality reduction processing on the music score sample embedding value and the pitch sample embedding value.

[0064] Step S103: Using the output of the linear transformation module as the input of the framework network module to obtain the first predicted sample feature corresponding to the music information.

[0065] In some embodiments, as Figure 2 As shown, the first device uses the output of the linear transformation module as the input of the framework network module to obtain the first predicted sample features corresponding to the music information.

[0066] Specifically, the essence of the LR module is to copy each input. The music score sample embedding value and the pitch sample embedding value include a large amount of repeated data, which leads to the need to delete the repeated data through the framework network module to finally obtain the first predicted sample feature.

[0067] Step S104: Obtain latent features corresponding to the first speech sample.

[0068] In some embodiments, as Figure 2 As shown, the first device inputs the first speech sample into a feature extraction module to obtain sample features corresponding to the first speech sample; and inputs the sample features into the second encoding module to obtain latent features corresponding to the first speech sample.

[0069] Step S105: input the latent features into a decoder to obtain a predicted speech sample corresponding to the latent features.

[0070] In some embodiments, as Figure 2As shown, the first device inputs the latent features into the decoder and outputs the predicted speech samples corresponding to the latent features.

[0071] In some optional embodiments, the first device may further input the predicted speech sample into a discriminator composed of an MSD and an MPD. The discriminator outputs a binary classification, respectively characterizing whether the predicted speech sample is real audio or machine-synthesized audio. The output of the discriminator is used to determine the quality of the speech synthesis model, which directly determines the quality of the generated predicted speech sample.

[0072] Step S106: Adjust the parameters of the decoder based on the first speech sample and the predicted speech sample; adjust the parameters of the linear transformation module and the framework network module based on the first predicted sample feature and the latent feature; and adjust the parameters of the pitch extraction module based on the pitch sample embedding value.

[0073] In some embodiments, the first device adjusts parameters of the speech synthesis model based on a loss function, wherein the loss function may include a first loss function, a second loss function, and a third loss function; optionally, a fourth loss function.

[0074] In a specific implementation, the first device confirms a first loss function based on the first speech sample and the predicted speech sample; and adjusts the parameters of the decoder based on the first loss function.

[0075] When implementing it specifically, Figure 2 As shown, the first device inputs the latent features into the flow module (Flow) included in the speech synthesis model to obtain the mapped latent features corresponding to the latent features; based on the mapped latent features and the first predicted sample features, a second loss function is determined; and based on the second loss function, the parameters of the linear transformation module, the framework network, the feature extraction module included in the speech synthesis model, the first encoding module included in the speech synthesis model, and the second encoding module included in the speech synthesis model are adjusted. The flow module is used to change the distribution of the latent features from complex to simple; the second loss function can be KL Loss.

[0076] When implementing it specifically, Figure 2 As shown, the first device inputs the pitch sample embedding value into the pitch prediction module included in the speech synthesis model, confirms that the output of the pitch prediction module is the predicted pitch corresponding to the pitch sample embedding value; determines a third loss function based on the predicted pitch and the labeled pitch corresponding to the pitch sample embedding value; and adjusts the parameters of the pitch extraction module and the pitch prediction module based on the third loss function. Optionally, the third loss function can be determined based on the MSE.

[0077] In a specific implementation, in response to the duration information including byte duration, the first device obtains the predicted initial duration and predicted final duration output by the first prediction module; based on the predicted initial duration, predicted final duration, marked initial duration and marked final duration corresponding to the duration information, the fourth loss function is confirmed; based on the fourth loss function, the parameters of the second prediction module are adjusted. That is to say, if the duration information does not include byte duration, but only includes initial duration and final duration, there is no need to confirm the fourth loss function, nor to adjust the parameters of the second prediction module. Optionally, the fourth loss function can be determined based on MSE.

[0078] In some optional embodiments, the first device can also confirm a fifth loss function based on the output of the discriminator and the cross-entropy loss function; the fifth loss function is used to adjust the parameters of at least one of the first encoding module, the second prediction module, the pitch extraction module, the first prediction module, the linear transformation module and the framework network module.

[0079] Thus, the speech synthesis model training method provided in this disclosure, as an end-to-end solution, can effectively address the current issues of unnatural and inaccurate pitch in singing synthesis. The input is MusicXML, and high-quality sound is synthesized by predicting market and pitch. Legato and vibrato features are also introduced to enable legato and vibrato prediction, while also supporting adjustable legato and vibrato, providing richer synthesis control and more natural sound.

[0080] Figure 3 shows an optional flow chart of the speech synthesis method provided by an embodiment of the present disclosure, Figure 4 A schematic diagram of a speech synthesis model in the inference stage provided by an embodiment of the present disclosure is shown.

[0081] The Flow module is a reversible module, such as Figure 2 As shown, when the input is a latent feature, the distribution of the latent feature can be changed from complex to simple; similarly, when the input is a first feature, the distribution of the first feature can be changed from simple to complex, so that the decoder can output the corresponding speech information.

[0082] Step S301: input the music information to be synthesized into the duration extraction module included in the speech synthesis model to obtain the music score embedding value.

[0083] In some embodiments, the music information to be synthesized includes the phonemes, duration information, pitch information, legato information, and vibrato information.

[0084] In some embodiments, as Figure 4As shown, the speech synthesis device (hereinafter referred to as the second device) inputs the phoneme into the first encoding module to obtain a phoneme embedding value (text embedding) corresponding to the phoneme; inputs the duration information into the second prediction module to obtain a duration embedding value corresponding to the duration information; inputs the legato information and the vibrato information into the first prediction module to obtain a first embedding value corresponding to the legato information and the vibrato information; it should be noted that the legato information and the vibrato information can be input into the first prediction module separately to obtain corresponding embedding values respectively, or can be input into the first prediction module simultaneously.

[0085] Specifically, the second device inputs the phoneme embedding value, the duration embedding value, the first embedding value, and the pitch information into the duration extraction module to obtain the musical score embedding value. In some optional embodiments, the second device may also input a user identifier (user ID) together with the phoneme embedding value, the duration embedding value, the first embedding value, and the pitch information into LR to obtain the musical score embedding value.

[0086] Step S302: input the music score embedding value and the pitch embedding value corresponding to the music score embedding value into a linear transformation module included in a speech synthesis model for dimensionality reduction.

[0087] In some embodiments, as Figure 4 As shown, the second device inputs the music score embedding value and the pitch embedding value corresponding to the music score embedding value into a linear transformation module, and before performing dimensionality reduction, the pitch embedding value corresponding to the music score embedding value can also be obtained.

[0088] During specific implementation, the second device inputs the music score embedding value into the pitch extraction module to obtain the pitch embedding value.

[0089] In some embodiments, the first device inputs the music score embedding value and the pitch embedding value into the linear transformation module, and performs dimensionality reduction processing on the music score embedding value and the pitch embedding value.

[0090] Step S303: Using the output of the linear transformation module as the input of the framework network module included in the speech synthesis model to obtain the first feature corresponding to the music information to be synthesized.

[0091] In some embodiments, as Figure 4 As shown, the second device uses the output of the linear transformation module as the input of the framework network module to obtain the first feature corresponding to the music information.

[0092] Specifically, the essence of the LR module is to copy (i.e., upsample) each input. The score embedding value and the pitch sample embedding value include a large amount of duplicate data, which leads to the need to delete the duplicate data through the framework network module to finally obtain the first feature.

[0093] Step S304: input the first feature into the decoder included in the speech synthesis model to obtain speech information corresponding to the music information to be synthesized.

[0094] In some embodiments, as Figure 4 As shown, before the second device inputs the first feature into the decoder, it also needs to input the first feature into the stream module to change the distribution of the first feature from simple to complex; then the output of the stream module is input into the decoder to output the voice information corresponding to the music information to be synthesized.

[0095] Thus, the speech synthesis method provided by this disclosure, as an end-to-end solution, can effectively address the current issues of unnatural and inaccurate synthesized speech in singing synthesis. It takes music XML as input and predicts market and pitch to produce high-quality sound. It also incorporates legato and vibrato features to enable legato and vibrato prediction, while also supporting adjustable legato and vibrato, providing richer synthesis control and more natural sound.

[0096] Figure 5 A schematic diagram of an optional structure of a training device for a speech synthesis model provided in an embodiment of the present disclosure is shown.

[0097] In some embodiments, the training device 500 for the speech synthesis model includes a first embedding value acquisition unit 501, a first linear transformation unit 502, a first acquisition unit 503, a second acquisition unit 504, a first decoding unit 505 and an adjustment unit 506.

[0098] The first embedding value acquisition unit 501 is used to input the music information corresponding to the first speech sample into the duration extraction module to obtain the embedding value of the music score sample;

[0099] The first linear transformation unit 502 is used to input the music score sample embedding value and the pitch sample embedding value corresponding to the music score sample embedding value into a linear transformation module for dimensionality reduction;

[0100] The first acquisition unit 503 is configured to use the output of the linear transformation module as the input of the framework network module to acquire a first prediction sample feature corresponding to the music information;

[0101] The second acquiring unit 504 is configured to acquire latent features corresponding to the first speech sample;

[0102] The first decoding unit 505 is configured to input the latent feature into a decoder to obtain a predicted speech sample corresponding to the latent feature;

[0103] The adjustment unit 506 is used to adjust the parameters of the decoder based on the first speech sample and the predicted speech sample; adjust the parameters of the linear transformation module and the framework network module based on the first predicted sample feature and the latent feature; and adjust the parameters of the pitch extraction module based on the pitch sample embedding value.

[0104] The first embedding value acquiring unit 501 is specifically configured to input the legato information and vibrato information included in the music information into a first prediction module, and acquire first sample embedding values corresponding to the legato information and the vibrato information;

[0105] Inputting the phonemes included in the music information into a first encoding module to obtain phoneme sample embedding values corresponding to the phonemes;

[0106] Inputting the duration information included in the music information into a second prediction module to obtain a duration sample embedding value corresponding to the duration information;

[0107] The first sample embedding value, the phoneme sample embedding value, the duration sample embedding value, and the pitch information included in the music information are input into the duration extraction module to obtain the score sample embedding value.

[0108] The first embedding value acquisition unit 501 is further configured to input the music score sample embedding value and the pitch sample embedding value corresponding to the music score sample embedding value into the linear transformation module for dimensionality reduction before inputting the music score sample embedding value and the pitch sample embedding value corresponding to the music score sample embedding value into the pitch extraction module to obtain the pitch sample embedding value corresponding to the music score sample embedding value.

[0109] The second acquisition unit 504 is specifically configured to input the first speech sample into the feature extraction module and the second encoding module included in the speech synthesis model to obtain latent features corresponding to the first speech sample.

[0110] The adjustment unit 506 is specifically configured to determine a first loss function based on the first speech sample and the predicted speech sample;

[0111] Parameters of the decoder are adjusted based on the first loss function.

[0112] The adjustment unit 506 is specifically configured to input the latent feature into the stream module included in the speech synthesis model to obtain a mapped latent feature corresponding to the latent feature;

[0113] Determining a second loss function based on the mapped latent feature and the first predicted sample feature;

[0114] Based on the second loss function, the parameters of the linear transformation module, the framework network, the feature extraction module included in the speech synthesis model, the first encoding module included in the speech synthesis model, and the second encoding module included in the speech synthesis model are adjusted.

[0115] The adjusting unit 506 is specifically configured to input the pitch sample embedding value into a pitch prediction module included in the speech synthesis model, and confirm that the output of the pitch prediction module is a predicted pitch corresponding to the pitch sample embedding value;

[0116] Determine a third loss function based on the predicted pitch and the labeled pitch corresponding to the pitch sample embedding value;

[0117] Based on the third loss function, parameters of the pitch extraction module and the pitch prediction module are adjusted.

[0118] The adjusting unit 506 is specifically configured to obtain the predicted initial consonant duration and the predicted final duration output by the first prediction module in response to the duration information including the byte duration;

[0119] Determine a fourth loss function based on the predicted initial consonant duration, the predicted final consonant duration, the marked initial consonant duration, and the marked final consonant duration corresponding to the duration information;

[0120] Based on the fourth loss function, parameters of the second prediction module are adjusted.

[0121] Figure 6 An optional structural diagram of a speech synthesis device provided in an embodiment of the present disclosure is shown, and the following description will be given according to each step.

[0122] In some embodiments, the speech synthesis apparatus 600 includes a second embedding value acquisition unit 601 , a second linear transformation unit 602 , a third acquisition unit 603 and a second decoding unit 604 .

[0123] The second embedding value acquisition unit 601 is used to input the music information to be synthesized into the duration extraction module included in the speech synthesis model to obtain the music score embedding value;

[0124] A second linear transformation unit 602 is configured to input the music score embedding value and the pitch sample embedding value corresponding to the music score embedding value into a linear transformation module included in a speech synthesis model for dimensionality reduction;

[0125] a third acquiring unit 603, configured to use the output of the linear transformation module as an input of a framework network module included in a speech synthesis model, and acquire a first feature corresponding to the music information to be synthesized;

[0126] The second decoding unit 604 is used to input the first feature into the decoder included in the speech synthesis model to obtain the speech information corresponding to the music information to be synthesized.

[0127] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device and a readable storage medium.

[0128] Figure 7 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0129] like Figure 7 As shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0130] Multiple components in the electronic device 800 are connected to the I / O interface 805, including an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0131] The computing unit 801 can be various general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 801 performs the various methods and processes described above, such as the training method of the speech synthesis model or the speech synthesis method. For example, in some embodiments, the training method of the speech synthesis model or the speech synthesis method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the training method of the speech synthesis model or the speech synthesis method described above can be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to execute the speech synthesis model training method or the speech synthesis method in any other appropriate manner (for example, by means of firmware).

[0132] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0133] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0134] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0135] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0136] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0137] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0138] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.

[0139] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. Throughout the present disclosure, "plurality" means two or more, unless otherwise specifically defined.

[0140] The above description is merely a specific embodiment of the present disclosure, but the scope of protection of the present disclosure is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this disclosure should be included in the scope of protection of the present disclosure. Therefore, the scope of protection of the present disclosure should be based on the scope of protection of the claims.

Claims

1. A method for training a speech synthesis model, characterized in that: The speech synthesis model includes a first prediction module, a second prediction module, a first encoding module, a duration extraction module, a linear transformation module, a skeleton network module, and a decoder; the output of the duration extraction module serves as the input of the linear transformation module, the output of the linear transformation module serves as the input of the skeleton network module, and the output of the skeleton network module serves as the input of the decoder after being processed by a stream module; the method includes: Inputting legato information and vibrato information included in the music information into a first prediction module to obtain first sample embedding values corresponding to the legato information and the vibrato information; inputting phonemes included in the music information into a first encoding module to obtain phoneme sample embedding values corresponding to the phonemes; inputting duration information included in the music information into a second prediction module to obtain duration sample embedding values corresponding to the duration information; inputting the first sample embedding value, the phoneme sample embedding value, the duration sample embedding value, and the pitch information included in the music information into the duration extraction module to obtain a score sample embedding value; Inputting the music score sample embedding value and the pitch sample embedding value corresponding to the music score sample embedding value into a linear transformation module for dimensionality reduction; Using the output of the linear transformation module as the input of the framework network module to obtain the first predicted sample feature corresponding to the music information; Obtaining latent features corresponding to the first speech sample; Inputting the latent features into a decoder to obtain a predicted speech sample corresponding to the latent features; Based on the first speech sample and the predicted speech sample, the parameters of the decoder are adjusted; based on the first predicted sample features and the latent features, the parameters of the linear transformation module and the framework network module are adjusted; based on the pitch sample embedding value, the parameters of the pitch extraction module are adjusted.

2. The method according to claim 1, characterized in that Before inputting the music score sample embedding value and the pitch sample embedding value corresponding to the music score sample embedding value into the linear transformation module for dimensionality reduction, the method further comprises: The music score sample embedding value is input into the pitch extraction module to obtain the pitch sample embedding value corresponding to the music score sample embedding value.

3. The method according to claim 1, characterized in that The obtaining of latent features corresponding to the first speech sample includes: The first speech sample is input into a feature extraction module and a second encoding module included in the speech synthesis model to obtain latent features corresponding to the first speech sample.

4. The method according to claim 1, wherein The adjusting the parameters of the decoder based on the first speech sample and the predicted speech sample includes: determining a first loss function based on the first speech sample and the predicted speech sample; Parameters of the decoder are adjusted based on the first loss function.

5. The method according to claim 3, characterized in that The adjusting the parameters of the linear transformation module and the framework network module based on the first prediction sample feature and the latent feature includes: Inputting the latent feature into the stream module included in the speech synthesis model to obtain a mapped latent feature corresponding to the latent feature; Determining a second loss function based on the mapped latent feature and the first predicted sample feature; Based on the second loss function, the parameters of the linear transformation module, the framework network, the feature extraction module included in the speech synthesis model, the first encoding module included in the speech synthesis model, and the second encoding module included in the speech synthesis model are adjusted.

6. The method according to claim 1, characterized in that The adjusting the parameters of the pitch extraction module based on the pitch sample embedding value includes: Inputting the pitch sample embedding value into a pitch prediction module included in the speech synthesis model, and confirming that the output of the pitch prediction module is a predicted pitch corresponding to the pitch sample embedding value; Determine a third loss function based on the predicted pitch and the labeled pitch corresponding to the pitch sample embedding value; Based on the third loss function, parameters of the pitch extraction module and the pitch prediction module are adjusted.

7. The method according to claim 1, characterized in that The method further comprises: In response to the duration information including the byte duration, obtaining the predicted initial consonant duration and the predicted final duration output by the first prediction module; Determine a fourth loss function based on the predicted initial consonant duration, the predicted final consonant duration, the marked initial consonant duration, and the marked final consonant duration corresponding to the duration information; Based on the fourth loss function, parameters of the second prediction module are adjusted.

8. A training device for a speech synthesis model, characterized in that: The speech synthesis model includes a first prediction module, a second prediction module, a first encoding module, a duration extraction module, a linear transformation module, a skeleton network module, and a decoder; the output of the duration extraction module serves as the input of the linear transformation module, the output of the linear transformation module serves as the input of the skeleton network module, and the output of the skeleton network module serves as the input of the decoder after being processed by the stream module; the device includes: The first embedding value acquisition unit is configured to input legato information and vibrato information included in the music information into the first prediction module to obtain first sample embedding values corresponding to the legato information and vibrato information; input phonemes included in the music information into the first encoding module to obtain phoneme sample embedding values corresponding to the phonemes; input duration information included in the music information into the second prediction module to obtain duration sample embedding values corresponding to the duration information; and input the first sample embedding value, the phoneme sample embedding value, the duration sample embedding value, and the pitch information included in the music information into the duration extraction module to obtain a score sample embedding value. A first linear transformation unit is configured to input the music score sample embedding value and the pitch sample embedding value corresponding to the music score sample embedding value into a linear transformation module for dimensionality reduction; A first acquisition unit is configured to use the output of the linear transformation module as the input of the framework network module to acquire a first prediction sample feature corresponding to the music information; A second acquisition unit, configured to acquire latent features corresponding to the first speech sample; A first decoding unit, configured to input the latent feature into a decoder to obtain a predicted speech sample corresponding to the latent feature; An adjustment unit is used to adjust the parameters of the decoder based on the first speech sample and the predicted speech sample; adjust the parameters of the linear transformation module and the framework network module based on the first predicted sample feature and the latent feature; and adjust the parameters of the pitch extraction module based on the pitch sample embedding value.

9. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Rhythm-controllable Chinese and English mixed speech synthesis method and system

    CN112802450A

  • Song synthesis method and device capable of keeping pitch

    CN113506560A