Singing voice synthesis method and related device
By using formant model and tone conversion model in the singing vocal synthesis technology, the language and tone information are decoupled, and the synthesis of any tone cross-language singing voice is achieved, which solves the problem that the existing technology cannot synthesize across languages, and improves the flexibility and diversity of synthesis.
Patent Information
- Application Number
- CN202310126243.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-07
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2043-02-07
AI Technical Summary
Existing singing vocal synthesis technology cannot synthesize singing vocals across languages because the language and the target vocal cannot be decoupled.
By inputting the syllable sequence and fundamental frequency marker sequence of the audio to be synthesized into the formant peak model, formant peak characterization information without tone information is obtained and input into the tone conversion model. Combining the pitch information, a Mel spectral feature including the tone information of the target object is generated, and finally the synthesized audio signal is generated through the vocoder.
The synthesis of singing across languages of any tone is realized, solving the problem of decoupling of language and tone, and improving the flexibility and diversity of singing synthesis.
Smart Images

Figure CN116072143B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a singing synthesis method and related devices. Background Art
[0002] With the continuous development of artificial intelligence in the field of music, the singing voice synthesis technology in music applications has also received more and more attention. Among them, singing voice synthesis technology is a new application of speech synthesis technology, which mainly synthesizes music score information and lyrics close to the singing voice of real people through computer programs. At present, when performing singing voice synthesis, a data-driven neural network model is generally used to realize it. The language (i.e., language type, such as Mandarin, Cantonese, Minnan dialect, etc.) supported by the method is related to the singing voice data corresponding to the target timbre, wherein the singing voice data includes the singing voice and the sound of the accompaniment instrument emitted by the designated person (professional singer) (when there is no accompaniment instrument, the singing voice data is the singing voice emitted by the designated person), and the language and the target timbre cannot be decoupled, thereby, the singing voice of any timbre across languages cannot be synthesized using this method. Therefore, how to realize the singing voice synthesis across languages has become a problem to be solved. Summary of the invention
[0003] The embodiments of the present application provide a singing voice synthesis method and related devices, which can synthesize singing voices of any timbre and across languages.
[0004] In a first aspect, an embodiment of the present application provides a singing synthesis method, the method comprising:
[0005] Inputting the syllable sequence and the fundamental frequency marker sequence of the audio to be synthesized into the formant model in the target acoustic model to obtain the formant representation information of the audio to be synthesized, where the formant representation information is representation information without timbre information;
[0006] Input the formant representation information and pitch information of the audio to be synthesized into the timbre conversion model in the target acoustic model to obtain the Mel spectrum feature, which includes the timbre information of the target object. The timbre conversion model is obtained based on the sample audio training of the target object;
[0007] The Mel-spectrogram features are input into the vocoder to obtain the synthesized audio signal.
[0008] In the embodiment of the present application, the formant model can be used to decouple the timbre information and language of the audio to be synthesized, and obtain the formant representation information that is unrelated to the timbre; the formant representation information is input into the timbre conversion model in the target acoustic model corresponding to the target object, and the Mel spectrum feature including the timbre information of the target object can be obtained; the Mel spectrum feature is input into the vocoder to obtain the synthesized audio signal. It can be seen that the embodiment of the present application can be used to synthesize singing voices of any timbre across languages.
[0009] In an optional embodiment, before inputting the syllable sequence and the fundamental frequency marker sequence of the audio to be synthesized into the formant model in the target acoustic model, the method further includes:
[0010] Acquire a training sample audio set, where the training sample audio set includes sample audios of multiple objects;
[0011] Training the initialized acoustic model based on each sample audio in the training sample audio set to obtain a trained acoustic model;
[0012] Using the formant model in the trained acoustic model as the formant model in the target acoustic model;
[0013] The initialized acoustic model includes an initialized formant model and an initialized timbre conversion model.
[0014] In an optional implementation, after obtaining the trained acoustic model, the method further includes:
[0015] Fixing the parameters of the formant model in the trained acoustic model, inputting the sample audio corresponding to the target object into the trained acoustic model, and retraining the timbre conversion model in the trained acoustic model to obtain the trained timbre conversion model;
[0016] The trained timbre conversion model is used as the timbre conversion model in the target acoustic model.
[0017] In an optional implementation, the initialized acoustic model is trained based on each sample audio in the training sample audio set to obtain a trained acoustic model, including:
[0018] Obtaining a syllable sequence, a fundamental frequency marker sequence, pitch information, timbre information, and acoustic features of each training sample audio in the training sample audio set, where the acoustic features are Mel-spectrogram features;
[0019] The initialized acoustic model is trained using the syllable sequence, fundamental frequency marker sequence, pitch information, timbre information and acoustic features of each sample audio to obtain a trained acoustic model.
[0020] In an optional implementation, the initialized acoustic model is trained using the syllable sequence, fundamental frequency marker sequence, pitch information, timbre information, and acoustic features of each sample audio to obtain a trained acoustic model, including:
[0021] Inputting the syllable sequence and fundamental frequency marker sequence of each sample audio into the initialized formant model to obtain the formant representation information of each sample audio;
[0022] Input the formant representation information, pitch information and timbre information of each sample audio into the initialized timbre conversion model to obtain the predicted Mel spectrum features of each sample audio;
[0023] Determine a loss value for each audio sample based on the predicted Mel-spectrogram feature of each audio sample;
[0024] When the loss value does not meet the training stop condition, the parameters of the initialized acoustic model including the initialized resonance peak model and the initialized timbre conversion model are adjusted according to the loss value to obtain a trained acoustic model.
[0025] In an optional implementation, determining the loss value of each audio sample based on the predicted mel spectrum feature of each audio sample includes:
[0026] Determine the minimum mean square error between the predicted Mel spectrum feature and the acoustic feature of each sample audio, and determine the discriminator error corresponding to the predicted Mel spectrum feature of each sample audio;
[0027] Based on the minimum mean square error and the discriminator error, the loss value of each sample audio is determined.
[0028] In an optional implementation, determining a discriminator error corresponding to the predicted mel spectrum feature of each sample audio includes:
[0029] Inputting the predicted Mel spectrum feature of each sample audio into the initialized discriminator to obtain a discrimination result of the predicted Mel spectrum feature of each sample audio, wherein the discrimination result includes whether the predicted Mel spectrum feature is a real Mel spectrum feature or a synthesized Mel spectrum feature;
[0030] Based on the discrimination result of the predicted Mel spectrum feature of each sample audio, a discriminator error corresponding to the predicted Mel spectrum feature of each sample audio is determined.
[0031] In an optional implementation, after determining the loss value of each sample audio, the method further includes:
[0032] When the loss value does not meet the conditions for stopping training, the parameters in the initialized discriminator are adjusted according to the loss value to obtain the trained discriminator.
[0033] In a second aspect, an embodiment of the present application provides a singing voice synthesis device, the device comprising:
[0034] A processing unit, used for inputting the syllable sequence and the fundamental frequency marker sequence of the audio to be synthesized into the formant model in the target acoustic model to obtain the formant representation information of the audio to be synthesized, where the formant representation information is the representation information without timbre information;
[0035] The processing unit is further used to input the formant representation information and pitch information of the audio to be synthesized into the timbre conversion model in the target acoustic model to obtain Mel-spectrogram features, where the synthesized Mel-spectrogram features include the timbre information of the target object, and the timbre conversion model is obtained based on the sample audio training of the target object;
[0036] The processing unit is also used to input the Mel spectrum features into the vocoder to obtain a synthesized audio signal.
[0037] Among them, the optional implementation methods of each unit in the singing synthesis device can be found in the description of the first aspect above, and will not be repeated here.
[0038] In a third aspect, an embodiment of the present application further provides a computer device, comprising: a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the method described in the first aspect is implemented.
[0039] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in the first aspect is implemented.
[0040] In a fifth aspect, the embodiments of the present application further provide a computer program product or a computer program, wherein the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method according to the first aspect provided in the embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0042] Figure 1 is a sound spectrum schematic diagram provided in an embodiment of the present application;
[0043] Figure 2 It is a flow chart of a singing synthesis method provided in an embodiment of the present application;
[0044] Figure 3 is a schematic diagram of a syllable boundary marking provided in an embodiment of the present application;
[0045] Figure 4a It is a schematic structural diagram of an acoustic model provided by an embodiment of the present application;
[0046] Figure 4b It is a schematic structural diagram of a discriminator provided by an embodiment of the present application;
[0047] Figure 5a It is a schematic flowchart of a method for training an acoustic model provided by an embodiment of the present application;
[0048] Figure 5b It is a schematic flowchart of a method for determining a trained acoustic model provided by an embodiment of the present application;
[0049] Figure 6 It is a schematic diagram of formant characterization information provided by an embodiment of the present application;
[0050] Figure 7 It is a schematic diagram of a singing synthesis device provided by an embodiment of the present application;
[0051] Figure 8 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0052] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0053] To facilitate the understanding of the embodiments disclosed in the present application, some concepts related to the embodiments of the present application will be first elaborated. The elaboration of these concepts includes but is not limited to the following content.
[0054] 1. Syllable
[0055] A syllable is the smallest speech unit in which a single vowel phoneme and a consonant phoneme are combined to pronounce. Among them, a single vowel phoneme can also form a syllable by itself. For example, a syllable in Chinese is a speech unit formed by the combination of an initial consonant and a final vowel. Among them, a single final vowel can also form a syllable by itself, such as "o" and "a", etc.
[0056] A phoneme is the smallest unit of Chinese pronunciation. For example, the pronunciation units of the Chinese character "hao" are the initial consonant "h" and the final vowel "ao". Similarly, a grapheme is the smallest unit of Chinese writing.
[0057] 2. Fundamental frequency
[0058] In speech, fundamental frequency refers to the basic frequency of speech. When the sound-producing body produces sound due to vibration, the sound can generally be decomposed into many simple sine waves. In other words, all natural sounds are basically composed of many sine waves with different frequencies. Among them, sounds of different frequencies can form a sound spectrum. See Figure 1 , Figure 1 is a schematic diagram of a sound spectrum provided in an embodiment of the present application. Figure 1 As shown in the figure, the lower frequency sine wave is the fundamental tone, while the other higher frequency sine waves are overtones. In the field of music, the fundamental tone is the main element to distinguish the pitch. The frequency of the fundamental tone can also be called the pitch, which determines the melody. The overtones determine the timbre of the human voice. When we usually talk about whether singing is out of tune, we mean whether the singer's fundamental frequency is consistent with the pitch information in the song score.
[0059] 3. Music score
[0060] Music score is a regular combination of various written symbols that record the pitch or rhythm of music. Common music score formats include MIDI, MusicXML, etc. In the singing synthesis stage, the lyrics content and pitch information of the audio to be synthesized need to be given. Optionally, the computer device can extract the lyrics content and pitch information of the audio to be synthesized from the music score.
[0061] At present, when performing singing synthesis, a data-driven neural network model is generally used to implement it. The languages supported by this singing synthesis method are related to the singing data corresponding to the target timbre, and the language and the target timbre cannot be decoupled. Therefore, this method cannot be used to synthesize singing with any timbre across languages.
[0062] Based on this, the embodiment of the present application provides a singing voice synthesis method and related devices, which can synthesize singing voices of any timbre across languages. Among them, the method includes: a computer device inputs the syllable sequence and fundamental frequency marker sequence of the audio to be synthesized into the resonance peak model in the target acoustic model to obtain the resonance peak representation information of the audio to be synthesized, and the resonance peak representation information is the representation information without timbre information; the resonance peak representation information and pitch information of the audio to be synthesized are input into the timbre conversion model in the target acoustic model to obtain Mel spectrum features, and the synthesized Mel spectrum features include the timbre information of the target object, and the timbre conversion model is obtained based on the sample audio training of the target object; the Mel spectrum features are input into the vocoder to obtain a synthesized audio signal. It can be seen that by using the embodiment of the present application, singing voices of any timbre across languages can be synthesized.
[0063] It should be noted that the computer devices mentioned above can be terminal devices or servers. The terminal devices include but are not limited to smart phones, tablet computers, laptops, desktop computers, smart speakers, smart watches, smart cars, etc. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, but is not limited to this.
[0064] In order to facilitate understanding of the embodiments of the present application, the specific implementation method of the above-mentioned singing synthesis method is described in detail below using a computer device as the execution subject.
[0065] See also Figure 2 , Figure 2 Schematic diagram of a singing voice synthesis method provided in an embodiment of the present application. Figure 2 As shown, the singing voice synthesis method may include but is not limited to the following steps:
[0066] S201: Input the syllable sequence and the fundamental frequency marker sequence of the audio to be synthesized into the formant model in the target acoustic model to obtain the formant representation information of the audio to be synthesized.
[0067] The resonance peak representation information is representation information without timbre information.
[0068] The syllable sequence is the result of word embedding representation of the syllables of the synthesized audio. Optionally, the syllable sequence can be a frame-level syllable sequence. Optionally, word embedding representation of the syllables of the synthesized audio can obtain a syllable sequence of dimension [T, 128], where T represents the number of audio frames and 128 represents a vector.
[0069] The fundamental frequency marking sequence (VUV Decision) is a sequence for marking the fundamental frequency sequence of the audio to be synthesized. Optionally, the dimension of the fundamental frequency marking sequence is [T, 1], where T represents the number of frames. Optionally, when marking the fundamental frequency sequence of the audio to be synthesized, if the fundamental frequency value of the i-th frame is greater than 0, the value of VUV[i, 1] is 1, otherwise it is 0.
[0070] Among them, inputting the syllable sequence and the fundamental frequency marker sequence into the resonance peak model together can implicitly provide consonant information and vowel information, realize the adaptation of the phoneme duration in the syllable, and reduce the problem of mispronunciation.
[0071] In an optional implementation, before step S201, the computer device may also pre-process the audio to be synthesized to obtain a syllable sequence and a fundamental frequency sequence of the audio to be synthesized.
[0072] In this embodiment, the computer device preprocesses the audio to be synthesized to obtain a syllable sequence of the audio to be synthesized, which may include: inputting the audio to be synthesized and the lyrics text corresponding to the audio to be synthesized into a trained singing alignment model for processing to obtain the first syllable annotation information corresponding to the audio to be synthesized; calling the speech learning software to align the sequence at the boundary of the first syllable annotation sequence of the audio to be synthesized based on the lyrics text corresponding to the audio to be synthesized to obtain the second syllable annotation information of the audio to be synthesized; and performing word embedding representation on the second syllable annotation information of the audio to be synthesized to obtain the syllable sequence of the audio to be synthesized.
[0073] Optionally, the speech learning software may be Praat software.
[0074] Optionally, for audio in different languages, different syllable dictionaries are used when marking syllables. For example, in Mandarin songs and Cantonese songs, different syllable dictionaries are used because the pronunciation of each Chinese character is different.
[0075] See also Figure 3 , Figure 3 is a schematic diagram of a syllable boundary marking provided in an embodiment of the present application, such as Figure 3 The figure is a schematic diagram of the result after a computer device calls Praat software to align the boundaries of syllable annotation information corresponding to an audio.
[0076] In this implementation, the computer device preprocesses the audio to be synthesized to obtain the fundamental frequency label sequence of the audio to be synthesized, which may include: extracting the fundamental frequency of the audio to be synthesized to obtain the fundamental frequency sequence of the audio to be synthesized; and marking the fundamental frequency sequence of the audio to be synthesized to obtain the fundamental frequency label sequence of the audio to be synthesized.
[0077] Optionally, when extracting the baseband of the synthesized audio, the computer device may use the pYin algorithm as a baseband extraction method. Optionally, the computer device may also use the DIO algorithm, the Yin algorithm, or the harvest algorithm in the world vocoder to extract the baseband of the synthesized audio, which is not limited here.
[0078] Optionally, the computer device performs marking processing on the fundamental frequency sequence of the audio to be synthesized to obtain the fundamental frequency marking sequence of the audio to be synthesized, which may include: marking the frame where the fundamental frequency value in the fundamental frequency sequence is not 0 as 1; marking the frame where the fundamental frequency value in the fundamental frequency sequence is 0 as 0; and obtaining the fundamental frequency marking sequence of the audio to be synthesized based on the marked fundamental frequency values.
[0079] In an optional embodiment, before step S201, the computer device also obtains a training sample audio set, which includes sample audios of multiple objects; trains the initialized acoustic model based on each sample audio in the training sample audio set to obtain a trained acoustic model; and uses the resonance peak model in the trained acoustic model as the resonance peak model in the target acoustic model; wherein the initialized acoustic model includes an initialized resonance peak model and an initialized timbre conversion model.
[0080] In an optional embodiment, after the computer device obtains the trained acoustic model, it also includes: fixing the parameters of the resonance peak model in the trained acoustic model, and inputting the sample audio corresponding to the target object into the trained acoustic model, training the timbre conversion model in the trained acoustic model to obtain the trained timbre conversion model; and using the trained timbre conversion model as the timbre conversion model in the target acoustic model.
[0081] S202: Input the formant representation information and pitch information of the audio to be synthesized into a timbre conversion model in the target acoustic model to obtain a mel spectrum feature, where the mel spectrum feature includes the timbre information of the target object.
[0082] The timbre conversion model in the target acoustic model is obtained based on sample audio training of the target object. Optionally, the timbre conversion model can also be regarded as a style transfer module.
[0083] In an optional implementation, the computer device further obtains pitch information of the audio to be synthesized. Optionally, the computer device obtains the pitch information of the audio to be synthesized, which may include: obtaining a music score of the audio to be synthesized; extracting the pitch information of the audio to be synthesized from the music score of the audio to be synthesized. Optionally, the format of the music score includes but is not limited to Musical Instrument Digital Interface (MIDI), Musical XML (Musical Extensible Markup Language), etc.
[0084] In an optional embodiment, a computer device inputs the resonance peak representation information and pitch information of the audio to be synthesized into a timbre conversion model in a target acoustic model to obtain a Mel-spectrogram feature, which may include: inputting the pitch information of the audio to be synthesized into a word embedding module of the pitch information to obtain a word embedding representation of the pitch information corresponding to the audio to be synthesized; inputting the resonance peak representation information and the word embedding representation of the pitch information of the audio to be synthesized into a timbre conversion model in the target acoustic model to obtain a Mel-spectrogram feature.
[0085] S203: Input the Mel spectrum features into a vocoder to obtain a synthesized audio signal.
[0086] That is, the vocoder model is used to convert the Mel-spectrogram features into an audio signal. Optionally, the vocoder can be a Mel-spectrogram Generative Adversarial Network (MelGAN) or a High-Fidelity Generative Adversarial Network (HifiGAN).
[0087] In the embodiment of the present application, the formant model can be used to decouple the timbre information and language of the audio to be synthesized, and obtain the formant representation information that is unrelated to the timbre; the formant representation information is input into the timbre conversion model in the target acoustic model corresponding to the target object, and the Mel spectrum feature including the timbre information of the target object can be obtained; the Mel spectrum feature is input into the vocoder to obtain the synthesized audio signal. It can be seen that the embodiment of the present application can be used to synthesize singing voices of any timbre across languages.
[0088] Optionally, during the process of music production, the synthesized audio signal obtained using the embodiments of the present application may be mixed to obtain a more professional music work.
[0089] See also Figure 4a , Figure 4a is a schematic diagram of the structure of an acoustic model provided in an embodiment of the present application, such as Figure 4a As shown, the acoustic model includes a resonance peak representation model, a timbre conversion model, a word embedding module for pitch information, and a word embedding module for timbre information.
[0090] The formant model is composed of N feed forward transformer blocks (Feed Forward Transformer Block, FFT Block). Optionally, N can be 5. The input of the formant model is the syllable sequence and fundamental frequency mark sequence of the audio, and the output is the formant representation information of the audio. In other words, the formant model is used to learn the formant representation information that is independent of timbre (or called syllable marking information that is independent of timbre) from the syllable sequence and fundamental frequency mark sequence of the audio, which represents the basic pronunciation content in the audio.
[0091] The word embedding module of pitch information is used to process the pitch information of the audio and obtain the word embedding representation of the pitch information of the input audio.
[0092] The word embedding module of the timbre information is used to embed the timbre information indicated by the singer identity (SingerID).
[0093] The timbre conversion model converts the representation features without timbre information (i.e., the resonance peak representation information) into the Mel-spectrogram features of a specific person's timbre. The input of the timbre conversion model is the resonance peak representation information, the pitch information, and the SingerID. The pitch information here is converted from the fundamental frequency sequence of dimension [T, 1] to the logarithmic domain, and then converted to [T, 128] through word embedding. In other words, the pitch information is information processed by the word embedding module of the pitch information. The output of the timbre conversion module is the predicted Mel-spectrogram features of dimension [T, 128]. In other words, the timbre conversion model converts the resonance peak representation information without timbre information into the Mel-spectrogram features with timbre information.
[0094] The timbre conversion model uses a style transfer model similar to styleGAN, which consists of M timbre blocks. Optionally, M can be 6. Each timbre block is composed of a one-dimensional convolution module and an adaptive instance normalization (Adaptive Instance Normalization, AdaIN) module. Among them, the input of the AdaIN module also includes a singer identity (Singer Identity, SingerID), SingerID is used to indicate the singer's timbre information. The role of AdaIN is to align the mean and variance of the audio's resonance peak representation information c (or the audio's content feature c) to the mean and variance of the target object's timbre feature s (or the target object's style feature s). The calculation formula is as follows (1):
[0095]
[0096] In formula (1), c represents the resonance peak representation information of the audio (content feature); s represents the timbre feature (style feature) of the target object; σ(s) represents the variance of the style feature; σ(c) represents the variance of the content feature; μ(s) represents the mean of the style feature; and μ(c) represents the mean of the content feature.
[0097] Figure 4a In the , the minimum mean square error (MSE Loss) and the discriminator error (Discriminator Loss) are used to determine the loss value of the Mel spectrum feature output by the timbre conversion model. Among them, the minimum mean square error is the error between the predicted Mel feature and the target Mel feature; the discriminator error refers to the error in judging the true or false of the predicted Mel spectrum feature.
[0098] The discriminator error is determined based on the discriminator, which is a multi-subband discriminant model. The computer device can divide the 128 subbands of the Mel spectrum feature into three frequency bands: low, medium, and high. The frequency bands of these three frequency bands are [0, 64], [32, 96], and [96, 128], respectively. Each subband is input into the discriminator to judge the authenticity of the Mel spectrum feature.
[0099] See also Figure 4b , Figure 4b is a schematic diagram of the structure of a discriminator provided in an embodiment of the present application. Figure 4b As shown, the discriminator is mainly composed of a 1D convolution module and a residual connection. The computer device can input the Mel spectrum of each subband in the Mel spectrum feature into the 1D convolution module to obtain the result of the 1D convolution, wherein the size of the convolution kernel in the 1D convolution module is 3, and the number of input and output channels is 64; the result of the 1D convolution is processed by the activation function to obtain the processed feature; the processed feature and the Mel spectrum of the subband are added by the residual connection, and after passing through 5 residual connection structures, a linear layer is passed to map the feature to [T, 1]; the discriminant result of the Mel spectrum feature is output based on [T, 1]. The discrimination result includes whether the Mel spectrum feature input into the discriminator is true (Ture) or the Mel spectrum feature input into the discriminator is fake (Fake) (or called synthetic).
[0100] When pre-training the acoustic model, the computer device may adjust the parameters in the resonance peak characterization model, the timbre conversion model, and the discriminator in the acoustic model based on the loss value to obtain the pre-trained acoustic model. Optionally, when the computer device adjusts the parameters in the resonance peak characterization model, the timbre conversion model, and the discriminator in the acoustic model based on the loss value, the learning rate may be 0.001, and the optimizer may be an Adaptive Moment Estimation (Adam) optimizer.
[0101] After the computer device obtains the pre-trained acoustic model, it obtains a trained formant model that is independent of timbre. Since the formant model can decouple timbre and language, it is the main reason for realizing cross-language singing synthesis. For the singing data of the target object, the computer device can fix the parameters of the formant model of the pre-trained model and the word embedding module of the pitch information, and only train the parameters of the timbre conversion model to obtain the acoustic model corresponding to the target object (or called the target acoustic model).
[0102] Based on the acoustic model and vocoder corresponding to the target object, a singing voice synthesis model corresponding to the target object can be obtained. In other words, a complete singing voice synthesis model includes two parts: the acoustic model and the vocoder.
[0103] The process of training the acoustic model corresponding to the target object (ie, the target acoustic model) is described in detail below.
[0104] See also Figure 5a , Figure 5a FIG. 1 is a flow chart of a method for training a target acoustic model provided in an embodiment of the present application. Figure 5a As shown, the method may include but is not limited to the following steps:
[0105] S501: Obtain a training sample audio set.
[0106] The training sample audio set includes sample audios of multiple objects. The sample audios of multiple objects include sample audios in at least two languages. Optionally, the sample audio corresponding to each object may be audio in a single language.
[0107] Taking Mandarin songs and Cantonese songs as examples, the training sample audio set may be a collection of Mandarin songs and Cantonese songs sung by multiple subjects, wherein each subject may sing songs in one language.
[0108] S502: Train the initialized acoustic model based on each sample audio in the training sample audio set to obtain a trained acoustic model.
[0109] Among them, the trained acoustic model includes a resonance peak model, a timbre conversion model, and a word embedding module for pitch information.
[0110] In an optional embodiment, the computer device trains an initialized acoustic model based on each sample audio in the training sample audio set to obtain a trained acoustic model, which may include: obtaining a syllable sequence, a fundamental frequency marker sequence, a pitch information, a timbre information and an acoustic feature of each training sample audio in the training sample audio set, wherein the acoustic feature is a Mel-spectrogram feature; and training the initialized acoustic model using the syllable sequence, the fundamental frequency marker sequence, the pitch information, the timbre information and the acoustic feature of each sample audio to obtain a trained acoustic model.
[0111] In this implementation, the computer device also obtains the lyrics text corresponding to each sample audio, and the computer device obtains the syllable sequence of each sample audio, which may include: inputting each sample audio and the lyrics text corresponding to the sample audio into a pre-trained singing alignment model for processing to obtain the first syllable annotation sequence of each sample audio; calling the speech learning software to align the sequence at the boundary of the first syllable annotation sequence of the sample audio based on the lyrics text corresponding to each sample audio to obtain the second syllable annotation information of each sample audio; embedding the second syllable annotation information of each sample audio to obtain the syllable sequence of each sample audio. Optionally, the dimension of the syllable sequence can be [T, 128].
[0112] Optionally, the speech learning software may be Praat software. Optionally, for audios of different languages, different syllable dictionaries are used when marking syllables. For example, in Mandarin songs and Cantonese songs, since the pronunciation of each Chinese character is different, the syllable dictionaries used are also different.
[0113] Optionally, the computer device processes the syllable annotation information of each sample audio to obtain the syllable sequence of each sample audio, which may include: expanding the syllable annotation information of each sample audio according to the corresponding number of audio frames to obtain the syllable sequence of each sample audio. For example, assuming that the syllable annotation information of a sample audio includes three syllables "na shi wo", and the number of audio frames corresponding to each syllable is [5, 5, 4] respectively, then the three syllables can be expanded into a syllable sequence of 14 frames, namely "na na na na na shi shi shi shi wo wo wo wo". In other words, the syllable sequence is a syllable sequence at the frame level.
[0114] In this implementation, the computer device obtains the fundamental frequency marking sequence of each sample audio, which may include: extracting the fundamental frequency of each sample audio to obtain the fundamental frequency sequence of each sample audio; marking the fundamental frequency sequence of each sample audio to obtain the fundamental frequency marking sequence of each sample audio.
[0115] Optionally, the computer device may extract the baseband of each sample audio using a DIO algorithm, a Yin algorithm, a pYin algorithm, or a harvest algorithm in a world vocoder, which is not limited here.
[0116] Optionally, the computer device performs labeling processing on the fundamental frequency sequence of each sample audio to obtain the fundamental frequency labeling sequence of each sample audio, which may include: marking the frame where the fundamental frequency value in the fundamental frequency sequence of each sample audio is not 0 as 1; marking the frame where the fundamental frequency value in the fundamental frequency sequence is 0 as 0; and obtaining the fundamental frequency labeling sequence of each sample audio based on the marked fundamental frequency value. Optionally, the dimension of the fundamental frequency labeling sequence is [T, 1], where T represents the number of frames. That is, when labeling the fundamental frequency sequence of each sample audio, if the fundamental frequency value of the i-th frame is greater than 0, the value of VUV[i, 1] is 1, otherwise it is 0.
[0117] In this implementation, the computer device obtains the pitch information and timbre information of each sample audio, which may include: obtaining the music score of each sample audio; and extracting the pitch information and timbre information of each sample audio from the music score of each sample audio.
[0118] In this implementation, the acoustic feature is a Mel-spectrogram feature because the frequency range that the human ear can hear is 20-20000 Hz, but the human ear does not have a linear perception relationship with the Hz scale unit, and the use of Mel-spectrogram features is more in line with the working principle of the human ear. At this time, the computer device obtains the acoustic features of each sample audio, that is, obtains the Mel-spectrogram features of each sample.
[0119] Optionally, the computer device obtains the acoustic features of each sample audio, that is, obtains the Mel-spectrogram features of each sample audio, which may include: performing frame windowing processing on each sample audio to obtain each frame windowed sample audio; performing Fourier transform on each frame windowed sample audio to obtain a linear spectrum corresponding to each sample audio; and using a Mel-scaled filter group to process the linear spectrum corresponding to each sample audio to obtain the Mel-spectrogram features of each sample audio.
[0120] In this implementation, the initialized acoustic model includes an initialized formant model, an initialized timbre conversion model, and a word embedding module of pitch information.
[0121] Combine the following Figure 5b In this implementation, the computer device uses the syllable sequence, fundamental frequency marker sequence, pitch information, timbre information and acoustic features of each sample audio to train the initialized acoustic model to obtain the trained acoustic model. Figure 5b , Figure 5b is a flowchart of a method for determining a trained acoustic model provided in an embodiment of the present application, which can also be considered as a sub-flow diagram of a method for training a target acoustic model provided in an embodiment of the present application. Figure 5b As shown, the method includes but is not limited to the following steps:
[0122] S5021. Input the syllable sequence and fundamental frequency marker sequence of each sample audio into the initialized formant model to obtain the formant representation information of each sample audio.
[0123] Among them, the formant representation information of each sample audio does not contain the timbre information of the sample audio. Therefore, the dimension of the formant representation information is the same as the dimension of the syllable sequence. For example, if the dimension of the syllable sequence of any sample audio is [T, 128], and the dimension of the fundamental frequency marker sequence is [T, 1], then after being processed by the formant model, the dimension of the formant representation information of the sample audio is [T, 128], where T represents the number of frames.
[0124] Among them, inputting the syllable sequence and the fundamental frequency marker sequence into the resonance peak model together can implicitly provide consonant information and vowel information, realize the adaptation of the phoneme duration in the syllable, and reduce the problem of mispronunciation.
[0125] See also Figure 6 , Figure 6 Schematic diagram of a resonance peak characterization information provided by an embodiment of the present application. Figure 6 As shown, the total number of frames of the resonance peak is 1200 frames, that is, T=1200. Therefore, the dimension of the resonance peak representation information is [1200, 128].
[0126] S5022: Input the formant representation information, pitch information, and timbre information of each sample audio into the initialized timbre conversion model to obtain the predicted Mel spectrum features of each sample audio.
[0127] The predicted mel spectrum features of each sample audio include timbre information.
[0128] In an optional embodiment, the computer device inputs the resonance peak representation information, pitch information and timbre information of each sample audio into an initialized timbre conversion model to obtain the predicted Mel-spectrogram features of each sample audio, which may include: inputting the pitch information of each sample audio into a word embedding module of the pitch information to obtain a word embedding representation of the pitch information of each sample audio; inputting the timbre information of each sample audio into a word embedding module of the timbre information to obtain a word embedding representation of the timbre information of each sample audio; inputting the resonance peak representation information, the word embedding representation of the pitch information and the word embedding representation of the timbre information of each sample audio into the initialized timbre conversion model to obtain the predicted Mel-spectrogram features of each sample audio.
[0129] S5023. Determine a loss value of each audio sample based on the predicted mel-spectrogram feature of each audio sample.
[0130] In an optional embodiment, the computer device determines the loss value of each sample audio based on the predicted Mel-spectrogram feature of each sample audio, which may include: determining the minimum mean square error between the predicted Mel-spectrogram feature and the acoustic feature of each sample audio, and determining the discriminator error corresponding to the predicted Mel-spectrogram feature of each sample audio; based on the minimum mean square error and the discriminator error, determining the loss value of each sample audio.
[0131] In this implementation, the computer device determines the discriminator error corresponding to the predicted Mel-spectrogram feature of each sample audio, which may include: inputting the predicted Mel-spectrogram feature of each sample audio into an initialized discriminator to obtain a discrimination result of the predicted Mel-spectrogram feature of each sample audio, the discrimination result including whether the predicted Mel-spectrogram feature is a real Mel-spectrogram feature, or the predicted Mel-spectrogram feature is a synthesized Mel-spectrogram feature; based on the discrimination result of the predicted Mel-spectrogram feature of each sample audio, determining the discriminator error corresponding to the predicted Mel-spectrogram feature of each sample audio.
[0132] S5024. When the loss value does not meet the condition for stopping training, the parameters of the initialized acoustic model including the initialized resonance peak model and the initialized timbre conversion model are adjusted according to the loss value to obtain a trained acoustic model.
[0133] In an optional implementation, when the loss value meets the training stop condition, the computer device can obtain a trained acoustic model based on the initialized formant model and the initialized timbre conversion model. The trained acoustic model also includes a word embedding module for pitch information and a word embedding module for timbre information.
[0134] In an optional embodiment, the loss value includes a first loss value and a second loss value; the computer device adjusts the parameters of the initialized acoustic model, including the initialized resonance peak model and the initialized timbre conversion model, according to the loss value to obtain a trained acoustic model, including: adjusting the parameters of the initialized resonance peak model and the initialized timbre conversion model according to the first loss value to obtain an acoustic model after adjusting the parameters, the adjusted acoustic model includes the resonance peak model after adjusting the parameters, the timbre conversion model after adjusting the parameters, and a word embedding module for pitch information; the syllable sequence, fundamental frequency marker sequence and pitch information of each sample audio are input into the acoustic model after adjusting the parameters to determine the second loss value; when the second loss value meets the condition for stopping training, the trained acoustic model is obtained.
[0135] It is understandable that the process of determining the trained acoustic model is a process that is constantly iterating.
[0136] In an optional implementation, when the loss value does not meet the training stop condition, the computer device further adjusts the parameters in the initialized discriminator according to the loss value to obtain a trained discriminator.
[0137] S503, fixing the parameters of the resonance peak model in the trained acoustic model, inputting the sample audio corresponding to the target object into the trained acoustic model, and retraining the timbre conversion model in the trained acoustic model to obtain the trained timbre conversion model.
[0138] For example, assuming that the training sample set includes 5 sample audios of 100 singers, that is, a total of 500 sample audios, the computer device first uses these 500 sample audios to train the initialized acoustic model to obtain a trained acoustic model. At this time, if you want to obtain an acoustic model exclusive to singer 5, the computer device can freeze the resonance peak model and the word embedding module of the pitch information in the trained acoustic model, and then input the 5 sample audios corresponding to singer 5 into the trained acoustic model, train the timbre conversion model in the trained acoustic model, and obtain a trained timbre conversion model.
[0139] Optionally, the process of the computer device training the timbre conversion model in the trained acoustic model can be understood as the process of fine-tuning the trained acoustic model, so as to further improve the performance of the model.
[0140] S504: Determine a target acoustic model based on the trained formant model in the acoustic model and the trained timbre conversion model.
[0141] Among them, the target acoustic model also includes a word embedding module for pitch information and a word embedding module for timbre information.
[0142] The computer device uses the resonance peak model in the trained acoustic model as the resonance peak model in the target acoustic model; and uses the trained timbre conversion model as the timbre conversion model in the target acoustic model.
[0143] That is to say, the resonance peak model in the trained acoustic model, the trained timbre conversion model and the word embedding module of pitch information together constitute a language-independent acoustic model exclusive to singer 5.
[0144] It can be seen that by adopting the embodiment of the present application, the acoustic model corresponding to the target object can be obtained by training the acoustic model. Thus, by using the acoustic model corresponding to the target object in combination with the vocoder, a singer who sings language A can sing songs in language B, for example, a singer who sings Mandarin can sing Cantonese songs, or a singer who sings Cantonese can sing Mandarin songs. Furthermore, on the one hand, it solves the problem that it is difficult for a single singer to cover multiple languages, and on the other hand, it can help ordinary users to sing songs across languages.
[0145] See also Figure 7 , Figure 7 701 is a schematic diagram of a singing voice synthesis device provided in an embodiment of the present application. The singing voice synthesis device described in this embodiment may include a processing unit 701 and a training unit 702.
[0146] The processing unit 701 is used to input the syllable sequence and the fundamental frequency marker sequence of the audio to be synthesized into the formant model in the target acoustic model to obtain the formant representation information of the audio to be synthesized, where the formant representation information is the representation information without timbre information;
[0147] The processing unit 701 is further used to input the formant representation information and pitch information of the audio to be synthesized into the timbre conversion model in the target acoustic model to obtain a mel spectrum feature, where the mel spectrum feature includes the timbre information of the target object, and the timbre conversion model is obtained based on the sample audio training of the target object;
[0148] The processing unit 701 is further configured to input the Mel spectrum feature into the vocoder to obtain a synthesized audio signal.
[0149] In an optional implementation, before the processing unit 701 inputs the syllable sequence and the fundamental frequency marker sequence of the audio to be synthesized into the formant model in the target acoustic model, the training unit 702 is used to:
[0150] Acquire a training sample audio set, where the training sample audio set includes sample audios of multiple objects;
[0151] Training the initialized acoustic model based on each sample audio in the training sample audio set to obtain a trained acoustic model;
[0152] Using the formant model in the trained acoustic model as the formant model in the target acoustic model;
[0153] The initialized acoustic model includes an initialized formant model and an initialized timbre conversion model.
[0154] In an optional implementation, after obtaining the trained acoustic model, the training unit 702 is further configured to:
[0155] Fixing the parameters of the formant model in the trained acoustic model, inputting the sample audio corresponding to the target object into the trained acoustic model, and retraining the timbre conversion model in the trained acoustic model to obtain the trained timbre conversion model;
[0156] The trained timbre conversion model is used as the timbre conversion model in the target acoustic model.
[0157] In an optional implementation, the training unit 702 is used to train the initialized acoustic model based on each sample audio in the training sample audio set to obtain the trained acoustic model, specifically for:
[0158] Obtaining a syllable sequence, a fundamental frequency marker sequence, pitch information, timbre information, and acoustic features of each training sample audio in the training sample audio set, where the acoustic features are Mel-spectrogram features;
[0159] The initialized acoustic model is trained using the syllable sequence, fundamental frequency marker sequence, pitch information, timbre information and acoustic features of each sample audio to obtain a trained acoustic model.
[0160] In an optional implementation, the training unit 702 is used to train the initialized acoustic model using the syllable sequence, fundamental frequency marker sequence, pitch information, timbre information and acoustic features of each sample audio to obtain the trained acoustic model, specifically for:
[0161] Inputting the syllable sequence and fundamental frequency marker sequence of each sample audio into the initialized formant model to obtain the formant representation information of each sample audio;
[0162] Input the formant representation information, pitch information and timbre information of each sample audio into the initialized timbre conversion model to obtain the predicted Mel spectrum features of each sample audio;
[0163] Determine a loss value for each audio sample based on the predicted Mel-spectrogram feature of each audio sample;
[0164] When the loss value does not meet the training stop condition, the parameters of the initialized acoustic model including the initialized resonance peak model and the initialized timbre conversion model are adjusted according to the loss value to obtain a trained acoustic model.
[0165] In an optional implementation, when the training unit 702 is used to determine the loss value of each sample audio based on the predicted mel spectrum feature of each sample audio, it is specifically used to:
[0166] Determine the minimum mean square error between the predicted Mel spectrum feature and the acoustic feature of each sample audio, and determine the discriminator error corresponding to the predicted Mel spectrum feature of each sample audio;
[0167] Based on the minimum mean square error and the discriminator error, the loss value of each sample audio is determined.
[0168] In an optional implementation, when the training unit 702 is used to determine the discriminator error corresponding to the predicted Mel spectrum feature of each sample audio, it is specifically used to:
[0169] Inputting the predicted Mel spectrum feature of each sample audio into the initialized discriminator to obtain a discrimination result of the predicted Mel spectrum feature of each sample audio, wherein the discrimination result includes whether the predicted Mel spectrum feature is a real Mel spectrum feature or a synthesized Mel spectrum feature;
[0170] Based on the discrimination result of the predicted Mel spectrum feature of each sample audio, a discriminator error corresponding to the predicted Mel spectrum feature of each sample audio is determined.
[0171] In an optional implementation, after determining the loss value of each sample audio, the training unit 702 is further configured to:
[0172] When the loss value does not meet the conditions for stopping training, the parameters in the initialized discriminator are adjusted according to the loss value to obtain the trained discriminator.
[0173] It can be understood that the specific implementation of each unit in the singing synthesis device described in the embodiment of the present application and the beneficial effects that can be achieved can be referred to the description of the aforementioned related method embodiments, and will not be repeated here.
[0174] See also Figure 8 , Figure 8 801, a user interface 802, a communication interface 803, and a memory 804. The processor 801, the user interface 802, the communication interface 803, and the memory 804 may be connected via a bus or other means, and the embodiment of the present application takes the connection via a bus as an example.
[0175] Among them, the processor 801 (or CPU (Central Processing Unit)) is the computing core and control core of the computer device, which can parse various instructions in the computer device and process various data of the computer device. For example, the CPU can be used to parse the power on and off instructions sent by the user to the computer device, and control the computer device to perform power on and off operations; for another example, the CPU can transmit various interactive data between the internal structures of the computer device, etc. The user interface 802 is a medium for realizing interaction and information exchange between the user and the computer device. Its specific embodiment can include a display screen (Display) for output and a keyboard (Keyboard) for input, etc. It should be noted that the keyboard here can be either a physical keyboard, a touch screen virtual keyboard, or a keyboard that combines a physical and a touch screen virtual. The communication interface 803 can optionally include a standard wired interface, a wireless interface (such as Wi-Fi, a mobile communication interface, etc.), which is controlled by the processor 801 for sending and receiving data.
[0176] The memory 804 (Memory) is a memory device in the computer device, which is used to store programs and data. It is understandable that the memory 804 here can include both the built-in memory of the computer device and the extended memory supported by the computer device. The memory 804 provides a storage space, which stores the operating system of the computer device, including but not limited to: Android system, iOS system, Windows Phone system, etc., which is not limited in this application.
[0177] In the embodiment of the present application, the processor 801 may execute the following operations by running the executable program code in the memory 804:
[0178] Inputting the syllable sequence and the fundamental frequency marker sequence of the audio to be synthesized into the formant model in the target acoustic model to obtain the formant representation information of the audio to be synthesized, where the formant representation information is representation information without timbre information;
[0179] Input the formant representation information and pitch information of the audio to be synthesized into the timbre conversion model in the target acoustic model to obtain the Mel spectrum feature, which includes the timbre information of the target object. The timbre conversion model is obtained based on the sample audio training of the target object;
[0180] The Mel-spectrogram features are input into the vocoder to obtain the synthesized audio signal.
[0181] In an optional implementation, before inputting the syllable sequence and the fundamental frequency marker sequence of the audio to be synthesized into the formant model in the target acoustic model, the processor 801 further executes:
[0182] Acquire a training sample audio set, where the training sample audio set includes sample audios of multiple objects;
[0183] Training the initialized acoustic model based on each sample audio in the training sample audio set to obtain a trained acoustic model;
[0184] Using the formant model in the trained acoustic model as the formant model in the target acoustic model;
[0185] The initialized acoustic model includes an initialized formant model and an initialized timbre conversion model.
[0186] In an optional implementation, after the processor 801 trains the initialized acoustic model based on each sample audio in the training sample audio set to obtain the trained acoustic model, it further executes:
[0187] Fixing the parameters of the formant model in the trained acoustic model, inputting the sample audio corresponding to the target object into the trained acoustic model, and retraining the timbre conversion model in the trained acoustic model to obtain the trained timbre conversion model;
[0188] The trained timbre conversion model is used as the timbre conversion model in the target acoustic model.
[0189] In an optional implementation, when the processor 801 trains the initialized acoustic model based on each sample audio in the training sample audio set to obtain the trained acoustic model, the processor 801 specifically performs:
[0190] Obtaining a syllable sequence, a fundamental frequency marker sequence, pitch information, timbre information, and acoustic features of each training sample audio in the training sample audio set, where the acoustic features are Mel-spectrogram features;
[0191] The initialized acoustic model is trained using the syllable sequence, fundamental frequency marker sequence, pitch information, timbre information and acoustic features of each sample audio to obtain a trained acoustic model.
[0192] In an optional implementation, when the processor 801 trains the initialized acoustic model using the syllable sequence, fundamental frequency marker sequence, pitch information, timbre information, and acoustic features of each sample audio to obtain the trained acoustic model, the processor 801 specifically performs:
[0193] Inputting the syllable sequence and fundamental frequency marker sequence of each sample audio into the initialized formant model to obtain the formant representation information of each sample audio;
[0194] Input the formant representation information, pitch information and timbre information of each sample audio into the initialized timbre conversion model to obtain the predicted Mel spectrum features of each sample audio;
[0195] Determine a loss value for each audio sample based on the predicted Mel-spectrogram feature of each audio sample;
[0196] When the loss value does not meet the training stop condition, the parameters of the initialized acoustic model including the initialized resonance peak model and the initialized timbre conversion model are adjusted according to the loss value to obtain a trained acoustic model.
[0197] In an optional implementation, when the processor 801 determines the loss value of each sample audio based on the predicted mel spectrum feature of each sample audio, it specifically performs:
[0198] Determine the minimum mean square error between the predicted Mel spectrum feature and the acoustic feature of each sample audio, and determine the discriminator error corresponding to the predicted Mel spectrum feature of each sample audio;
[0199] Based on the minimum mean square error and the discriminator error, the loss value of each sample audio is determined.
[0200] In an optional implementation, when determining the discriminator error corresponding to the predicted Mel-spectrogram feature of each sample audio, the processor 801 specifically performs:
[0201] Inputting the predicted Mel spectrum feature of each sample audio into the initialized discriminator to obtain a discrimination result of the predicted Mel spectrum feature of each sample audio, wherein the discrimination result includes whether the predicted Mel spectrum feature is a real Mel spectrum feature or a synthesized Mel spectrum feature;
[0202] Based on the discrimination result of the predicted Mel spectrum feature of each sample audio, a discriminator error corresponding to the predicted Mel spectrum feature of each sample audio is determined.
[0203] In an optional implementation, after determining the loss value of each audio sample, the processor 801 further executes:
[0204] When the loss value does not meet the conditions for stopping training, the parameters in the initialized discriminator are adjusted according to the loss value to obtain the trained discriminator.
[0205] In a specific implementation, the processor 801, user interface 802, communication interface 803 and memory 804 described in the embodiments of the present application can execute the implementation method of the computer device described in the singing synthesis method provided in the embodiments of the present application, and can also execute the implementation method described in the singing synthesis device provided in the embodiments of the present application, which will not be repeated here.
[0206] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. The computer program includes program instructions. When the program instructions are executed by a processor, the singing synthesis method provided by the embodiment of the present application is implemented. For details, please refer to the implementation methods provided in the above steps, which will not be repeated here.
[0207] The embodiment of the present application also provides a computer program product or a computer program, wherein the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method as described in the embodiment of the present application. The specific implementation method can be referred to the above description, which will not be repeated here.
[0208] It should be noted that, for the above-mentioned various method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0209] A person skilled in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0210] The above disclosure is only part of the embodiments of the present application, which certainly cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.
Claims
1. A singing voice synthesis method, characterized in that: The method comprises: Inputting the syllable sequence and the fundamental frequency marker sequence of the audio to be synthesized into the formant model in the target acoustic model to obtain the formant representation information of the audio to be synthesized, wherein the formant representation information is representation information without timbre information; the formant model is obtained by training an initialized acoustic model based on sample audios of multiple objects in the acquired training sample audio set, wherein the formant model in the trained acoustic model is used as the formant model in the target acoustic model, and the initialized acoustic model includes an initialized formant model and an initialized timbre conversion model; Inputting the formant representation information and pitch information of the audio to be synthesized into the timbre conversion model in the target acoustic model to obtain a mel-spectrogram feature, wherein the synthesized mel-spectrogram feature includes the timbre information of the target object, and the timbre conversion model is obtained by training based on the sample audio of the target object; the timbre conversion model in the target acoustic model is obtained by training based on the sample audio of the target object in the training sample audio set; The Mel spectrum features are input into a vocoder to obtain a synthesized audio signal.
2. The method according to claim 1, characterized in that Before inputting the syllable sequence and the fundamental frequency marker sequence of the audio to be synthesized into the formant model in the target acoustic model, the method further comprises: Acquire a training sample audio set, wherein the training sample audio set includes sample audios of multiple objects; Training the initialized acoustic model based on each sample audio in the training sample audio set to obtain a trained acoustic model; Using the formant model in the trained acoustic model as the formant model in the target acoustic model; The initialized acoustic model includes an initialized formant model and an initialized timbre conversion model.
3. The method according to claim 2, characterized in that After obtaining the trained acoustic model, the method further includes: Fixing the parameters of the formant model in the trained acoustic model, inputting the sample audio corresponding to the target object into the trained acoustic model, training the timbre conversion model in the trained acoustic model, and obtaining the trained timbre conversion model; The trained timbre conversion model is used as the timbre conversion model in the target acoustic model.
4. The method according to claim 2, characterized in that: The step of training the initialized acoustic model based on each sample audio in the training sample audio set to obtain a trained acoustic model includes: Acquire a syllable sequence, a fundamental frequency marker sequence, pitch information, timbre information, and an acoustic feature of each training sample audio in the training sample audio set, wherein the acoustic feature is a Mel-spectrogram feature; The initialized acoustic model is trained using the syllable sequence, fundamental frequency marker sequence, pitch information, timbre information and acoustic features of each sample audio to obtain the trained acoustic model.
5. The method according to claim 4, characterized in that The method of training the initialized acoustic model by using the syllable sequence, fundamental frequency marker sequence, pitch information, timbre information and acoustic features of each sample audio to obtain the trained acoustic model comprises: Inputting the syllable sequence and fundamental frequency marker sequence of each sample audio into the initialized formant model to obtain the formant representation information of each sample audio; Inputting the formant representation information, pitch information and timbre information of each sample audio into an initialized timbre conversion model to obtain the predicted Mel spectrum features of each sample audio; Determine a loss value of each audio sample based on the predicted Mel spectrum feature of each audio sample; When the loss value does not meet the condition for stopping training, the parameters of the initialized resonance peak model and the initialized timbre conversion model included in the initialized acoustic model are adjusted according to the loss value to obtain the trained acoustic model.
6. The method according to claim 5, characterized in that The step of determining the loss value of each audio sample based on the predicted Mel spectrum feature of each audio sample includes: Determine the minimum mean square error between the predicted Mel spectrum feature and the acoustic feature of each sample audio, and determine the discriminator error corresponding to the predicted Mel spectrum feature of each sample audio; Based on the minimum mean square error and the discriminator error, a loss value of each sample audio is determined.
7. The method according to claim 6, characterized in that The determining of the discriminator error corresponding to the predicted Mel spectrum feature of each sample audio includes: Inputting the predicted Mel-spectrogram feature of each sample audio into an initialized discriminator to obtain a discrimination result of the predicted Mel-spectrogram feature of each sample audio, wherein the discrimination result includes whether the predicted Mel-spectrogram feature is a real Mel-spectrogram feature or whether the predicted Mel-spectrogram feature is a synthesized Mel-spectrogram feature; Based on the discrimination result of the predicted Mel spectrum feature of each sample audio, a discriminator error corresponding to the predicted Mel spectrum feature of each sample audio is determined.
8. The method according to any one of claims 5 to 7, characterized in that: After determining the loss value of each sample audio, the method further includes: When the loss value does not meet the training stop condition, the parameters in the initialized discriminator are adjusted according to the loss value to obtain a trained discriminator.
9. A computer device, characterized in that: include: A processor, a communication interface and a memory, wherein the processor, the communication interface and the memory are connected to each other, wherein the memory stores an executable program code, and the processor is used to call the executable program code to execute the method as described in any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor is caused to perform the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Voice processing method and device, equipment, medium and program product
CN113178187A
Singing synthesis method and device, computer device and storage medium
CN113555001A