Vocal conversion model training methods, song timbre conversion methods and related products
By training a singing voice conversion model, text features are used to replace timbre features, solving the problem of finding or recording example recordings in existing technologies and achieving efficient singing voice conversion.
Patent Information
- Application Number
- CN202411636646.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-15
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2044-11-15
AI Technical Summary
Existing singing voice conversion technology requires users to find or record ready-made recordings with specified timbres, resulting in low singing voice conversion efficiency.
By acquiring sample songs and vocal description texts of the singers, the audio encoding module and text encoding module in the singing conversion model are used for training. This allows text features to replace vocal features, enabling the training of the singing conversion model and reducing reliance on sample recordings.
It allows for voice conversion without requiring users to pre-record or find sample recordings, significantly improving the efficiency of vocal conversion.
Smart Images

Figure CN119673185B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of vocal synthesis technology, and in particular to a vocal conversion model training method, a song timbre conversion method, a computer device, a computer-readable storage medium, and a computer program product. Background Technology
[0002] With the development of computer technology, vocal conversion technology has become increasingly popular. Vocal conversion is a process of converting singing voices in audio from one sound source to another target sound source, while maintaining the content and rhythm of the song.
[0003] In related technologies, vocal conversion involves obtaining characteristic information representing timbre from example recordings and then using this information to control the vocal conversion process. However, while effective, this method requires users to first find or record existing recordings with the specified timbre, reducing the efficiency of vocal conversion. Summary of the Invention
[0004] Therefore, it is necessary to provide a singing voice conversion model training method, a song timbre conversion method, a computer device, a computer-readable storage medium, and a computer program product to address the above-mentioned technical problems.
[0005] Firstly, this application provides a method for training a singing voice conversion model, including:
[0006] Obtain sample songs and corresponding vocal description text of the singers whose voices are paired with the sample songs;
[0007] The timbre features of the sample song are obtained by the audio encoding module in the singing conversion model, and the text features corresponding to the timbre description text of the singer are obtained by the text encoding module in the singing conversion model.
[0008] Based on the difference between the timbre features and the text features, the model parameters of the singing voice conversion model are adjusted to obtain a trained singing voice conversion model. The trained singing voice conversion model is used to extract the text features corresponding to the input timbre description text through the text encoding module, and to obtain the target timbre features based on the text features corresponding to the timbre description text, and to perform song timbre conversion.
[0009] In one embodiment, adjusting the model parameters of the singing voice conversion model based on the difference between the timbre features and the text features to obtain a trained singing voice conversion model includes:
[0010] Based on the text features and the pronunciation information of the sample song, obtain the phoneme pronunciation feature distribution information of the sample song;
[0011] The audio feature distribution information used to decode the sample song is obtained, and the distribution mapping module in the singing conversion model performs a distribution mapping between the phoneme pronunciation feature distribution information and the audio feature distribution information to obtain the distribution mapping result;
[0012] Based on the differences between the timbre features and the text features, and the differences between the distribution mapping results and the reference distribution mapping results, the model parameters of the singing voice conversion model are adjusted to obtain a trained singing voice conversion model.
[0013] In one embodiment, adjusting the model parameters of the singing conversion model based on the difference between the timbre features and the text features, and the difference between the distribution mapping result and the reference distribution mapping result, includes:
[0014] The distribution decoding module in the singing conversion model decodes the corresponding predicted song based on the audio feature distribution information.
[0015] A first loss value is determined based on the difference between the timbre features and the text features; a second loss value is determined based on the difference between the distribution mapping result and the reference distribution mapping result; and a third loss value is determined based on the difference between the predicted song and the sample song.
[0016] The model loss value is determined based on the first loss value, the second loss value, and the third loss value, and the model parameters of the singing conversion model are adjusted based on the model loss value.
[0017] In one embodiment, the step of decoding the corresponding predicted song by the distribution decoding module in the singing conversion model based on the audio feature distribution information includes:
[0018] The audio feature distribution information, the text features, and the pitch information of the sample song are input into the distribution decoding module in the singing conversion model, and the predicted song is obtained based on the output of the distribution decoding module.
[0019] In one embodiment, the reference distribution mapping result includes the phoneme pronunciation feature distribution information or the audio feature distribution information;
[0020] The difference between the distribution mapping result and the reference distribution mapping result is determined through the following steps:
[0021] If the distribution mapping result is the audio feature distribution mapping result obtained by the distribution mapping module through forward mapping of the audio feature distribution information, then the difference between the audio feature distribution mapping result and the phoneme pronunciation feature distribution information is determined.
[0022] If the distribution mapping result is the phoneme feature distribution mapping result obtained by the distribution mapping module through reverse mapping of the phoneme pronunciation feature distribution information, then the difference between the phoneme feature distribution mapping result and the audio feature distribution information is determined.
[0023] In one embodiment, the pronunciation information of the sample song includes the pitch information of the sample song and the phoneme features of the sample song;
[0024] The step of obtaining the phoneme pronunciation feature distribution information of the sample song based on the text features and the pronunciation information of the sample song includes:
[0025] The text features, the pitch information of the sample song, and the phoneme features of the sample song are input into the text encoding module to obtain the phoneme prior distribution output by the text encoding module;
[0026] The prior distribution of phonemes is used as the phoneme pronunciation feature distribution information of the sample song.
[0027] Secondly, this application also provides a method for converting the timbre of a song, including:
[0028] Obtain the song to be converted, as well as timbre configuration information, including at least a timbre description text;
[0029] The pronunciation information and timbre configuration information of the song to be converted are input into the trained singing voice conversion model to obtain the song with timbre conversion output by the singing voice conversion model.
[0030] The singing voice conversion model is trained according to the singing voice conversion model training method described in any of the above items.
[0031] In one embodiment, the timbre configuration information also includes source audio with a specified timbre;
[0032] The step of inputting the pronunciation information and timbre configuration information of the song to be converted into a trained singing voice conversion model to obtain the timbre-converted song output by the singing voice conversion model includes:
[0033] The text encoding module in the singing conversion model extracts features from the timbre description text to obtain the text features of the timbre description text, and the audio encoding module in the singing conversion model extracts features from the source audio to obtain the timbre features of the source audio;
[0034] Based on the feature weights of the text features of the timbre description text and the timbre features of the source audio, feature fusion is performed on the text features of the timbre description text and the timbre features of the source audio to obtain the target timbre features;
[0035] Based on the target timbre characteristics and the pronunciation information of the song to be converted, the song with converted timbre is output.
[0036] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the singing voice conversion model training method as described in any of the preceding claims or the steps of the song timbre conversion method as described in any of the preceding claims.
[0037] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the singing voice conversion model training method or the steps of the song timbre conversion method as described in any of the preceding claims.
[0038] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the singing voice conversion model training method as described in any of the preceding claims or the steps of the song timbre conversion method as described in any of the preceding claims.
[0039] The aforementioned singing voice conversion model training method, song timbre conversion method, computer equipment, computer-readable storage medium, and computer program product train the audio encoding module and text encoding module in the singing voice conversion model using paired sample songs and singer timbre descriptions. This enables the text features output by the text encoding module to be as close as possible to the timbre features output by the audio encoding module. Thus, after training, text features can be used to replace timbre features. When timbre conversion is needed, only any timbre description text related to the target timbre needs to be input, eliminating the need for users to record or search for example recordings in advance, effectively improving the efficiency of singing voice conversion. Attached Figure Description
[0040] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0041] Figure 1 This is a flowchart illustrating a singing voice conversion model training method in one embodiment;
[0042] Figure 2 This is a schematic diagram of the structure of a singing voice conversion model in one embodiment;
[0043] Figure 3This is a flowchart illustrating a song timbre conversion method in one embodiment;
[0044] Figure 4a This is a flowchart illustrating another song timbre conversion method in one embodiment;
[0045] Figure 4b This is a flowchart illustrating another song timbre conversion method in one embodiment;
[0046] Figure 4c This is a flowchart illustrating another song timbre conversion method in one embodiment;
[0047] Figure 5 This is an internal structural diagram of a computer device in one embodiment;
[0048] Figure 6 This is an internal structural diagram of another computer device in one embodiment. Detailed Implementation
[0049] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0050] To enable those skilled in the art to better understand this application, the relevant technologies are described below.
[0051] Vocal conversion is a process of transforming singing sounds in audio from one sound source to another target sound source while preserving the song's content and rhythm. In related technologies, feature information representing timbre characteristics is obtained from example recordings during vocal conversion, and this information is then used to control the conversion process. However, while effective, this method requires the user to first find or record an existing recording with the specified timbre. To alleviate this problem, the model can decouple the relationship between audio, singing style, and the singer, breaking down multiple dimensions. However, this method requires complex design and meticulous model training, and model iteration and maintenance are costly. Therefore, related technologies suffer from low efficiency in vocal conversion.
[0052] Based on this, this application provides a singing voice conversion model training method, a song timbre conversion method, a computer device, a computer-readable storage medium, and a computer program product to at least solve the above-mentioned technical problems.
[0053] In one embodiment, Figure 1As shown, a method for training a singing voice conversion model is provided. This embodiment illustrates the application of this method to a terminal. It is understood that this method can also be applied to a server, and further to a system including both a terminal and a server, and is implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps:
[0054] S101, Obtain the sample song and the vocal timbre description text of the singer paired with the sample song.
[0055] The sample songs are songs containing vocals, which can be sung by a real person or synthesized using vocal synthesis technology. In some examples, the sample songs can be recorded as audio files, such as WAV audio files. The sample songs can be dry vocal data containing the singer's voice, and can be obtained by removing noise and blank segments from the audio of the song containing the singer's voice.
[0056] A singer's timbre description text can be text information used to describe the characteristics of a singer's timbre, such as "sweet girl's voice" or "mature baritone."
[0057] In this step, multiple training samples can be obtained for training the singing conversion model. Each training sample includes a sample song and a vocal timbre description text paired with the sample song. The vocal timbre description text paired with the sample song refers to the text used to describe the vocal timbre characteristics of the singer corresponding to the sample song.
[0058] S102, the timbre features of the sample song are obtained by the audio encoding module in the singing conversion model, and the text features corresponding to the timbre description text of the singer are obtained by the text encoding module in the singing conversion model.
[0059] After obtaining the sample song and the paired singer's timbre description text, on the one hand, the sample song can be input into the audio encoding module of the singing conversion model, and the audio encoding module can obtain the timbre features of the sample song. For example, the Mel spectrum corresponding to the sample song can be extracted and input into the Mel encoder of the audio encoding module, and the output can be used as the timbre features. On the other hand, the singer's timbre description text can be input into the text encoding module of the singing conversion model, and the text encoding module can obtain the text features corresponding to the singer's timbre description text. In some examples, the singer's timbre description text can be preprocessed by a pre-trained language model, and the preprocessed information can be input into the text encoding module for feature extraction.
[0060] S103, Based on the difference between timbre features and text features, adjust the model parameters of the singing conversion model to obtain a trained singing conversion model; the trained singing conversion model is used to extract the text features corresponding to the input timbre description text through the text encoding module, and obtain the target timbre features based on the text features corresponding to the timbre description text to perform song timbre conversion.
[0061] Specifically, this application proposes to add a text-controlled timbre function to the singing conversion model, that is, to control the timbre used in the singing conversion process through the singer's timbre description text. In this regard, after obtaining the timbre features and text features, the difference between the timbre features and text features can be determined, and the model parameters of the singing conversion model can be adjusted according to the difference, so that for the same sample song, the timbre features output by the audio encoding module and the text features output by the text encoding module can be as close as possible. Thus, the text features corresponding to the singer's timbre description text can be used as the singer's timbre features, replacing the timbre features extracted from the sample song itself.
[0062] To address this, the model parameters of the singing conversion model can be adjusted based on the differences between timbre features and text features. In some embodiments, the module parameters of other modules in the singing conversion model, except for the audio encoding module and the text encoding module, can be fixed, while the module parameters of the audio encoding module and the text encoding module are adjusted according to the differences between timbre features and text features. When the training termination condition is met, a trained singing conversion model can be obtained.
[0063] When the singing conversion model is trained, since the timbre features and text features output by the audio encoding module and the text encoding module for the paired sample songs and singer timbre description texts are already the same or quite similar, the timbre description texts can be input into the trained singing conversion model. The model then extracts the text features corresponding to the input timbre description texts through the text encoding module, and obtains the target timbre features based on the text features corresponding to the timbre texts, and performs song timbre conversion.
[0064] In the above-mentioned singing voice conversion model training method, by using paired sample songs and singer timbre descriptions to train the audio encoding module and text encoding module in the singing voice conversion model, the text features output by the text encoding module can be made as close as possible to the timbre features output by the audio encoding module. Thus, after training, the text features can be used to replace the timbre features. When timbre conversion is needed, only any timbre description text related to the target timbre needs to be input, without the need for the user to record or find sample recordings in advance, effectively improving the efficiency of singing voice conversion.
[0065] In some exemplary embodiments, in step S103, according to the differences between the timbre features and the text features, adjusting the model parameters of the singing voice conversion model to obtain a trained singing voice conversion model may include the following steps:
[0066] S1031. Obtain the phoneme pronunciation feature distribution information of the sample song according to the text features and the pronunciation information of the sample song.
[0067] In specific implementation, the pronunciation information corresponding to the sample song can be obtained. This pronunciation information is information that is independent of the singer and related to the content of the sample song. In some examples, the pronunciation information of the sample song may include information related to the pitch of the sample song and the phonemes corresponding to the lyrics text. Here, a phoneme is the smallest unit of Chinese pronunciation. For example, the pronunciation units corresponding to the character "好" include "h" and "ao"; pitch is also called fundamental frequency. When a sounding body vibrates to emit sound, the sound can generally be decomposed into many simple sine waves. That is to say, all natural sounds can be understood as composed of many sine waves with different frequencies. Among them, the sine wave with the lowest frequency is the fundamental tone, and the other sine waves with higher frequencies are overtones.
[0068] Since there are differences in the phoneme pronunciation methods of different singers when singing songs, in this step, after obtaining the text features of the singer's timbre description text, the text features and the pronunciation information of the sample song can be fused, and the phoneme pronunciation feature distribution information of the sample song can be obtained according to the fusion result. Among them, the phoneme pronunciation feature distribution information can reflect the distribution of various phoneme pronunciation features. In some embodiments, the phoneme pronunciation feature distribution information may be the prior distribution of phoneme pronunciation features.
[0069] S1032. Obtain the audio feature distribution information for decoding the sample song, and perform distribution mapping on the phoneme pronunciation feature distribution information and the audio feature distribution information by the distribution mapping module in the singing voice conversion model to obtain a distribution mapping result.
[0070] In practical applications, embodiments of the present application can also obtain the audio feature distribution information for decoding the sample song. Among them, the audio feature distribution information can characterize the distribution of audio features. It can be understood as high-dimensional features obtained from the latent space. The audio feature distribution information can also be called the latent variable of the sample song. In some embodiments, the audio feature distribution information includes the posterior distribution of audio features. By decoding the encoded audio feature distribution information, the waveform corresponding to the singing voice signal can be obtained.
[0071] In one embodiment, obtaining audio feature distribution information for decoding a sample song may include the following steps: obtaining a spectrogram corresponding to the sample song; inputting the spectrogram into a posterior distribution encoder to obtain the posterior distribution of audio features corresponding to the sample song output by the posterior distribution encoder; and using the posterior distribution of audio features as the audio feature distribution information of the sample song.
[0072] In practice, a spectrogram, such as a linear-scale spectrogram, can be obtained for the sample song. This spectrogram can then be input into the posterior distribution encoder in the song conversion model to obtain the posterior distribution of the audio features corresponding to the sample song, which is used as the audio feature distribution information of the sample song.
[0073] Although both phoneme articulation feature distribution information and audio feature distribution information are distribution information, phoneme articulation features are mainly related to the timbre of a specific singer and the phonemes, pitch, and other articulation information involved in the sample song. Phoneme articulation features are relatively stable and simple, while audio features encompass multiple dimensions of information in the audio signal, such as spectral characteristics, time-domain characteristics, and cepstral characteristics. These information are numerous and have complex interrelationships. Due to the high complexity and diversity of audio signals themselves, the distribution of audio features also exhibits greater complexity. It can be understood that there is a certain difference in complexity between phoneme articulation feature distribution information and audio feature distribution information; audio feature distribution information is more complex than phoneme articulation feature distribution information.
[0074] Since audio feature distribution information can decode sample songs, after generating phoneme pronunciation feature distribution information of the sample song based on text features and pronunciation information, the song conversion model can attempt to construct a distribution mapping relationship between the phoneme pronunciation feature distribution information and the audio feature distribution information through the distribution mapping module in the model. This distribution mapping relationship allows for distribution mapping between the phoneme pronunciation feature distribution information and the audio feature distribution information, yielding the distribution mapping result. In some exemplary embodiments, the distribution mapping module can be a processing module using normalizing flow techniques, such as a flow model. The distribution mapping module transforms a simple prior distribution into a more complex posterior distribution through a reversible transformation, thereby improving the expressive power of the prior distribution and better representing the distribution characteristics of real samples.
[0075] It is understandable that when the module parameters of the distribution mapping module are appropriate, the distribution mapping module can correctly map simple phoneme pronunciation feature distribution information into more complex audio feature distribution information, thereby using the mapped audio feature distribution information to decode the corresponding song.
[0076] S1033. Based on the differences between timbre features and text features, as well as the differences between the distribution mapping results and the reference distribution mapping results, adjust the model parameters of the singing voice conversion model to obtain a trained singing voice conversion model.
[0077] Specifically, embodiments of this application can also obtain a reference distribution mapping result, which can be a predetermined and correct distribution mapping. By comparing the distribution mapping result output by the distribution mapping module in the singing conversion model with the reference distribution mapping result, it can be determined whether the current module parameters of the distribution mapping module are appropriate, thereby adjusting the singing conversion model accordingly.
[0078] Then, based on the differences between timbre features and text features, as well as the differences between the distribution mapping results and the reference distribution mapping results, the model parameters of the singing conversion model are adjusted. For example, based on the differences between timbre features and text features, as well as the differences between the distribution mapping results and the reference distribution mapping results, the model loss value is determined. The model parameters are iteratively adjusted based on the model loss value. When the training termination condition is met, the trained singing conversion model is obtained.
[0079] In some optional embodiments, the modules in the singing conversion model can be trained simultaneously. For example, the text encoding module, audio encoding module, and distribution mapping module can be trained simultaneously. Alternatively, different modules in the singing conversion model can be trained sequentially. For instance, the module parameters of all modules except the timbre encoding module and the text encoding module can be fixed first. Then, based on the difference between timbre features and text features, the module parameters of the text encoding module and the audio encoding module can be adjusted. When the training termination condition is met, the timbre encoding module and the text encoding module can be fixed again, and the other modules can be trained. It can be understood that, whether the modules are trained simultaneously or sequentially, the model parameters of the singing conversion model can be adjusted based on the difference between timbre features and text features, as well as the difference between the distribution mapping result and the reference distribution mapping result.
[0080] In this embodiment, based on existing text features and sample song pronunciation information, a distribution mapping relationship can be constructed between the distribution mapping of simple phoneme pronunciation features corresponding to the singer and the distribution of audio features. This helps the singing conversion model combine the user-input timbre description text to generate more natural and expressive singing audio, enhancing the diversity and naturalness of the singing conversion process.
[0081] In one embodiment, step S1033, adjusting the model parameters of the singing conversion model based on the differences between timbre features and text features, and the differences between the distribution mapping result and the reference distribution mapping result, may include the following steps:
[0082] The distribution decoding module in the song conversion model decodes the corresponding predicted song based on the audio feature distribution information; the first loss value is determined based on the difference between timbre features and text features; the second loss value is determined based on the difference between the distribution mapping result and the reference distribution mapping result; and the third loss value is determined based on the difference between the predicted song and the sample song. The model loss value is determined based on the first, second, and third loss values, and the model parameters of the song conversion model are adjusted based on the model loss value.
[0083] In one embodiment, Figure 2 The diagram illustrates the structure of a vocal conversion model, which includes a text encoder and an audio encoder for encoding the singer's timbre features, a stream model as a distributed mapping module, and a distributed encoder and a distributed decoder. For paired WAV format vocal dry data (i.e., sample songs) and singer timbre description text, the corresponding timbre features and text features can be obtained respectively through the audio encoder and text encoder.
[0084] Since the audio waveform is decoded using audio feature distribution information in this embodiment, the correct decoding of the audio feature distribution information will affect the conversion effect of the song's timbre. Therefore, the distribution decoding module in the song conversion model can decode the corresponding predicted song based on the audio feature distribution information, then determine the first loss value based on the difference between the timbre features and the text features, determine the second loss value based on the difference between the distribution mapping result and the reference distribution mapping result, and determine the third loss value based on the difference between the predicted song and the sample song.
[0085] Specifically, the first loss value is positively correlated with the difference between timbre features and text features; the second loss value is positively correlated with the difference between the distribution mapping result and the reference distribution mapping result; and the difference between the predicted song and the sample song is positively correlated with the third loss value. In some exemplary embodiments, the first loss value for timbre features and text features can be determined based on a regression loss function (such as the mean squared error loss function (MSELoss)), and the difference between the predicted song and the sample song can be determined based on the difference between the audio waveform of the predicted song and the audio waveform of the sample song.
[0086] Furthermore, the model loss value can be determined based on the first loss value, the second loss value, and the third loss value. The model parameters of the singing conversion model can be adjusted based on the model loss value, for example, based on the backpropagation algorithm, the model parameters can be adjusted based on the model loss value.
[0087] In this embodiment, by determining the model loss value based on the first loss value, the second loss value, and the third loss value, on the one hand, the features output by the text encoding module and the audio encoding module can be made as close as possible, so that the text features output by the text encoding module can replace the timbre features output by the audio encoding module in the subsequent process. On the other hand, the distribution mapping relationship between the phoneme feature distribution information and the audio feature distribution information can be effectively constructed, so that a complex audio feature distribution can be fitted based on a simple phoneme feature distribution, thereby improving the naturalness and detail of the synthesized singing voice. Furthermore, the decoding accuracy of the distribution decoder can be effectively improved, making the decoding result more accurate.
[0088] In one embodiment, the distribution decoding module in the singing conversion model decodes the corresponding predicted song based on audio feature distribution information, which may include the following steps:
[0089] The audio feature distribution information, text features, and pitch information of the sample song are input into the distribution decoding module in the singing conversion model, and the predicted song is obtained based on the output of the distribution decoding module.
[0090] In practical applications, when singers with different timbres perform the same song, the audio waveforms will exhibit some differences. To address this, in this embodiment, in addition to inputting audio feature distribution information into the distribution decoding module, text features and the pitch information of the sample song can also be input together into the distribution decoding module. The distribution decoding module then combines these three types of information to decode the predicted song. By inputting text features and the pitch information of the sample song together into the distribution decoding module, the feature information related to the singer's timbre and the song's pitch can be enhanced during the decoding process, helping to improve the correspondence between the decoded predicted song and the singer's timbre and the song's pitch.
[0091] In one embodiment, the reference distribution mapping result includes phoneme pronunciation feature distribution information or audio feature information; in step S1033, the difference between the distribution mapping result and the reference distribution mapping result can be determined through the following steps:
[0092] If the distribution mapping result is the audio feature distribution mapping result obtained by the distribution mapping module through forward mapping of the audio feature distribution information, then the difference between the audio feature distribution mapping result and the phoneme pronunciation feature distribution information is determined; if the distribution mapping result is the phoneme feature distribution mapping result obtained by the distribution mapping module through reverse mapping of the phoneme pronunciation feature distribution information, then the difference between the phoneme feature distribution mapping result and the audio feature distribution information is determined.
[0093] Specifically, the audio feature distribution mapping result is obtained by the distribution mapping module through forward mapping of the audio feature distribution information; the phoneme feature distribution mapping result is obtained by the distribution mapping module through reverse mapping of the phoneme pronunciation feature distribution information.
[0094] In practical applications, the distribution mapping module in the singing conversion model is a reversible mapping module. That is, the distribution mapping module can map phoneme pronunciation feature distribution information to audio feature distribution information, and also map audio feature distribution information to phoneme pronunciation feature distribution information.
[0095] After obtaining the phoneme pronunciation feature distribution information of the sample song based on text features and pronunciation information, distribution mapping can be performed as needed. Specifically, if the distribution mapping module performs a forward mapping on the audio distribution feature information, that is, maps the audio feature distribution information Za to Zt', then the mapped phoneme pronunciation feature distribution information can be used as the distribution mapping result, and the phoneme pronunciation feature distribution information Zt generated based on text features and pronunciation information can be used as the reference distribution mapping result. The difference between the distribution mapping result Zt' and the phoneme pronunciation feature distribution information Zt can then be determined.
[0096] If the distribution mapping module performs reverse mapping on the phoneme pronunciation feature distribution information, that is, maps the phoneme pronunciation feature distribution information Zt to the audio feature distribution information Za', then the mapped audio feature distribution information Za' can be used as the distribution mapping result, and the pre-acquired audio feature distribution information Za can be used as the reference distribution mapping result. Then, the difference between the phoneme feature distribution mapping result Za' and the audio feature information Za can be determined.
[0097] In this embodiment, the distribution mapping module performs reversible distribution mapping on the phoneme feature distribution information and the audio feature distribution information. This enables the distribution mapping module to accurately construct the relationship between simple phoneme feature distribution information and complex audio feature distribution information, ensuring that key information is not lost during data transformation and effectively improving the performance and stability of the singing conversion model.
[0098] In some exemplary embodiments, the pronunciation information of the sample song includes the pitch information and phoneme features of the sample song; in step S1031, obtaining the phoneme pronunciation feature distribution information of the sample song based on the text features and the pronunciation information of the sample song may include the following steps:
[0099] The text features, pitch information of the sample song, and phoneme features of the sample song are input into the text encoding module to obtain the phoneme prior distribution output by the text encoding module; the phoneme prior distribution is used as the phoneme pronunciation feature distribution information of the sample song.
[0100] In practical applications, pitch information and phoneme features of sample songs can be obtained. In some exemplary embodiments, frequency domain analysis methods such as Short-Time Fourier Transform (STFT) can be used to convert the audio signal of the sample song to the time-frequency domain, thereby extracting pitch information. Alternatively, linear predictive cepstral coefficients (LPCC) or Mel frequency cepstral coefficients (MFCC) can be used to extract pitch information. Furthermore, the prior probabilities of phonemes in the sample song can be obtained through a phoneme feature extraction model and used as phoneme features.
[0101] Then, the text features, pitch information, and phoneme features of the sample song can be input into the text encoding module. The text encoding module then obtains the prior phoneme distribution based on the text features, pitch information, and phoneme features, and uses this prior phoneme distribution as the phoneme pronunciation feature distribution information of the sample song. In this embodiment, by inputting the pitch information and phoneme features of the sample song together into the text encoding module, the acoustic characteristics of a singer producing the corresponding phoneme sound according to the pitch of the sample song can be effectively simulated, providing a foundation for subsequent timbre conversion.
[0102] In one embodiment, Figure 3 As shown, a method for song timbre conversion is provided. This embodiment illustrates the application of this method to a terminal. It is understood that this method can also be applied to a server, and further to a system including both a terminal and a server, and is implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps:
[0103] S301, Obtain the song to be converted, and timbre configuration information including at least timbre description text.
[0104] In practice, users can obtain the songs to be converted, i.e., the songs to be converted. In some examples, the songs sung by users may be out of tune or lack singing skills. In this case, songs sung by other singers with good singing skills or accurate pitch can be used as the songs to be converted.
[0105] On the other hand, timbre configuration information can also be set, which includes at least timbre description text. The timbre description text can be any natural language statement entered according to the user's timbre setting requirements. This natural language statement can describe the characteristics of the singer's timbre, such as "bright", "magnetic", "round", "delicate" and other descriptive text.
[0106] S302, input the pronunciation information and timbre configuration information of the song to be converted into the trained singing voice conversion model, and obtain the song with timbre conversion output by the singing voice conversion model.
[0107] The singing voice conversion model is trained according to the singing voice conversion model training method in any of the above embodiments.
[0108] In this step, the timbre configuration information and the pronunciation information corresponding to the song to be converted can be input together into the trained vocal conversion model to obtain the song with timbre conversion output by the vocal conversion model. In some embodiments, the pronunciation information corresponding to the song to be converted can be extracted by the vocal conversion model, or the pronunciation information can be extracted by other models.
[0109] For example, such as Figure 4a As shown, the pitch information, phonetic posterior grams (PPGs), and timbre configuration information of the song to be converted can be input into the singing conversion model. For the timbre description text in the timbre configuration information, the text encoding module can obtain the corresponding text features. Then, the text features corresponding to the timbre description text are used as the target timbre features and input together with the pronunciation information of the song to be converted into the text encoding module. This yields the phonetic pronunciation feature distribution information output by the text encoding module. After passing through the distribution mapping module, the phonetic pronunciation feature distribution information is converted into audio feature distribution information. Finally, after decoding by the distribution decoder, the timbre-converted song can be obtained.
[0110] In this embodiment, the song to be converted, along with timbre configuration information including at least a timbre description text, can be obtained. Then, the pronunciation information and timbre configuration information of the song to be converted are input into a trained vocal conversion model to obtain the timbre-converted song output by the vocal conversion model. This embodiment allows the trained vocal conversion model to determine the target timbre features using the text features corresponding to the timbre description text by simply inputting the timbre description text. When timbre conversion is needed, only any timbre description text related to the target timbre needs to be input, eliminating the need for the user to record or search for example recordings beforehand, effectively improving vocal conversion efficiency.
[0111] The song timbre conversion method of this application can transform a user's off-key or unskilled singing voice into one with vocal technique; in addition, this application can efficiently fine-tune the synthesized timbre using natural language input, thereby achieving the effect of voice beautification.
[0112] In one embodiment, the timbre configuration information further includes source audio with a specified timbre; in step S302, the pronunciation information and timbre configuration information of the song to be converted are input into the trained singing conversion model to obtain the timbre-converted song output by the singing conversion model, which may include the following steps:
[0113] S3021, the text encoding module in the singing conversion model extracts features from the timbre description text to obtain the text features of the timbre description text, and the audio encoding module in the singing conversion model extracts features from the source audio to obtain the timbre features of the source audio.
[0114] like Figure 4b As shown, in some embodiments, the phoneme configuration information can simultaneously include timbre description text and source audio with a specified timbre. After the phoneme configuration information is input into the singing conversion model, on the one hand, the text encoding module in the singing conversion model can extract features from the timbre description text to obtain the text features of the timbre description text; on the other hand, the audio encoding module in the singing conversion model can extract features from the source audio to obtain the timbre features of the source audio.
[0115] S3022, Based on the feature weights of the text features of the timbre description text and the timbre features of the source audio, feature fusion is performed on the text features of the timbre description text and the timbre features of the source audio to obtain the target timbre features.
[0116] After obtaining the text features of the timbre description text and the timbre features of the source audio, their respective feature weights can be obtained and fused. The fused feature is then used as the target timbre feature. For example, if the fusion is performed with a feature weight of 3:7 for the text features of the timbre description text and the timbre features of the source audio, the target timbre feature = 0.3 * text features of the timbre description text + 0.7 * timbre features of the source audio.
[0117] S3023 outputs the timbre-converted song based on the target timbre characteristics and the pronunciation information of the song to be converted.
[0118] After obtaining the target timbre features, the singing conversion model can output a timbre-converted song based on the target timbre features and the pronunciation information of the song to be converted. The specific output process can be referred to the aforementioned embodiments, and will not be repeated here.
[0119] In this embodiment, by performing feature fusion on the text features of the timbre description text and the timbre features of the source audio according to their respective feature weights, the target timbre features are obtained. This allows for the simultaneous adjustment of the singer's timbre through both text and audio, greatly improving the flexibility and efficiency of vocal timbre conversion.
[0120] In other embodiments, the timbre configuration information may also include only the source audio, for example, such as Figure 4cAs shown, the target timbre features can be determined solely based on the timbre features output by the audio encoding module. Those skilled in the art can select the type of timbre configuration information input into the singing conversion model according to the actual situation.
[0121] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0122] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 5 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores dry vocal data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network. When executed by the processor, the computer program implements a vocal conversion model training method or a song timbre conversion method.
[0123] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 6As shown, the computer device includes a processor, memory, input / output interfaces, a communication interface, a display unit, and an input device. The processor, memory, and input / output interfaces are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interfaces are used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a singing voice conversion model training method or a song timbre conversion method. The display unit is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device casing, or an external keyboard, touchpad or mouse.
[0124] Those skilled in the art will understand that Figure 5 and Figure 6 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0125] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0126] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0127] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0128] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0129] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0130] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0131] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A method for training a singing voice conversion model, characterized in that, The method includes: Obtain sample songs and corresponding vocal description text of the singers whose voices are paired with the sample songs; The timbre features of the sample song are obtained by the audio encoding module in the singing conversion model, and the text features corresponding to the singer's timbre description text are obtained by the first text encoding module in the singing conversion model. The text features, the pitch information of the sample song, and the phoneme features of the sample song are input into the second text encoding module. Based on the phoneme prior distribution output by the second text encoding module, the phoneme pronunciation feature distribution information of the sample song is obtained. The audio feature distribution information used to decode the sample song is obtained, and the distribution mapping module in the singing conversion model performs a distribution mapping between the phoneme pronunciation feature distribution information and the audio feature distribution information to obtain the distribution mapping result; Based on the differences between the timbre features and the text features, and the differences between the distribution mapping result and the reference distribution mapping result, the model parameters of the singing voice conversion model are adjusted to obtain a trained singing voice conversion model. The trained singing voice conversion model is used to extract the text features corresponding to the input timbre description text through the first text encoding module, and obtain the target timbre features based on the text features corresponding to the timbre description text to perform song timbre conversion.
2. The method according to claim 1, characterized in that, The step of adjusting the model parameters of the singing conversion model based on the differences between the timbre features and the text features, and the differences between the distribution mapping result and the reference distribution mapping result, includes: The distribution decoding module in the singing conversion model decodes the corresponding predicted song based on the audio feature distribution information. A first loss value is determined based on the difference between the timbre features and the text features; a second loss value is determined based on the difference between the distribution mapping result and the reference distribution mapping result; and a third loss value is determined based on the difference between the predicted song and the sample song. The model loss value is determined based on the first loss value, the second loss value, and the third loss value, and the model parameters of the singing conversion model are adjusted based on the model loss value.
3. The method according to claim 2, characterized in that, The distribution decoding module in the singing conversion model decodes the corresponding predicted song based on the audio feature distribution information, including: The audio feature distribution information, the text features, and the pitch information of the sample song are input into the distribution decoding module in the singing conversion model, and the predicted song is obtained based on the output of the distribution decoding module.
4. The method according to claim 1, characterized in that, The reference distribution mapping result includes the phoneme pronunciation feature distribution information or the audio feature information; The difference between the distribution mapping result and the reference distribution mapping result is determined through the following steps: If the distribution mapping result is the audio feature distribution mapping result obtained by the distribution mapping module through forward mapping of the audio feature distribution information, then the difference between the audio feature distribution mapping result and the phoneme pronunciation feature distribution information is determined. If the distribution mapping result is the phoneme feature distribution mapping result obtained by the distribution mapping module through reverse mapping of the phoneme pronunciation feature distribution information, then the difference between the phoneme feature distribution mapping result and the audio feature distribution information is determined.
5. A method for converting the timbre of a song, characterized in that, The method includes: Obtain the song to be converted, as well as timbre configuration information, including at least a timbre description text; The pronunciation information and timbre configuration information of the song to be converted are input into the trained singing voice conversion model to obtain the song with timbre conversion output by the singing voice conversion model. The singing voice conversion model is trained according to the singing voice conversion model training method as described in any one of claims 1 to 4.
6. The method according to claim 5, characterized in that, The timbre configuration information also includes source audio with a specified timbre; The step of inputting the pronunciation information and timbre configuration information of the song to be converted into a trained singing voice conversion model to obtain the timbre-converted song output by the singing voice conversion model includes: The first text encoding module in the singing conversion model extracts features from the timbre description text to obtain the text features of the timbre description text, and the audio encoding module in the singing conversion model extracts features from the source audio to obtain the timbre features of the source audio; Based on the feature weights of the text features of the timbre description text and the timbre features of the source audio, feature fusion is performed on the text features of the timbre description text and the timbre features of the source audio to obtain the target timbre features; Based on the target timbre characteristics and the pronunciation information of the song to be converted, the song with converted timbre is output.
7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the singing voice conversion model training method according to any one of claims 1 to 4 or the steps of the song timbre conversion method according to any one of claims 5 to 6.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the singing voice conversion model training method according to any one of claims 1 to 4 or the steps of the song timbre conversion method according to any one of claims 5 to 6.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the singing voice conversion model training method according to any one of claims 1 to 4 or the steps of the song timbre conversion method according to any one of claims 5 to 6.
Citation Information
Patent Citations
System and method for voice-to-voice conversion
CN111201565A
Construction method of voice synthesizer, and voice synthesis method and device
CN113823257A