Phoneme posterior graph model training method and device, medium and program product
By training the phoneme posterior graph model to extract the lead singer's vocal content from the sound audio, the problem of poor singing conversion effect in the prior art is solved, and high-quality singing conversion in complex harmony environments is achieved.
Patent Information
- Application Number
- CN202510870662.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-08-26
AI Technical Summary
The prior art is difficult to accurately distinguish and extract the unique pronunciation information of each harmony, resulting in poor singing voice conversion effect, inaccurate restoration of the singing content, and may cause timbre distortion and blurred melody.
By obtaining the standard phoneme posterior graph features and sound audio of the lead singer audio, the phoneme posterior graph model is used for feature extraction and training, and the model is adjusted in combination with the loss function until the convergence condition is reached, and the training phoneme posterior graph model can be generated, so that the lead singer's singing content can be extracted from the sound audio.
It improves the accuracy of the lead singer's vocal content extraction, enhances the quality of vocal conversion, and can retain the main vocal information in a complex harmony environment to avoid chaotic pronunciation.
Smart Images

Figure CN120544585A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of audio processing technology, and in particular to a phoneme posterior graph model training method, equipment, medium and program product. Background Art
[0002] Singing Voice Conversion (SVC) is a technology that converts one person's singing voice into another's, while preserving the content and melody of the original voice. Singing Voice Conversion is primarily based on the content and melody of the original voice, as well as the timbre of the target voice.
[0003] Current audio processing techniques for capturing vocal content often rely on transcription and self-supervised learning. However, these methods struggle to accurately distinguish and independently extract the unique pronunciation information of individual harmonies, instead tending to capture the average pronunciation characteristics of multiple voices. Consequently, they are unable to accurately extract the desired vocal content, reducing the effectiveness of vocal conversion. Summary of the Invention
[0004] In view of this, the present invention aims to provide a method, device, medium, and program product for training a phoneme posterior graph model, which can improve the accuracy of lead vocal content extraction and the quality of vocal conversion. The specific scheme is as follows:
[0005] In a first aspect, the present application discloses a method for training a phoneme posterior graph model, comprising:
[0006] Obtaining standard phoneme posterior graph features corresponding to the lead singer audio, and harmony audio containing other human voices except the lead singer synthesized based on the lead singer audio;
[0007] Extracting features of the harmony audio using the to-be-trained phoneme posterior graph model to obtain predicted phoneme posterior graph features;
[0008] A first loss value is determined based on the standard phoneme posterior graph features and the predicted phoneme posterior graph features, and the phoneme posterior graph model to be trained is adjusted according to the first loss value until the convergence condition is reached, so as to obtain a trained phoneme posterior graph model, so as to use the trained phoneme posterior graph model to determine the lead vocal content in the song.
[0009] Optionally, adjusting the phoneme posterior graph model to be trained according to the first loss value includes:
[0010] Based on the predicted phoneme posterior graph features, the decoder is used to convert the harmony audio into a singing voice, and a first spectrum corresponding to the harmony audio is obtained;
[0011] Obtaining a second spectrum corresponding to the lead vocal audio, and determining a second loss value based on a difference between the first spectrum and the second spectrum;
[0012] The phoneme posterior graph model to be trained is adjusted according to the first loss value and the second loss value.
[0013] Optionally, using a decoder to convert the harmony audio into singing voice to obtain a first spectrum corresponding to the harmony audio includes:
[0014] Determining the lead singing voice content corresponding to the harmony audio according to the predicted phoneme posterior graph features;
[0015] Generate a fusion feature based on the singing melody feature, the singing timbre feature and the lead singing voice content corresponding to the harmony audio;
[0016] The fusion feature is decoded by using a decoder to obtain a first spectrum corresponding to the harmony audio.
[0017] Optionally, after determining the second loss value based on the difference between the first spectrum and the second spectrum, the method further includes:
[0018] Adjusting the decoder according to the second loss value until a convergence condition is reached to obtain a trained decoder;
[0019] A singing voice conversion model is obtained based on the trained phoneme posterior graph model and the trained decoder.
[0020] Optionally, the harmony audio is audio of a target scene generated based on the lead vocal audio;
[0021] The target scene includes one or more of a short-time high-volume chorus scene, a delayed harmony scene, a duet scene, and a chorus scene.
[0022] Optionally, the harmony audio of the short-term high-volume chorus scene is obtained by fusing vocal data into a target period of the lead vocal audio of the song;
[0023] The harmony audio of the delayed harmony scene is obtained by shifting and superimposing the lead vocal audio in the same song;
[0024] The harmony audio of the duet scene is obtained by selectively superimposing the lead vocal audios of different songs;
[0025] The harmony audio of the chorus scene is obtained by superimposing the lead vocal audios of at least two songs.
[0026] In a second aspect, the present application discloses a singing voice conversion method, comprising:
[0027] Inputting the singing voice file to be converted into a trained phoneme posterior graph model; the trained phoneme posterior graph model is trained according to the aforementioned phoneme posterior graph model training method;
[0028] According to the output of the trained phoneme posterior graph model, the main singing voice content in the song file to be converted is determined so that singing conversion is performed based on the main singing voice content.
[0029] Optionally, performing singing voice conversion based on the lead singing voice content includes:
[0030] Generate fusion features based on the singing melody information of the singing file to be converted, the lead singing voice content and the target timbre file;
[0031] The fusion feature is decoded into a spectrum using a trained decoder, and the spectrum is converted into audio to achieve singing conversion of the singing file to be converted.
[0032] In a third aspect, the present application discloses an electronic device, comprising:
[0033] Memory, used to store computer programs;
[0034] A processor is used to execute the computer program to implement the aforementioned singing voice conversion method or phoneme posterior graph model training method.
[0035] In a fourth aspect, the present application discloses a computer-readable storage medium for storing a computer program; wherein the computer program, when executed by a processor, implements the aforementioned singing voice conversion method, or the phoneme posterior graph model training method.
[0036] In a fifth aspect, the present application discloses a computer program product, including a computer program, which, when executed by a processor, implements the aforementioned singing voice conversion method or phoneme posterior graph model training method.
[0037] In the present application, the standard phoneme posterior graph features corresponding to the lead singer audio and the harmony audio containing the voices of other people except the lead singer synthesized based on the lead singer audio are obtained; the harmony audio is feature extracted using the phoneme posterior graph model to be trained to obtain the predicted phoneme posterior graph features; a first loss value is determined based on the standard phoneme posterior graph features and the predicted phoneme posterior graph features, and the phoneme posterior graph model to be trained is adjusted according to the first loss value until the convergence condition is reached to obtain the trained phoneme posterior graph model, so as to use the trained phoneme posterior graph model to determine the lead singer vocal content in the song. The phoneme posterior graph model is trained by utilizing the differences in phoneme posterior graph features between the lead singer audio and the harmony audio. The harmony audio is an audio generated based on the lead singer audio and contains other human voices except the lead singer. The trained phoneme posterior graph model obtained through training has the ability to extract the lead singer audio from the harmony audio. Using the trained phoneme posterior graph model to extract the lead singer content in the singing file to be converted can improve the accuracy of the lead singer content extraction and thus improve the quality of singing conversion. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0039] Figure 1 A flowchart of a phoneme posterior graph model training method provided in this application;
[0040] Figure 2 A schematic diagram of a specific singing voice conversion system provided in this application;
[0041] Figure 3 A specific singing voice conversion method flow chart provided in this application;
[0042] Figure 4 A schematic diagram of the structure of a phoneme posterior graph model training device provided in this application;
[0043] Figure 5 This is a structural diagram of an electronic device provided in this application. DETAILED DESCRIPTION
[0044] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0045] In the related art, the technical solutions adopted for obtaining the content of a piece of singing voice are mostly text transcription or self-supervised learning methods. Text transcription usually uses speech recognition technology to convert the singing voice in the audio into text form, and then analyzes and processes the singing content based on the text information, which has poor recognition effect. The self-supervised method uses the characteristics of the audio itself to build a model to automatically learn and extract valuable information from a large amount of audio data to achieve the acquisition of singing content; however, this type of method has obvious limitations. When multiple parts in a harmony sing different content, there will be pronunciation confusion. At this time, the model will retain the average pronunciation information of multiple voices; that is, it is difficult for the model to accurately distinguish and independently extract the unique pronunciation information of each harmony, but tends to capture the average pronunciation characteristics of multiple voices. Ultimately, the converted audio effect is poor. Not only can it not accurately restore the original rich and unique detailed information of each harmony, it may also cause many problems such as timbre distortion, melody ambiguity, and reduced clarity of lyrics, which greatly affects the quality and usability of the audio.
[0046] The present application discloses a method for training a phoneme posterior graph model. Figure 1 As shown, the method may include the following steps:
[0047] Step S11: obtaining the standard phoneme posterior graph features corresponding to the lead singer audio, and the harmony audio containing other human voices except the lead singer synthesized based on the lead singer audio.
[0048] First, obtain the standard phoneme posterior graph features corresponding to the lead vocal audio. Specifically, the phoneme posterior graph (PPG) features can be obtained by inputting the lead vocal audio into an encoder. Harmony audio is also obtained. This harmony audio is synthesized based on the lead vocal audio and includes other vocals besides the lead vocal.
[0049] In some implementations, the harmony audio is generated based on the lead vocal audio for target scenarios. Specifically, to improve model capabilities, ensure the richness of training samples, and meet the data diversity and scalability requirements of subsequent model training, the existing chorus audio is divided into multiple scenarios, and harmony audio for different scenarios is generated from the lead vocal audio. These target scenarios include, but are not limited to, short, high-volume chorus scenarios, delayed harmony scenarios, duet scenarios, and chorus scenarios.
[0050] Among them, the harmony audio of the short-time high-volume chorus scene is obtained by fusing vocal data into the target time period of the lead singer audio of the song; the harmony audio of the delayed harmony scene is obtained by shifting and superimposing the lead singer audio in the same song; the harmony audio of the duet scene is obtained by selectively superimposing the lead singer audio in different songs; the harmony audio of the chorus scene is obtained by superimposing the lead singer audio in at least two songs.
[0051] Specifically, by fusing vocal data with a target time period of a song's lead vocal audio, harmony audio for short, high-volume chorus scenes is generated. Short, high-energy scenes refer to songs with high-volume but short-duration harmonies. For this, a selection of vocal data is directly fused with the lead vocal audio. Delayed harmony scenes are generated by shifting and overlaying lead vocal audio from the same song. In delayed harmony scenes, some songs have delayed harmonies with the same content. For this, vocals from the same song are shifted, compressed, equalized, and reverberated before being overlaid on the original vocals. Harmony audio for duets is generated by selectively overlaying lead vocals from different songs. Duets refer to songs with multiple vocalists singing duets with completely different melodies, timbres, and content. For this, vocals from two different songs are selected and overlaid using a corresponding algorithm to ensure some overlap and some solos. By superimposing the lead vocal audio of at least two songs, the harmony audio of the chorus scene is obtained, that is, a completely mixed scene, in which there are multiple tracks of chorus audio with different pitches, timbres and contents superimposed on each other. For this situation, the vocals of multiple songs are processed by an effector and then superimposed.
[0052] By generating a variety of harmony datasets for various scenarios and categories for model training, we target a wide range of harmony scenarios, including but not limited to short, high-energy harmonies, delayed harmonies, duet harmonies, and fully aliased harmonies. This comprehensive and targeted dataset provides a sufficient, high-quality data foundation for subsequent model training, significantly enriching the model's training samples and enabling it to learn the characteristic patterns of various complex harmony scenarios, improving the model's ability to extract lead vocal content.
[0053] Step S12: extracting features of the harmony audio using the phoneme posterior graph model to be trained to obtain predicted phoneme posterior graph features.
[0054] The harmonic audio is fed into the trained phoneme posterior graph model, and the output of the trained phoneme posterior graph model is used as the predicted phoneme posterior graph feature. The PPG feature is a probability distribution of phonemes associated with the speech content, and further processing can reveal the speech content.
[0055] Step S13: Determine a first loss value based on the standard phoneme posterior graph features and the predicted phoneme posterior graph features, adjust the phoneme posterior graph model to be trained according to the first loss value until the convergence condition is reached, and obtain the trained phoneme posterior graph model, so as to use the trained phoneme posterior graph model to determine the lead singer's vocal content in the song.
[0056] A first loss function is used to determine a first loss value based on the difference between the standard phoneme posterior graph features and the predicted phoneme posterior graph features, and the to-be-trained phoneme posterior graph model is trained based on the first loss value to obtain the trained phoneme posterior graph model. The phoneme posterior graph model is trained using the standard phoneme posterior graph features as a standard so that the trained phoneme posterior graph model can extract the phoneme posterior graph features of the lead vocal from the harmony audio, and further processing can then obtain the lead vocal content in the harmony audio.
[0057] For example Figure 2 As shown, the dotted line represents the portion used only during training. The encoder used to extract the standard phoneme posterior graph features can be the whisper model's encoder, meaning it does not participate in speech conversion. The trained phoneme posterior graph model is used to predict the lead vocal PPG features to be retained from the harmony audio superimposed on the lead vocal. During the operational phase, the trained phoneme posterior graph model encoder extracts clear lead vocal content while suppressing harmony information for subsequent singing conversion.
[0058] After obtaining the lead singer audio, generate harmony audio based on the lead singer audio, determine the PPG features of the lead singer audio, use the phoneme posterior graph model to be trained to extract the PPG features of the harmony audio, and use the first loss function to train the phoneme posterior graph model to be trained based on these two PPG features. The trained phoneme posterior graph model obtained after training has the ability to extract the lead singer audio from the harmony audio; the above-mentioned first loss function can be a loss function such as MSELoss (mean square error loss function) for measuring the difference between the predicted value and the true value of the model, which is not limited in this embodiment.
[0059] After training, the phoneme posterior graph model can retain the content information of the main vocal part and suppress the harmony part in the presence of multi-track harmonies with different contents. Using this feature information for singing voice conversion can retain the content information of the lead singer in the presence of complex harmonies and avoid pronunciation confusion.
[0060] In a specific implementation, adjusting the phoneme posterior graph model to be trained according to the first loss value may include: using a decoder to convert the harmony audio into singing voice based on the predicted phoneme posterior graph features to obtain a first spectrum corresponding to the harmony audio; obtaining a second spectrum corresponding to the lead singer audio, and determining a second loss value based on the difference between the first spectrum and the second spectrum; adjusting the phoneme posterior graph model to be trained according to the first loss value and the second loss value. After determining the second loss value based on the difference between the first spectrum and the second spectrum, it also includes: adjusting the decoder according to the second loss value until the convergence condition is reached to obtain a trained decoder; and obtaining a singing voice conversion model based on the trained phoneme posterior graph model and the trained decoder.
[0061] To further improve the effectiveness and quality of singing voice conversion, the difference between the first spectrum corresponding to the harmony audio and the second spectrum corresponding to the lead vocal audio is used to train the decoder and phoneme posterior graph model used for singing voice conversion. This improves the accuracy of the phoneme posterior graph model in extracting the lead vocal content and enhances the decoder's singing voice conversion capabilities. The aforementioned second loss value can be determined using the L1 loss function.
[0062] Among them, the harmony audio is converted into singing voice using a decoder to obtain a first spectrum corresponding to the harmony audio, including: determining the lead singing voice content corresponding to the harmony audio based on the predicted phoneme posterior graph features; generating a fusion feature based on singing melody features, singing timbre features and the lead singing voice content corresponding to the harmony audio; and decoding the fusion feature using a decoder to obtain the first spectrum corresponding to the harmony audio.
[0063] Specifically, you can use a singing melody extractor (such as CQT (Constant-Q Transform) feature) to extract singing melody information; extract the timbre features in the target timbre file; and use the trained PPG model to extract the lead singing content in the singing file to be converted. For example Figure 2 As shown, the fusion is performed based on the singing melody information, the lead vocal content, and the timbre characteristics. Specifically, the melody features, the lead vocal content, and the singing timbre are first unified in length using a length regulator, then linearly mapped to unify the dimensions before being fused to generate fused features. The fused features are then transmitted to the decoder to be restored to audio. The fused features are then transmitted to the trained decoder, which first decodes the fused features into a spectrum. The vocoder then converts the spectrum into audio, resulting in the final converted audio corresponding to the singing file to be converted.
[0064] The generation process of the above-mentioned second spectrum includes: obtaining the lead singer's vocal content of the lead singer audio based on the standard phoneme posterior graph features of the lead singer audio; using the same singing melody features and singing timbre features as when generating the first spectrum to generate the fusion features of the lead singer audio; using the trained decoder to decode the fusion features of the lead singer audio to obtain the second spectrum corresponding to the lead singer audio.
[0065] Specifically, the decoder and PPG model are trained using the loss between the second spectrum and the generated first spectrum, enabling the model to extract PPG information that meets the overall system requirements. This information is then combined with other features and restored to audio via the decoder. The PPG model is also trained using differences in phoneme posterior graph features, enabling it to retain PPG information of the primary vocal component within harmony data, maintaining clarity of the lead vocal in challenging scenarios with multiple harmonies and varying content. Based on this multi-scenario harmony dataset and the comprehensive training of the phoneme posterior graph model and decoder, the singing voice conversion system is able to accurately extract clear PPG features and accurately reflect the pronunciation of the primary vocal in complex multi-track audio containing diverse content and layered harmonies, providing a foundation for subsequent processing. Specifically, the system effectively identifies and preserves the key features of the primary vocal component within layered harmonic audio, while also appropriately suppressing the harmonies. This ensures accurate extraction of the lead vocal content in complex audio environments, resolving the common articulation issues with complex harmonies in singing voice conversion models.
[0066] In this embodiment, the trained phoneme posterior graph model is used to determine the vocal content of the lead singer in a song, including: inputting a singing file to be converted into the trained phoneme posterior graph model; the trained phoneme posterior graph model outputs a phoneme posterior graph feature corresponding to the singing file to be converted, which is the phoneme posterior graph feature of the lead singer in the singing file to be converted. The phoneme posterior graph feature is a phoneme probability distribution related to the speech content, and further processing can be used to obtain the speech content, i.e., the lead singer vocal content. Then, singing conversion is performed based on the lead singer vocal content.
[0067] Vocal conversion is performed based on the lead vocal content, including: vocal conversion based on the vocal melody information of the vocal file to be converted, the lead vocal content, and the target timbre file. That is, the user inputs the vocal file to be converted and the target timbre file recorded by the user, and the system automatically generates a converted song that meets the user's needs. For example, if a user wants to convert a professional song with harmony by a singer to their own timbre, they only need to provide the original song audio and the reference timbre clip recorded by themselves. The system will automatically complete the conversion and generate a professional song audio sung by the user's timbre; users can experience different singing styles.
[0068] Furthermore, singing conversion based on the lead vocal content can also include: singing conversion based on the singing melody information of the singing file to be converted, the lead vocal content, the target timbre file and the target pitch value, so that the song can be adjusted to a pitch that is more suitable for one's own vocal range to meet personalized singing conversion needs.
[0069] As can be seen from the above, in this embodiment, the standard phoneme posterior graph features corresponding to the lead singer audio and the harmony audio containing the rest of the human voice except the lead singer synthesized based on the lead singer audio are obtained; the harmony audio is feature extracted using the phoneme posterior graph model to be trained to obtain the predicted phoneme posterior graph features; a first loss value is determined based on the standard phoneme posterior graph features and the predicted phoneme posterior graph features, and the phoneme posterior graph model to be trained is adjusted according to the first loss value until the convergence condition is reached, so as to obtain the trained phoneme posterior graph model, so as to use the trained phoneme posterior graph model to determine the lead singer vocal content in the song. The phoneme posterior graph model is trained by utilizing the differences in phoneme posterior graph features between the lead singer audio and the harmony audio. The harmony audio is an audio generated based on the lead singer audio and contains other human voices except the lead singer. The trained phoneme posterior graph model obtained through training has the ability to extract the lead singer audio from the harmony audio. Using the trained phoneme posterior graph model to extract the lead singer content in the singing file to be converted can improve the accuracy of the lead singer content extraction and thus improve the quality of singing conversion.
[0070] Based on the above embodiment, the present application also discloses a singing voice conversion method, see Figure 3As shown, the method may include the following steps:
[0071] Step S21: inputting the singing voice file to be converted into the trained phoneme posterior graph model.
[0072] The trained phoneme posterior graph model is obtained by training using the above-mentioned phoneme posterior graph model training method.
[0073] Step S22: Determine the main singing voice content in the song file to be converted based on the output of the trained phoneme posterior graph model, so as to perform singing conversion based on the main singing voice content.
[0074] The singing file to be converted is input into the trained phoneme posterior graph model. The PPG features corresponding to the singing file to be converted output by the trained phoneme posterior graph model are the PPG features of the lead singer in the singing file to be converted. The PPG features are the probability distribution of phonemes related to the speech content. Through further processing, the speech content, that is, the lead singer's singing content, can be obtained. Then, the singing conversion is performed based on the lead singer's singing content. The trained phoneme posterior graph model can retain the content information of the main vocal part and suppress the harmony part in the presence of multiple tracks of different contents. Using this feature information for singing voice conversion can retain the content information of the lead singer in the presence of complex harmony and avoid pronunciation confusion.
[0075] The singing voice conversion is performed based on the lead singing voice content, and the singing voice conversion can be performed based on the singing melody information of the singing voice file to be converted, the lead singing voice content and the target timbre file. Specifically, the singing voice conversion based on the lead singing voice content includes: generating fusion features based on the singing melody information of the singing voice file to be converted, the lead singing voice content and the target timbre file; using the trained decoder to decode the fusion features into a spectrum, and converting the spectrum into audio to achieve the singing voice conversion of the singing voice file to be converted. The training process of the above-mentioned trained decoder includes: using the decoder to perform singing voice conversion on the harmony audio based on the predicted phoneme posterior graph features to obtain the first spectrum corresponding to the harmony audio; obtaining the second spectrum corresponding to the lead singing audio, and determining the second loss value based on the difference between the first spectrum and the second spectrum; adjusting the phoneme posterior graph model to be trained according to the above-mentioned first loss value and second loss value, and adjusting the decoder according to the second loss value until the convergence condition is reached to obtain a trained decoder.
[0076] The system automatically generates a converted song tailored to the user's needs by providing the user with a vocal file to be converted and a target sound file they have recorded themselves. For example, if a user wishes to convert a professional, harmonized song by a vocalist to their own sound, they simply provide the original song audio and a clip of their own recorded sound. The system then automatically converts the sound and generates a professional audio clip of the song sung with the user's sound, allowing users to experience different singing styles.
[0077] Furthermore, singing conversion based on the lead vocal content can also include: singing conversion based on the singing melody information of the singing file to be converted, the lead vocal content, the target timbre file and the target pitch value, so that the song can be adjusted to a pitch that is more suitable for one's own vocal range to meet personalized singing conversion needs.
[0078] For the specific process of the above steps, please refer to the corresponding content disclosed in the above embodiments, and will not be repeated here.
[0079] As can be seen from the above, in this embodiment, the singing file to be converted is input into the trained phoneme posterior graph model; the trained phoneme posterior graph model is trained according to the aforementioned phoneme posterior graph model training method; based on the output of the trained phoneme posterior graph model, the lead singing content in the singing file to be converted is determined so as to perform singing conversion based on the lead singing content. The phoneme posterior graph model is trained by utilizing the difference in phoneme posterior graph features between the lead singer audio and the harmony audio, and the harmony audio is an audio generated based on the lead singer audio and includes other human voices except the lead singer. The trained phoneme posterior graph model thus obtained has the ability to extract the lead singer audio from the harmony audio. Using the trained phoneme posterior graph model to extract the lead singing content in the singing file to be converted can improve the accuracy of extracting the lead singing content, thereby improving the quality of singing conversion.
[0080] Correspondingly, the present application also discloses a phoneme posterior graph model training device, see Figure 4 As shown, the device includes:
[0081] A feature acquisition module 11 is used to obtain standard phoneme posterior graph features corresponding to the lead singer audio, and a harmony audio containing other human voices except the lead singer synthesized based on the lead singer audio;
[0082] A feature extraction module 12 is configured to extract features from the harmony audio using the phoneme posterior graph model to be trained to obtain predicted phoneme posterior graph features;
[0083] The training module 13 is used to determine a first loss value based on the standard phoneme posterior graph features and the predicted phoneme posterior graph features, and adjust the phoneme posterior graph model to be trained according to the first loss value until the convergence condition is reached, so as to obtain the trained phoneme posterior graph model, so as to use the trained phoneme posterior graph model to determine the lead singer's vocal content in the song.
[0084] As can be seen from the above, in this embodiment, the standard phoneme posterior graph features corresponding to the lead singer audio and the harmony audio containing the rest of the human voice except the lead singer synthesized based on the lead singer audio are obtained; the harmony audio is feature extracted using the phoneme posterior graph model to be trained to obtain the predicted phoneme posterior graph features; a first loss value is determined based on the standard phoneme posterior graph features and the predicted phoneme posterior graph features, and the phoneme posterior graph model to be trained is adjusted according to the first loss value until the convergence condition is reached, so as to obtain the trained phoneme posterior graph model, so as to use the trained phoneme posterior graph model to determine the lead singer vocal content in the song. The phoneme posterior graph model is trained by utilizing the differences in phoneme posterior graph features between the lead singer audio and the harmony audio. The harmony audio is an audio generated based on the lead singer audio and contains other human voices except the lead singer. The trained phoneme posterior graph model obtained through training has the ability to extract the lead singer audio from the harmony audio. Using the trained phoneme posterior graph model to extract the lead singer content in the singing file to be converted can improve the accuracy of the lead singer content extraction and thus improve the quality of singing conversion.
[0085] In some specific embodiments, the training module 13 includes:
[0086] a spectrum determining unit, configured to perform singing voice conversion on the harmony audio using a decoder based on the predicted phoneme posterior graph feature to obtain a first spectrum corresponding to the harmony audio;
[0087] a second loss value determining unit, configured to obtain a second spectrum corresponding to the lead vocal audio, and determine a second loss value based on a difference between the first spectrum and the second spectrum;
[0088] A model training unit is used to adjust the phoneme posterior graph model to be trained according to the first loss value and the second loss value.
[0089] In some specific embodiments, the spectrum determining unit includes:
[0090] a singing content determining unit, configured to determine the lead singing singing content corresponding to the harmony audio according to the predicted phoneme posterior graph features;
[0091] A fusion feature generating unit, configured to generate a fusion feature based on a singing melody feature, a singing timbre feature, and a lead singing voice content corresponding to the harmony audio;
[0092] The spectrum determination unit is configured to decode the fusion feature using a decoder to obtain a first spectrum corresponding to the harmony audio.
[0093] In some specific embodiments, the training module 13 includes:
[0094] a model adjustment unit, configured to, after determining a second loss value based on a difference between the first spectrum and the second spectrum, adjust the decoder according to the second loss value until a convergence condition is met, thereby obtaining a trained decoder;
[0095] The singing voice conversion model determining unit is used to obtain a singing voice conversion model based on the trained phoneme posterior graph model and the trained decoder.
[0096] In some specific embodiments, the harmony audio is audio of a target scene generated based on the lead vocal audio;
[0097] The target scene includes one or more of a short-time high-volume chorus scene, a delayed harmony scene, a duet scene, and a chorus scene.
[0098] In some specific embodiments, the harmony audio of the short-term high-volume chorus scene is obtained by fusing vocal data into a target period of the main vocal audio of the song;
[0099] The harmony audio of the delayed harmony scene is obtained by shifting and superimposing the lead vocal audio in the same song;
[0100] The harmony audio of the duet scene is obtained by selectively superimposing the lead vocal audios of different songs;
[0101] The harmony audio of the chorus scene is obtained by superimposing the lead vocal audios of at least two songs.
[0102] Furthermore, the present application also discloses an electronic device, see Figure 5 The contents in the drawings should not be considered as any limitation on the scope of use of the present application.
[0103] Figure 5 This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of the present application. The electronic device 20 may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the singing voice conversion method disclosed in any of the aforementioned embodiments, or the relevant steps in the phoneme posterior graph model training method.
[0104] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device. The communication protocol it follows is any communication protocol that can be applied to the technical solution of this application and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world. Its specific interface type can be selected according to specific application needs and is not specifically limited here.
[0105] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or CD, etc. The resources stored thereon include an operating system 221, a computer program 222 and data 223 including the lead singer's vocal content, etc. The storage method can be temporary storage or permanent storage.
[0106] The operating system 221 is used to manage and control the hardware devices and computer program 222 on the electronic device 20, so as to enable the processor 21 to calculate and process the massive data 223 in the memory 22. The operating system 221 can be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of performing the singing voice conversion method performed by the electronic device 20 as disclosed in any of the aforementioned embodiments, the computer program 222 can further include computer programs capable of performing other specific tasks.
[0107] Furthermore, an embodiment of the present application also discloses a computer storage medium, in which computer executable instructions are stored. When the computer executable instructions are loaded and executed by a processor, the singing voice conversion method or the phoneme posterior graph model training method steps disclosed in any of the aforementioned embodiments are implemented.
[0108] Furthermore, an embodiment of the present application also discloses a computer program product, including a computer program, which, when executed by a processor, implements the singing voice conversion method or the phoneme posterior graph model training method steps disclosed in any of the aforementioned embodiments.
[0109] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. Reference can be made to the descriptions of the identical or similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and the relevant parts can be referred to the descriptions of the methods.
[0110] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0111] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0112] The above is a detailed introduction to the phoneme posterior graph model training method, equipment, medium and program product provided by the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
Claims
1. A method for training a phoneme posterior graph model, characterized in that: include: Obtaining standard phoneme posterior graph features corresponding to the lead singer audio, and harmony audio containing other human voices except the lead singer synthesized based on the lead singer audio; Extracting features of the harmony audio using the to-be-trained phoneme posterior graph model to obtain predicted phoneme posterior graph features; A first loss value is determined based on the standard phoneme posterior graph features and the predicted phoneme posterior graph features, and the phoneme posterior graph model to be trained is adjusted according to the first loss value until the convergence condition is reached, so as to obtain a trained phoneme posterior graph model, so as to use the trained phoneme posterior graph model to determine the lead vocal content in the song.
2. The phoneme posterior graph model training method according to claim 1, characterized in that: Adjusting the to-be-trained phoneme posterior graph model according to the first loss value includes: Based on the predicted phoneme posterior graph features, the decoder is used to convert the harmony audio into a singing voice, and a first spectrum corresponding to the harmony audio is obtained; Obtaining a second spectrum corresponding to the lead vocal audio, and determining a second loss value based on a difference between the first spectrum and the second spectrum; The phoneme posterior graph model to be trained is adjusted according to the first loss value and the second loss value.
3. The phoneme posterior graph model training method according to claim 2, wherein: Converting the harmony audio into a singing voice using a decoder to obtain a first spectrum corresponding to the harmony audio includes: Determining the lead singing voice content corresponding to the harmony audio according to the predicted phoneme posterior graph features; Generate a fusion feature based on the singing melody feature, the singing timbre feature and the lead singing voice content corresponding to the harmony audio; The fusion feature is decoded by using a decoder to obtain a first spectrum corresponding to the harmony audio.
4. The phoneme posterior graph model training method according to claim 2, wherein: After determining the second loss value based on the difference between the first spectrum and the second spectrum, the method further includes: Adjusting the decoder according to the second loss value until a convergence condition is reached to obtain a trained decoder; A singing voice conversion model is obtained based on the trained phoneme posterior graph model and the trained decoder.
5. The method for training a phoneme posterior graph model according to any one of claims 1 to 4, wherein: The harmony audio is the audio of the target scene generated based on the lead vocal audio; The target scene includes one or more of a short-time high-volume chorus scene, a delayed harmony scene, a duet scene, and a chorus scene.
6. The method for training a phoneme posterior graph model according to claim 5, wherein: The harmony audio of the short-term high-volume chorus scene is obtained by fusing the vocal data to the target period of the lead singer audio of the song; The harmony audio of the delayed harmony scene is obtained by shifting and superimposing the lead vocal audio in the same song; The harmony audio of the duet scene is obtained by selectively superimposing the lead vocal audios of different songs; The harmony audio of the chorus scene is obtained by superimposing the lead vocal audios of at least two songs.
7. A singing voice conversion method, characterized in that: include: Input the singing voice file to be converted into the trained phoneme posterior graph model; The trained phoneme posterior graph model is obtained by training according to the phoneme posterior graph model training method according to any one of claims 1 to 6; According to the output of the trained phoneme posterior graph model, the main singing voice content in the song file to be converted is determined so that singing conversion is performed based on the main singing voice content.
8. The singing voice conversion method according to claim 7, wherein: Performing singing voice conversion based on the lead singing voice content includes: Generate fusion features based on the singing melody information of the singing file to be converted, the lead singing voice content and the target timbre file; The fusion feature is decoded into a spectrum using a trained decoder, and the spectrum is converted into audio to achieve singing conversion of the singing file to be converted.
9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor for executing the computer program to implement the phoneme posterior graph model training method according to any one of claims 1 to 6, or the singing voice conversion method according to claim 7 or 8.
10. A computer-readable storage medium, characterized in that Used to store computer programs; wherein when the computer program is executed by a processor, the phoneme posterior graph model training method as described in any one of claims 1 to 6, or the singing voice conversion method as described in claim 7 or 8 is implemented.
11. A computer program product, characterized in that It includes a computer program, which, when executed by a processor, implements the phoneme posterior graph model training method as described in any one of claims 1 to 6, or the singing voice conversion method as described in claim 7 or 8.