Song generation method and apparatus, electronic device, and storage medium

CN116645938BActive Publication Date: 2026-09-25IFLYTEK CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310384260.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-06
Publication Date
2026-09-25
Estimated Expiration
2043-04-06

AI Technical Summary

Technical Problem

[0003]现有技术中,通常采用歌唱合成技术,利用声优预先录制的歌声数据库中的声源,经过一定的算法设计,合成可听的歌曲,但是采用歌唱合成技术制作的合成歌曲存在比较明显的机械感、缺乏真人歌曲中的细腻情感表达及细致演绎,例如咬字、吐息等处理,导致制作出的歌曲质量和效果较低

Benefits of technology

[0058]本申请提出的歌曲制作方法,包括:对源歌曲的声学特征进行编码,得到源歌曲编码信息,源歌曲编码信息包括源歌曲的发音内容信息、源歌手的音色信息以及歌唱细节信息,歌唱细节信息至少包括情感信息、韵律信息、语气信息中的至少一种;基于源歌手的音色特征,从源歌曲编码信息中剔除源歌手的音色信息,得到第一歌曲编码信息;将第一歌曲编码信息与目标歌手的音色特征进行融合,得到第二歌曲编码信息;通过对源歌曲的基频信息以及第二歌曲编码信息进行解码,生成目标歌曲。采用本申请的技术方案,可以只将真人演唱的源歌曲中的音色替换为目标歌手的音色,保留源歌曲中的情感、韵律或语气等细节,从而能够提高歌曲制作的拟人度,提高制作出的歌曲质量和效果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116645938B_ABST
    Figure CN116645938B_ABST
Patent Text Reader

Abstract

The application provides a song generation method and device, electronic equipment and a storage medium. The method comprises the following steps: encoding the acoustic characteristics of a source song to obtain source song encoding information, wherein the source song encoding information comprises pronunciation content information of the source song, timbre information of a source singer, and singing detail information comprising at least one of emotional information, rhythm information and tone information; removing the timbre information of the source singer from the source song encoding information based on the timbre characteristics of the source singer to obtain first song encoding information; fusing the first song encoding information with the timbre characteristics of a target singer to obtain second song encoding information; and decoding the fundamental frequency information of the source song and the second song encoding information to generate a target song. According to the scheme, the timbre of the source song sung by a real person is replaced by the timbre of the target singer, and the details such as emotion, rhythm or tone in the source song are retained, so that the personification degree of song production can be improved, and the quality and effect of the produced song can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a song generation method, apparatus, electronic device, and storage medium. Background Technology

[0002] In recent years, with the upgrading and evolution of digital artificial intelligence technology and the rise and development of the metaverse ecosystem, virtual human technology has matured continuously, the core market size has continued to expand, and more and more virtual humans have entered the public eye, creating new value in industries such as e-commerce, media, finance, culture and tourism, education, and pan-entertainment. Virtual idol singers, as a major category, have a wider range of applications and greater development potential and commercial value compared to real-life idol singers. For the creation of virtual idol singers, song production is particularly important.

[0003] In existing technologies, singing synthesis technology is commonly used. It utilizes the sound sources in a database of pre-recorded vocals by voice actors and synthesizes audible songs through certain algorithm design. However, synthesized songs produced using singing synthesis technology have a relatively obvious mechanical feel and lack the delicate emotional expression and detailed performance of real songs, such as the processing of pronunciation and breathing, resulting in lower quality and effect of the produced songs. Summary of the Invention

[0004] Based on the defects and shortcomings of the prior art, this application proposes a song generation method, apparatus, electronic device and storage medium, which can improve the quality and effect of the produced songs.

[0005] The technical solution proposed in this application is as follows:

[0006] According to a first aspect of the embodiments of this application, a song generation method is provided, including:

[0007] The acoustic features of the source song are encoded to obtain source song encoding information. The source song encoding information includes the pronunciation content information of the source song, the timbre information of the source singer, and the singing details information. The singing details information includes at least one of emotional information, rhythmic information, and intonation information.

[0008] Based on the timbre characteristics of the source singer, the timbre information of the source singer is removed from the source song encoding information to obtain the first song encoding information;

[0009] The first song encoding information is fused with the timbre characteristics of the target singer to obtain the second song encoding information;

[0010] The target song is generated by decoding the baseband information of the source song and the encoding information of the second song.

[0011] Optionally, the process of encoding the acoustic features of the source song to obtain source song encoding information; removing the source singer's timbre information from the source song encoding information based on the source singer's timbre features to obtain first song encoding information; fusing the first song encoding information with the target singer's timbre features to obtain second song encoding information; and generating the target song by decoding the fundamental frequency information of the source song and the second song encoding information, includes:

[0012] The acoustic features of the source song are encoded using a pre-trained song conversion model to obtain source song encoding information. Based on the timbre features of the source singer, the timbre information of the source singer is removed from the source song encoding information to obtain first song encoding information. The first song encoding information is fused with the timbre features of the target singer to obtain second song encoding information. Finally, the target song is generated by decoding the fundamental frequency information of the source song and the second song encoding information.

[0013] Optionally, the song conversion model includes:

[0014] The coding network is used to encode the acoustic features of the source song to obtain the source song coding information;

[0015] A reversible probability distribution model is used to remove the source singer's timbre information from the source song encoding information based on the source singer's timbre characteristics to obtain first song encoding information, and to fuse the first song encoding information with the target singer's timbre characteristics to obtain second song encoding information;

[0016] A decoding network is used to generate a target song by decoding the baseband information of the source song and the encoding information of the second song.

[0017] Optionally, the song conversion model further includes:

[0018] Timbre coding networks are used to encode the acoustic features of a singer's speech to obtain the singer's timbre characteristics.

[0019] Optionally, the training process of the song conversion model aims to minimize the sum of a first loss function and a second loss function. The first loss function is determined based on the difference between the result of the song conversion model recovering the sample source song using the first encoding information and the sample source song. The second loss function is determined based on the difference between the first encoding information and the second encoding information.

[0020] The first encoding information is obtained by encoding the acoustic features of the sample source song by the song conversion model. The second encoding information is obtained by fusing the third encoding information with the timbre information of the source singer corresponding to the sample source song by the song conversion model. The third encoding information is obtained by encoding the features of the sample source song that are unrelated to the speaker's timbre by a preset prior encoding network.

[0021] Optionally, the song conversion model includes an encoding network, a reversible probability distribution model, a decoding network, and a timbre encoding network;

[0022] The training process for the song conversion model includes:

[0023] The acoustic features of the sample source song are encoded through the encoding network to obtain the first probability distribution corresponding to the sample source song, and the timbre features of the source singer are obtained by encoding the acoustic features of the sample source song through the timbre encoding network.

[0024] The features of the sample source song that are independent of the speaker's timbre are input into a preset prior coding network to obtain the third probability distribution of the sample source song.

[0025] The third probability distribution and the timbre features of the source singer are input into the reversible probability distribution model, so that the reversible probability distribution model fuses the third probability distribution with the timbre information of the source singer to obtain the second probability distribution corresponding to the sample source song. The first probability distribution and the fundamental frequency of the sample source song are input into the decoding network to obtain the decoded song.

[0026] A first loss function is determined by comparing the decoded song and the sample source song, and a second loss function is determined by comparing the first probability distribution and the second probability distribution.

[0027] The model parameters of the song conversion model are adjusted with the goal of minimizing the sum of the first loss function and the second loss function.

[0028] Optionally, the fundamental frequency information of the source song includes a standard fundamental frequency and a converted fundamental frequency corresponding to the standard fundamental frequency, wherein the standard fundamental frequency is the fundamental frequency extracted from the source song, and the converted fundamental frequency is obtained by performing a modulation domain conversion on the standard fundamental frequency;

[0029] The step of generating a target song by decoding the fundamental frequency information of the source song and the encoding information of the second song includes:

[0030] For each base frequency in the base frequency information of the source song, the base frequency is decoded with the second song encoding information to obtain a decoded song corresponding to each base frequency in the base frequency information of the source song.

[0031] Select the target song from the various decoded songs.

[0032] Optionally, the vocal characteristics of the target singer include the vocal characteristics of a virtual singer, which are obtained by fusing the vocal characteristics of real singers with different vocal features.

[0033] Optionally, the process of acquiring the virtual singer's vocal characteristics includes:

[0034] Identify at least one intended vocal characteristic of the virtual singer;

[0035] Acquire the timbre characteristics of real singers that match each intended timbre characteristic;

[0036] According to preset weights, the timbre features of real singers that match each desired timbre feature are weighted and fused to obtain the timbre features of virtual singers.

[0037] Optionally, the preset weights include multiple different weight combinations;

[0038] The virtual singer's timbre features are obtained by weighted fusion of the timbre features of real singers that match each intended timbre feature according to preset weights, including:

[0039] According to each weight combination, the timbre features of real singers that match each intention timbre feature are weighted and fused to obtain the candidate fused timbre features corresponding to each weight combination.

[0040] Based on each candidate fusion timbre feature, the song to be converted is timbre converted, and based on the song conversion result, the target fusion timbre feature is selected from the candidate fusion timbre features as the timbre feature of the virtual singer.

[0041] According to a second aspect of the embodiments of this application, a song generation method is provided, including:

[0042] The acoustic features of the source song, the timbre features of the target singer, and the fundamental frequency information of the source song are input into a pre-trained song conversion model. The song conversion model removes the timbre information of the source singer from the source song to obtain source song encoding features that do not contain the timbre information of the source singer. The source song encoding features and the timbre features of the target singer are then used to generate a target song that contains the timbre of the target singer.

[0043] The training process of the song conversion model aims to minimize the sum of the first loss function and the second loss function. The first loss function is determined based on the difference between the result of the song conversion model recovering the sample source song using the first encoding information and the sample source song. The second loss function is determined based on the difference between the first encoding information and the second encoding information.

[0044] The first encoding information is obtained by encoding the acoustic features of the sample source song by the song conversion model. The second encoding information is obtained by fusing the third encoding information with the timbre information of the source singer corresponding to the sample source song by the song conversion model. The third encoding information is obtained by encoding the acoustic features of the sample source song that are unrelated to the speaker's timbre by a preset prior encoding network.

[0045] According to a third aspect of the embodiments of this application, a song generation apparatus is provided, comprising:

[0046] The encoding module is used to encode the acoustic features of the source song to obtain source song encoding information. The source song encoding information includes the pronunciation content information of the source song, the timbre information of the source singer, and the singing details information. The singing details information includes at least one of the following: emotional information, rhythmic information, and tone information.

[0047] The timbre removal module is used to remove the timbre information of the source singer from the source song encoding information based on the timbre characteristics of the source singer, so as to obtain the first song encoding information;

[0048] The timbre fusion module is used to fuse the first song encoding information with the timbre characteristics of the target singer to obtain the second song encoding information;

[0049] The decoding module is used to generate the target song by decoding the baseband information of the source song and the encoding information of the second song.

[0050] According to a fourth aspect of the embodiments of this application, a song generation apparatus is provided, comprising:

[0051] The feature input module is used to input the acoustic features of the source song, the timbre features of the target singer, and the fundamental frequency information of the source song into a pre-trained song conversion model, so that the song conversion model removes the timbre information of the source singer from the source song, obtains the source song encoding features that do not contain the timbre information of the source singer, and uses the source song encoding features and the timbre features of the target singer to generate a target song that contains the timbre of the target singer;

[0052] The training process of the song conversion model aims to minimize the sum of the first loss function and the second loss function. The first loss function is determined based on the difference between the result of the song conversion model recovering the sample source song using the first encoding information and the sample source song. The second loss function is determined based on the difference between the first encoding information and the second encoding information.

[0053] The first encoding information is obtained by encoding the acoustic features of the sample source song by the song conversion model. The second encoding information is obtained by fusing the third encoding information with the timbre information of the source singer corresponding to the sample source song by the song conversion model. The third encoding information is obtained by encoding the acoustic features of the sample source song that are unrelated to the speaker's timbre by a preset prior encoding network.

[0054] According to a fifth aspect of the embodiments of this application, an electronic device is provided, including: a memory and a processor;

[0055] The memory is connected to the processor and is used to store programs;

[0056] The processor is used to implement the above-described song generation method by running the program in the memory.

[0057] According to a sixth aspect of the embodiments of this application, a storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, it implements the above-described song generation method.

[0058] The song production method proposed in this application includes: encoding the acoustic features of a source song to obtain source song encoding information, which includes the pronunciation content information of the source song, the timbre information of the source singer, and singing detail information, wherein the singing detail information includes at least one of emotional information, rhythmic information, and tone information; based on the timbre features of the source singer, removing the timbre information of the source singer from the source song encoding information to obtain first song encoding information; fusing the first song encoding information with the timbre features of the target singer to obtain second song encoding information; and generating the target song by decoding the fundamental frequency information of the source song and the second song encoding information. By adopting the technical solution of this application, only the timbre in a source song sung by a real person can be replaced with the timbre of the target singer, while retaining details such as emotion, rhythm, or tone in the source song, thereby improving the anthropomorphism of the song production and enhancing the quality and effect of the produced song. Attached Figure Description

[0059] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0060] Figure 1 This is a flowchart illustrating a song production method provided in an embodiment of this application;

[0061] Figure 2 This is a schematic diagram of the application structure of the song conversion model provided in the embodiments of this application;

[0062] Figure 3 This is a schematic diagram of the processing flow for training the song conversion model provided in an embodiment of this application;

[0063] Figure 4 This is a schematic diagram of the training structure of the song conversion model provided in the embodiments of this application;

[0064] Figure 5 This is a schematic diagram of the processing flow for determining the timbre characteristics of a virtual singer provided in an embodiment of this application;

[0065] Figure 6 This is a flowchart illustrating another song production method provided in an embodiment of this application;

[0066] Figure 7 This is a schematic diagram of the structure of a song production device provided in an embodiment of this application;

[0067] Figure 8 This is a schematic diagram of another song production device provided in an embodiment of this application;

[0068] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0069] The technical solutions of this application are applicable to the application scenarios of digital virtual humans, especially to the application scenarios of song production for virtual singers. By adopting the technical solutions of this application, the anthropomorphism of song production can be improved, and the quality and effect of the produced songs can be enhanced.

[0070] With the upgrading and evolution of digital artificial intelligence technology and the rise and development of the metaverse ecosystem, virtual singers have increasingly broad development prospects in the pan-entertainment industry. For virtual singers, song production is essential. Current technology typically uses vocal synthesis to create songs. In this process, a voice actor, acting as the "vest person," records a certain number of sound sources and stores them in a vocal database. Then, through a specific algorithm design, sound sources are selected from the vocal database and synthesized using vocal synthesis technology. Here, the "vest person" refers to the person who manipulates the virtual character for live streaming, and also broadly refers to the worker who provides the sound source. However, songs synthesized using various sound sources recorded by the "vest person" often have a noticeable mechanical feel, lacking the delicate emotional expression and nuanced performance of real human songs, such as pronunciation and breathing techniques. As a result, the synthesized songs still have a significant gap in quality compared to real human voices, leading to lower overall quality and effect.

[0071] Furthermore, each virtual singer corresponds to a unique timbre. In existing technologies, the timbre of a virtual singer is typically used as the timbre of a person behind the virtual singer. When creating a song for that virtual singer, the song is synthesized using the voice source recorded by the person behind the virtual singer. If the voice source of that person cannot be used (for example, if the person behind the virtual singer terminates the contract and requests that the right to use their voice source no longer be granted), it will result in the inability to create a song for that virtual singer, affecting the stability of song production.

[0072] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0073] Exemplary methods

[0074] This application provides a song production method, which can be executed by an electronic device. This electronic device can be any device with data and instruction processing capabilities, such as a computer, a smart terminal, or a server. See also... Figure 1 As shown, the method includes:

[0075] S101. Encode the acoustic features of the source song to obtain the source song encoding information.

[0076] This embodiment aims to convert a source song sung by a source singer into a target song sung by a target singer. The source singer is a real singer, whose singing can express the song's emotions, rhythm, tone, and other details. The target singer can be a singer who cannot accurately express the song's details, or it can be a virtual singer. Converting the source song sung by the source singer into a target song sung by the target singer allows the target song to reflect the target singer's timbre while retaining the emotions, rhythm, tone, and other details of the source song.

[0077] Specifically, in this embodiment, the acoustic features of the source song are first encoded to obtain the source song encoding information. The acoustic features of the source song can be speech waveform features, linear amplitude spectrum features, Mel spectrum features, etc., and this embodiment does not impose specific limitations. To ensure the accuracy of the acoustic features of the source song, the source song is preferably dry vocal data, i.e., pure human voice without music and without post-processing. The source singer needs to record the live performance of the source song in a clean recording studio environment to ensure that the obtained source song is a high-fidelity live performance. Since a live singer's performance can reveal singing details, such as emotion, rhythm, or tone, the source song performed by the source singer not only contains the pronunciation content information and the singer's timbre information, but also the singing details. Therefore, the source song encoding information after encoding the acoustic features also includes the pronunciation content information, the singer's timbre information, and singing details, where the singing details include at least one of emotional information, rhythmic information, and tone information.

[0078] The source singer is preferably a real person who can be selected, and this real person can be a singer with certain singing skills to ensure the quality of the source song, thereby improving the quality of the converted target song.

[0079] S102. Based on the timbre characteristics of the source singer, remove the timbre information of the source singer from the source song encoding information to obtain the first song encoding information.

[0080] In this embodiment, when generating a song, in order to ensure that the target song of the target singer retains the singing details of the source song, it is only necessary to replace the timbre of the source song with the timbre of the target singer. Therefore, this embodiment first needs to determine the timbre characteristics of the source singer, and then, based on the timbre characteristics of the source singer, remove the timbre information of the source singer from the source song encoding information corresponding to the source song, thereby obtaining the first song encoding information that does not contain timbre information (i.e., is unrelated to timbre). The first song encoding information only removes the timbre information, so it still contains the singing details of the source singer.

[0081] To determine the timbre characteristics of the source singer, this embodiment can use a timbre encoding network for timbre feature extraction. The acoustic features of the source song or other songs sung by the source singer are input into the timbre encoding network, which can then extract the timbre features from these acoustic features, thereby obtaining the timbre characteristics of the source singer.

[0082] S103. The encoding information of the first song is fused with the timbre characteristics of the target singer to obtain the encoding information of the second song.

[0083] Specifically, in this embodiment, after obtaining the first song encoding information that does not contain timbre information, the timbre features of the target singer are fused into the first song encoding information to obtain the second song encoding information. This second song encoding information contains both the timbre information of the target singer and the singing details of the source singer. When the target singer is a real singer, the timbre features of the target singer can be extracted from the acoustic features of the song sung by the target singer using a timbre encoding network. When the target singer is a virtual singer, the virtual singer's image positioning can be preset, and the virtual singer's timbre features can be determined based on the timbre features of a real singer that match the virtual singer's image positioning. The timbre features of the real singer can also be extracted from the acoustic features of the song sung by the real singer using a timbre encoding network.

[0084] S104. Generate the target song by decoding the baseband information of the source song and the encoding information of the second song.

[0085] Specifically, in this embodiment, after obtaining the second song encoding information containing the timbre information of the target singer and the singing details of the source singer, it is necessary to decode the second song encoding information. However, during the decoding process of the second song encoding information, problems such as fundamental frequency jitter and discontinuity may occur, affecting the pitch accuracy of the generated song. In the song production process, pitch accuracy is one of the basic factors affecting the quality of the produced song, and the fundamental frequency information of the song is the data that reflects the pitch accuracy of the song. In song production, the pitch accuracy of the song can be guaranteed based on the corresponding fundamental frequency information. Therefore, in order to ensure that the produced song has a stable pitch accuracy, this embodiment needs to extract the fundamental frequency of the source song to obtain the fundamental frequency information of the source song. While decoding the second song encoding information, the pitch accuracy is controlled using the fundamental frequency information of the source song, so that the target song generated after decoding has a stable and continuous fundamental frequency, which conforms to the pitch accuracy corresponding to the fundamental frequency information of the source song, thereby improving the pitch accuracy of the target song. Furthermore, the generated target song not only contains the timbre of the target singer, but also the singing details of the source singer, that is, it retains the delicate emotions and performance style of the source song, which can ensure the anthropomorphism of the song and improve the quality and effect of the produced song.

[0086] For the fundamental frequency extraction of the source song, this embodiment can use existing speech signal processing tools, such as STRAIGHT, Praat, WORLD, etc., or a pre-trained neural network-based fundamental frequency extraction model. This fundamental frequency extraction model is an existing model, and will not be described in detail in this embodiment.

[0087] Furthermore, when using the aforementioned fundamental frequency extraction method to extract the fundamental frequency of the source song, inaccurate extraction may occur, such as errors in voiced / unvoiced consonants, half / second harmonic frequencies, and inaccurate extraction of guttural sounds, thus affecting the quality of the produced song. Therefore, after extracting the fundamental frequency information of the source song, this embodiment needs to check and correct the fundamental frequency information to ensure that it is the same as the fundamental frequency of the source song. In this embodiment, the fundamental frequency check and correction can be performed manually to obtain the corrected fundamental frequency information.

[0088] Furthermore, in order to achieve a better audiovisual experience from the generated target song, further post-processing can be performed on the target song, such as mixing and arrangement. The song's effects can also be adjusted and improved to produce a song that meets release standards.

[0089] As described above, the song generation method proposed in this application encodes the acoustic features of a source song to obtain source song encoding information. This source song encoding information includes the pronunciation content information of the source song, the timbre information of the source singer, and singing detail information. The singing detail information includes at least one of emotional information, rhythmic information, and intonation information. Based on the timbre features of the source singer, the timbre information of the source singer is removed from the source song encoding information to obtain first song encoding information. The first song encoding information is then fused with the timbre features of the target singer to obtain second song encoding information. Finally, the target song is generated by decoding the fundamental frequency information of the source song and the second song encoding information. Using the technical solution of this embodiment, only the timbre in a source song sung by a real person can be replaced with the timbre of the target singer, while retaining details such as emotion, rhythm, or intonation in the source song. This improves the anthropomorphism of the song production and enhances the quality and effect of the produced song.

[0090] As an optional implementation, another embodiment of this application discloses a song generation method, including:

[0091] The acoustic features of the source song are encoded using a pre-trained song conversion model to obtain source song encoding information. Based on the timbre features of the source singer, the timbre information of the source singer is removed from the source song encoding information to obtain first song encoding information. The first song encoding information is fused with the timbre features of the target singer to obtain second song encoding information. Finally, the target song is generated by decoding the fundamental frequency information of the source song and the second song encoding information.

[0092] Specifically, in this embodiment, a song conversion model can be pre-trained. Then, the acoustic features of the source song, the timbre features of the source singer, the timbre features of the target singer, and the fundamental frequency information of the source song are all input into the song conversion model. The model encodes the acoustic features of the source song to obtain source song encoding information. Based on the timbre features of the source singer, the timbre information is removed from the source song encoding information to obtain first song encoding information. The first song encoding information is then fused with the timbre features of the target singer to obtain second song encoding information. Finally, the fundamental frequency information of the source song and the second song encoding information are decoded to generate the target song. The process of generating the target song using the song conversion model has been specifically described in the above embodiments and will not be repeated in this embodiment.

[0093] See Figure 2 As shown, in this embodiment, the song conversion model includes: an encoding network, a reversible probability distribution model, and a decoding network. During the application of the song conversion model, the acoustic features of the source song are... The input is fed into an encoding network, which processes the acoustic features of the source song. Encoding is performed to obtain the source song encoding information z, and this source song encoding information z is then transmitted to the reversible probability distribution model. The acoustic features of the source song are... Features can be speech waveform features, linear amplitude spectrum features, Mel spectrum features, etc. In this embodiment, the timbre features s0 of the source singer and s0 of the target singer are also input into the reversible probability distribution model. The reversible probability distribution model first performs a forward transform on the source song encoding information z based on the timbre features s0 of the source singer, removing the timbre information of the source singer from the source song encoding information z, thus obtaining first song encoding information that does not contain timbre information. Then, based on the timbre features s0 of the target singer, it performs an inverse transform on the first song encoding information that does not contain timbre information, fusing the timbre features s0 of the target singer with the first song encoding information to obtain second song encoding information that contains the timbre information of the target singer. Encode the second song information The data is transmitted to the decoding network. The decoding network processes the source song's fundamental frequency information f and the second song's encoding information. Decode the target song to generate its acoustic features. This allows the acquisition of the target song. In this embodiment, the encoding and decoding networks are the same as those in the variational autoencoder. The timbre encoding network can be an existing network for timbre extraction from speech, and the reversible probability distribution model can be the Glow model or the NICE stream model, etc., which will not be elaborated on in this embodiment.

[0094] Furthermore, in this embodiment, the song conversion model may also include a timbre encoding network. This timbre encoding network can encode the acoustic features of the singer's speech to obtain the singer's timbre features. In this embodiment, the acoustic features of the source song are... The input is fed into a timbre coding network, which processes the acoustic features of the source song. Encoding is performed to obtain the timbre features s0 of the source singer, and these timbre features s0 are then input into a reversible probability distribution model. The acoustic features of the source song input into the timbre encoding network are... It can also be speech waveform features, linear amplitude spectrum features, Mel spectrum features, etc., and the acoustic features of the source song input into the timbre coding network. Acoustic features of the source song input into the encoding network It can be the same type of acoustic feature or different types of acoustic features, for example, the acoustic features of the source song input into the timbre coding network. The Mel-spectral features are used for inputting the acoustic features of the source song into the encoding network. These are speech waveform features, or acoustic features of the source song input into the timbre coding network. Acoustic features of the source song input into the encoding network All of these are speech waveform features.

[0095] As an optional implementation, another embodiment of this application discloses that the training process of the song conversion model aims to minimize the sum of a first loss function and a second loss function. The first loss function is used to train the encoding and decoding capabilities of the encoding and decoding networks, and the second loss function is used to train the timbre conversion capability of the reversible probability distribution model. The first loss function is determined based on the difference between the result of the song conversion model recovering the sample source song using the first encoded information and the sample source song itself. The second loss function is determined based on the difference between the first and second encoded information. The first encoded information is obtained by encoding the acoustic features of the sample source song using the encoding network in the song conversion model. The second encoded information is obtained by fusing the third encoded information with the timbre information of the source singer corresponding to the sample source song using the reversible probability distribution model in the song conversion model. The third encoded information is obtained by encoding features of the sample source song that are unrelated to the speaker's timbre using a pre-set prior encoding network.

[0096] In this embodiment, the features of the source song that are independent of the speaker's timbre can be the acoustic features of a song with the same content as the source song sung by a real singer, excluding timbre features. This real singer can be the singer of the source song or another singer. Extracting the acoustic features of the song that do not include timbre features can be done using an existing speech recognition acoustic model to extract the bottleneck features corresponding to the song. These bottleneck features are the acoustic features of the song that do not include timbre features. Alternatively, the features of the source song that are independent of the speaker's timbre can also be the text features of the speech content corresponding to the source song. In this case, the third encoded information obtained by encoding the text features of the speech content corresponding to the source song by the preset prior coding network needs to be length-aligned with the first encoded information to obtain the aligned third encoded information.

[0097] As an optional implementation, see [link to relevant documentation]. Figure 3 and Figure 4 As shown, in another embodiment of this application, the training process of a song conversion model is disclosed, including:

[0098] S301. The acoustic features of the sample source song are encoded through an encoding network to obtain the first probability distribution corresponding to the sample source song. The acoustic features of the sample source song are encoded through a timbre encoding network to obtain the timbre features of the source singer.

[0099] Specifically, before training the song conversion model, a large amount of dry vocal data needs to be collected as sample source songs for training the model. The collected dry vocal data must be clean, recorded in a high-fidelity recording studio environment, preferably in 48kHz / 16bit format. When collecting the dry vocal data, it is preferable to collect data from hundreds of singers, with a balanced number of male and female singers. Each singer needs to record at least three songs, and the recorded songs should be segmented into phrases, with each phrase lasting at least five seconds. These segmented phrases are then used as sample source songs, thus obtaining a large number of sample source songs.

[0100] After acquiring the sample source song, this embodiment needs to extract the acoustic features of the sample source song. Among them, the acoustic features of the sample source songs This can include speech waveform features, linear amplitude spectrum features, Mel spectrum features, etc. The extracted acoustic features of the source songs will be used. The input is fed into the encoding network of the song conversion model, which utilizes the acoustic features of the sample source song. Encoding is performed to obtain the first probability distribution z' corresponding to the sample source song, which is the first encoding information in the above embodiment.

[0101] This embodiment also requires extracting the acoustic features of the source song. Among them, the acoustic features of the sample source songs Acoustic characteristics of the sample source songs The acoustic features can be of the same type or different types, i.e., the acoustic features of the source songs in the sample. It can also be speech waveform features, linear amplitude spectrum features, Mel spectrum features, etc. The extracted acoustic features of the source songs are then analyzed. The input is fed into the timbre encoding network in the song conversion model, which then uses the timbre encoding network to analyze the acoustic features of the sample source song. Encoding is performed to extract the acoustic features of the sample source song. The timbre features in the sample song are used to obtain the timbre features s0′ of the source singer.

[0102] S302. Input the features of the source song that are independent of the speaker's timbre into a preset prior coding network to obtain the third probability distribution of the source song.

[0103] This embodiment also requires pre-collecting features y that are independent of the speaker's timbre corresponding to the sample source song, including: acoustic features that do not contain timbre features corresponding to songs sung by real singers with the same content as the sample source song, or text features of the speech content corresponding to the sample source song. The features y that are independent of the speaker's timbre corresponding to the sample source song are input into a preset prior coding network to obtain the third probability distribution y' corresponding to the sample source song, which is the third coding information in the above embodiment. In this embodiment, the third probability distribution y' output by the prior coding network can be a Gaussian mixture density function. The parameter π k μ k σ k Where M is the number of Gaussian distributions, and π k μ k σ k represents the weights, mean, and standard deviation of the k-th Gaussian distribution.

[0104] When the feature y, which is independent of the speaker's timbre, corresponding to the sample source song input into the preset prior coding network is a text feature of the speech content corresponding to the sample source song, the third probability distribution corresponding to the sample source song output by the prior coding network needs to be length-aligned with the first probability distribution z' corresponding to the sample source song output by the coding network to obtain the aligned third probability distribution y'. The length alignment between the third probability distribution and the first probability distribution z' expands the probability distribution sequence according to duration information, thus achieving alignment. This duration information can be obtained manually based on the sample source song and its corresponding speech content (i.e., lyrics), or it can be automatically obtained using forced alignment (FA).

[0105] S303. Input the third probability distribution and the timbre features of the source singer into the reversible probability distribution model so that the reversible probability distribution model can fuse the third probability distribution with the timbre information of the source singer to obtain the second probability distribution corresponding to the sample source song. Then, input the first probability distribution and the fundamental frequency of the sample source song into the decoding network to obtain the decoded song.

[0106] In this embodiment, the third probability distribution y' output by the prior coding network and the timbre features s0' of the source singer of the singing sample song are input into the reversible probability distribution model. Based on the timbre features s0' of the source singer of the singing sample song, the reversible probability distribution model performs an inverse transformation on the third probability distribution y' corresponding to the source song, fusing the third probability distribution y' with the timbre information of the source singer of the singing sample song, thereby obtaining a second probability distribution containing timbre information.

[0107] This embodiment also requires inputting the first probability distribution z' corresponding to the sample source song output by the encoding network into the decoding network for decoding. Since the fundamental frequency range of the singing voice is large, the decoding network recovers the robust original waveform (wherein, the original waveform is the acoustic feature of the sample source song input into the encoding network). This is quite difficult and may result in issues such as fundamental frequency jitter and discontinuity. Therefore, it is necessary to further input the pre-extracted fundamental frequency f′ of the sample source song into the decoding network to synthesize the speech waveform (i.e., generate the acoustic features corresponding to the decoded song) when decoding the first probability distribution z′. The control of the base frequency in the decoding network allows the voice waveform recovered by the decoding network to have a stable and continuous base frequency.

[0108] Specifically, the decoding network decodes the first probability distribution z' based on the pre-extracted fundamental frequency f′ of the sample source song, thereby obtaining the acoustic features corresponding to the decoded song. Thus, the acoustic characteristics can be obtained The song to be decoded is determined. The method for extracting the fundamental frequency of the sample source song is the same as that used in the previous embodiment, and will not be described in detail here.

[0109] S304. By comparing the decoded song and the sample source song, determine the first loss function, and by comparing the first probability distribution and the second probability distribution, determine the second loss function.

[0110] This embodiment compares the decoded song with the sample source song to determine the first loss function as the reconstruction loss function of the song conversion model; that is, it calculates the acoustic features of the sample source song. Acoustic characteristics of decoding songs The loss function between the first probability distribution z' and the second probability distribution. This embodiment also compares the loss function between the first probability distribution z' and the second probability distribution. Determine the second loss function, i.e., calculate the first probability distribution z' and the second probability distribution. The KL divergence between the two is used as the second loss function.

[0111] S305. Adjust the model parameters of the song conversion model with the goal of minimizing the sum of the first loss function and the second loss function.

[0112] In this embodiment, the sum of the first loss function and the second loss function is used as the total loss function of the song conversion model. The model parameters are adjusted to minimize the total loss function, thereby improving the acoustic characteristics of the source songs. Acoustic characteristics of decoding songs As close as possible to the first probability distribution z' and the second probability distribution Get as close as possible.

[0113] As an optional implementation, another embodiment of this application discloses the fundamental frequency information of the source song, including a standard fundamental frequency and a converted fundamental frequency corresponding to the standard fundamental frequency. The standard fundamental frequency is extracted from the source song, and the converted fundamental frequency is obtained by modulating the standard fundamental frequency.

[0114] Specifically, different singers may require different fundamental frequencies. To adapt to the target singer's timbre and achieve the best listening experience, this embodiment can perform a modulation conversion on the standard fundamental frequency of the source song to obtain at least one converted fundamental frequency. Since the song's intervals conform to the twelve-tone equal temperament, unlike traditional modulation methods based on a single Gaussian, this embodiment uses either a pitch shift or a pitch drop to perform modulation conversion on the standard fundamental frequency of the source song to ensure accurate pitch. The specific formula is as follows:

[0115]

[0116] Among them, f trg f represents the fundamental frequency after modulation-domain conversion. src The baseband frequency of the original song before key conversion is represented by n, where n is the pitch shift value, and -11 ≤ n ≤ 11, and n is an integer. When n > 0, it indicates a key conversion using a rising pitch; when n < 0, it indicates a key conversion using a falling pitch. In this embodiment, three sets of n values ​​are preferably set: -1, 0, and 1, to obtain the conversion baseband frequency corresponding to each n value. When n is 0, the corresponding conversion baseband frequency is the standard baseband frequency.

[0117] Further, in the above embodiment, step S104, which generates the target song by decoding the fundamental frequency information of the source song and the encoding information of the second song, specifically includes:

[0118] First, for each fundamental frequency in the fundamental frequency information of the source song, decode that fundamental frequency and the second song encoding information to obtain the decoded song corresponding to each fundamental frequency in the fundamental frequency information of the source song.

[0119] Specifically, regarding the encoding information of the second song During decoding, each baseband frequency in the source song's baseband information needs to be decoded once to obtain the decoded song corresponding to each baseband frequency. Specifically, based on each baseband frequency, the encoding information of the second song is... The decoding process is the same as the process described in step S104 of the above embodiment, and will not be described in detail in this embodiment.

[0120] Second, select the target song from the various decoded songs.

[0121] Specifically, after obtaining the decoded songs corresponding to each baseband frequency, one decoded song needs to be selected as the target song. When selecting the target song, the best-sounding song can be chosen based on its perceived quality. For example, several listeners can be selected to rate the listenability of the decoded songs corresponding to each baseband frequency, and the average score of each decoded song can be calculated. The decoded song with the highest average score can then be selected as the target song. In this embodiment, the specific execution process of further post-processing the target song is the same as that of step S104 in the above embodiment, and will not be described in detail here.

[0122] This embodiment improves the listening experience and overall quality of the produced song by performing a modulation domain conversion on the standard fundamental frequency of the source song and selecting the song with the best listening effect from the decoded songs corresponding to multiple fundamental frequencies.

[0123] As an optional implementation, in this embodiment, the target singer includes a real singer or a virtual singer. When the target singer is a virtual singer, the timbre characteristics of the target singer include: the timbre characteristics of the virtual singer. The timbre characteristics of the virtual singer can be the timbre characteristics of a real singer that match the virtual singer's image positioning, or they can be obtained by fusing the timbre characteristics of real singers with different timbre features. The timbre characteristics of the real singer can also be obtained by using the timbre encoding network in the song conversion model of the above embodiment to encode the timbre of the input real singer's voice (such as a song sung by a real singer).

[0124] As an optional implementation, see [link to relevant documentation]. Figure 5 As shown, another embodiment of this application discloses a process for obtaining the vocal characteristics of a virtual singer when the target singer is a virtual singer, including:

[0125] S501. Determine at least one intentional timbre characteristic of the virtual singer.

[0126] In this embodiment, a pre-defined image positioning is set for the virtual singer. Different image positioning characteristics correspond to different intended vocal characteristics. For example, if the virtual singer's image positioning is: a young female voice, neutral, bright, and slightly magnetic, then the virtual singer's image positioning characteristics include: a neutral young female voice, a bright young female voice, and a slightly magnetic young female voice. In this case, the virtual singer's intended vocal characteristics include: a neutral young female voice, a bright young female voice, and a slightly magnetic young female voice. When setting the image positioning of the virtual singer, at least one characteristic is set; therefore, the virtual singer's intended vocal characteristics are at least one.

[0127] S502. Obtain the timbre characteristics of real singers that match each intended timbre characteristic.

[0128] In this embodiment, after determining the intended vocal characteristics of the virtual singer, it is necessary to identify a real singer who matches the intended vocal characteristics and collect the dry vocal recordings of the songs performed by the real singers who match the intended vocal characteristics. The dry vocal recordings of the real singers are the recordings of any song performed by the real singer in a clean recording studio environment. When the intended vocal characteristics of the virtual singer include: a neutral young female voice, a bright young female voice, and a slightly magnetic young female voice, the selected real singers can be real people with a neutral young female voice, a bright young female voice, and a magnetic young female voice. The dry vocal recordings of the above three real singers are collected, namely, dry vocal recording y1 of a neutral young female voice, dry vocal recording y2 of a bright young female voice, and dry vocal recording y3 of a magnetic young female voice. Furthermore, the duration of each collected dry vocal recording needs to be at least five seconds.

[0129] After collecting the dry vocal samples of a real singer's song that match each intended timbre characteristic, the timbre encoding network in the song conversion model described above can be used to encode the acoustic features of the input dry vocal samples of the real singer's song, thereby obtaining the timbre characteristics of the real singer that match each intended timbre characteristic.

[0130] S503. According to the preset weights, the timbre features of real singers that match each intentional timbre feature are weighted and fused to obtain the timbre features of virtual singers.

[0131] In this embodiment, preset weights are set, which include the weights corresponding to the timbre features of real singers that match each desired timbre feature, and the sum of all the weights in the preset weights is 1. According to the preset weights, the timbre features of real singers that match each desired timbre feature are weighted and fused to obtain the timbre features of the virtual singer.

[0132] Furthermore, this step specifically includes:

[0133] First, according to each weight combination, the timbre features of real singers that match each intentional timbre feature are weighted and fused to obtain the candidate fused timbre features corresponding to each weight combination.

[0134] The preset weights can include at least one set of weight combinations. Each weight combination includes the weights corresponding to the timbre features of a real singer that match each intended timbre characteristic, and the sum of the weights in each weight combination is 1. For example, the timbre feature s1 corresponds to the dry voice y1 of a neutral young female voice song, and the weight λ1 corresponds to the timbre feature s1; the timbre feature s2 corresponds to the dry voice y2 of a bright young female voice song, and the weight λ2 corresponds to the timbre feature s2; the timbre feature s3 corresponds to the dry voice y3 of a magnetic young female voice song, and the weight λ3 corresponds to the timbre feature s3, where λ1, λ2, and λ3 are all greater than or equal to 0. At this point, the timbre feature weight combinations can be {λ1=0.1, λ2=0.1, λ3=0.8}, {λ1=0.2, λ2=0.6, λ3=0.2}, {λ1=0.4, λ2=0.3, λ3=0.3}, {λ1=0.3, λ2=0.2, λ3=0.5}, etc. This embodiment can iterate through the combined values ​​of λ1, λ2, and λ3 (i.e., weight combinations) with a certain precision. In the example above, the weight combination is determined with a precision of 0.1. This embodiment can also determine the weight combination with a precision of 0.01. This embodiment does not limit the precision.

[0135] In this embodiment, the timbre features of real singers matching each desired timbre characteristic are weighted and fused according to each weight combination to obtain candidate fused timbre features corresponding to each weight combination. For example, the calculation formula for candidate fused timbre features is:

[0136] s=λ1*s1+λ2*s2+λ3*s3

[0137] Where s is the candidate fused timbre feature, which is the timbre feature of the real singer that matches each intention timbre feature, including s1, s2, and s3. When the weight combination is {λ1, λ2, λ3}, the timbre feature is obtained by weighted fusion.

[0138] Second, based on each candidate fusion timbre feature, the song to be converted is timbre converted, and based on the song conversion result, the target fusion timbre feature is selected from the candidate fusion timbre features as the timbre feature of the virtual singer.

[0139] To ensure that the selected virtual singer's target timbre is more listenable and better matches the virtual singer's image positioning, this embodiment can prepare a dry version of a song as the song to be converted. The timbre of this song is then converted, specifically, the timbre features in the acoustic characteristics of the song to be converted are transformed into various candidate fusion timbre features, thus obtaining the song conversion results corresponding to each candidate fusion timbre feature. In this embodiment, the song conversion method described in the above embodiments (such as the song conversion model in the above embodiments) can be used to convert the timbre features in the acoustic characteristics of the song to be converted into candidate fusion timbre features. Then, a group of listeners are organized to listen to the song conversion results corresponding to each candidate fusion timbre feature. Scores are then given based on multiple dimensions, including the degree of matching with the virtual singer's image positioning and the listenability of the timbre. Based on the scoring results, the virtual singer's timbre features are selected from all candidate fusion timbre features, thereby improving the matching degree between the virtual singer's target timbre and the virtual singer's image positioning, as well as the listenability of the virtual singer's target timbre, and ultimately improving the quality and effect of the produced song.

[0140] Furthermore, since the person behind the voice owns the copyright to their voice, if a virtual singer's target voice uses the voice of a certain person behind the voice, then with the termination of the contract by that person, the virtual singer will also lose the right to use that person's voice, thus affecting the virtual singer's song production. However, this embodiment uses multiple voice features to fuse into a new voice feature. This fused voice feature is not the voice of any person behind the voice, and does not involve copyright issues, thus improving the stability of virtual singer's song production.

[0141] This application also discloses a song generation method, which can be executed by an electronic device. The electronic device can be any device with data and instruction processing capabilities, such as a computer, a smart terminal, or a server. See also... Figure 6 As shown, the method includes:

[0142] S601. Input the acoustic features of the source song, the timbre features of the target singer, and the fundamental frequency information of the source song into a pre-trained song conversion model so that the song conversion model removes the timbre information of the source singer from the source song, obtains the source song encoding features that do not contain the timbre information of the source singer, and uses the source song encoding features and the timbre features of the target singer to generate a target song that contains the timbre of the target singer.

[0143] In this embodiment, the training process of the song conversion model aims to minimize the sum of the first loss function and the second loss function. The first loss function is determined based on the difference between the result of the song conversion model recovering the sample source song using the first encoding information and the sample source song itself. The second loss function is determined based on the difference between the first encoding information and the second encoding information. The first encoding information is obtained by the song conversion model encoding the acoustic features of the sample source song. The second encoding information is obtained by the song conversion model fusing the third encoding information with the timbre information of the source singer corresponding to the sample source song. The third encoding information is obtained by a preset prior encoding network encoding the acoustic features of the sample source song that are unrelated to the speaker's timbre. In this step, the specific execution method of the song conversion model application and the training process of the song conversion model are the same as in the above embodiment, and will not be described in detail in this embodiment.

[0144] Exemplary device

[0145] Corresponding to the above-described song generation method, this application also discloses a song generation apparatus, see [link to apparatus]. Figure 7 As shown, the device includes:

[0146] The encoding module 100 is used to encode the acoustic features of the source song to obtain source song encoding information. The source song encoding information includes the pronunciation content information of the source song, the timbre information of the source singer, and the singing details information. The singing details information includes at least one of the following: emotional information, rhythmic information, and tone information.

[0147] The timbre removal module 110 is used to remove the timbre information of the source singer from the source song encoding information based on the timbre characteristics of the source singer, so as to obtain the first song encoding information;

[0148] The timbre fusion module 120 is used to fuse the first song encoding information with the timbre characteristics of the target singer to obtain the second song encoding information;

[0149] The decoding module 130 is used to generate the target song by decoding the base frequency information of the source song and the encoding information of the second song.

[0150] The song generation apparatus proposed in this application includes an encoding module 100 that encodes the acoustic features of a source song to obtain source song encoding information. This source song encoding information includes the pronunciation content information of the source song, the timbre information of the source singer, and singing detail information. The singing detail information includes at least one of emotional information, rhythmic information, and intonation information. A timbre removal module 110 removes the timbre information of the source singer from the source song encoding information based on the timbre features of the source singer, obtaining first song encoding information. A timbre fusion module 120 fuses the first song encoding information with the timbre features of the target singer to obtain second song encoding information. A decoding module 130 decodes the fundamental frequency information of the source song and the second song encoding information to generate a target song. Using the technical solution of this embodiment, only the timbre in a source song sung by a real person can be replaced with the timbre of the target singer, while retaining details such as emotion, rhythm, or intonation in the source song. This improves the anthropomorphism of the song production and enhances the quality and effect of the produced song.

[0151] As an optional implementation, another embodiment of this application also discloses a song generation apparatus, which further includes an input module.

[0152] The model application module is used to encode the acoustic features of the source song using a pre-trained song conversion model to obtain source song encoding information. Based on the timbre features of the source singer, the timbre information of the source singer is removed from the source song encoding information to obtain first song encoding information. The first song encoding information is fused with the timbre features of the target singer to obtain second song encoding information. Finally, the target song is generated by decoding the fundamental frequency information of the source song and the second song encoding information.

[0153] As an optional implementation, another embodiment of this application also discloses a song conversion model, including:

[0154] The coding network is used to encode the acoustic features of the source song to obtain the source song coding information;

[0155] A reversible probability distribution model is used to remove the source singer's timbre information from the source song encoding information based on the source singer's timbre features to obtain the first song encoding information, and to fuse the first song encoding information with the target singer's timbre features to obtain the second song encoding information.

[0156] The decoding network is used to generate the target song by decoding the baseband information of the source song and the encoding information of the second song.

[0157] As an optional implementation, another embodiment of this application also discloses a song conversion model, which further includes:

[0158] Timbre coding networks are used to encode the acoustic features of a singer's speech to obtain the singer's timbre characteristics.

[0159] As an optional implementation, another embodiment of this application also discloses that the training process of the song conversion model aims to minimize the sum of the first loss function and the second loss function. The first loss function is determined based on the difference between the result of the song conversion model recovering the sample source song using the first encoding information and the sample source song. The second loss function is determined based on the difference between the first encoding information and the second encoding information.

[0160] The first encoding information is obtained by encoding the acoustic features of the sample source song by the song conversion model. The second encoding information is obtained by fusing the third encoding information with the timbre information of the source singer corresponding to the sample source song by the song conversion model. The third encoding information is obtained by encoding the features of the sample source song that are unrelated to the speaker's timbre by a preset prior encoding network.

[0161] As an optional implementation, another embodiment of this application also discloses that the song conversion model includes an encoding network, a reversible probability distribution model, a decoding network, and a timbre encoding network, and the song generation device further includes: a model training module, wherein the model training module is used for:

[0162] The acoustic features of the sample source song are encoded by an encoding network to obtain the first probability distribution corresponding to the sample source song. The timbre features of the source singer are obtained by encoding the acoustic features of the sample source song by a timbre encoding network.

[0163] The features of the sample source song that are unrelated to the speaker's timbre are input into a pre-defined prior coding network to obtain the third probability distribution of the sample source song.

[0164] The third probability distribution and the timbre features of the source singer are input into the reversible probability distribution model so that the reversible probability distribution model can fuse the third probability distribution with the timbre information of the source singer to obtain the second probability distribution corresponding to the sample source song. The first probability distribution and the fundamental frequency of the sample source song are input into the decoding network to obtain the decoded song.

[0165] The first loss function is determined by comparing the decoded song and the sample source song, and the second loss function is determined by comparing the first probability distribution and the second probability distribution.

[0166] The model parameters of the song conversion model are adjusted with the goal of minimizing the sum of the first and second loss functions.

[0167] As an optional implementation, another embodiment of this application also discloses the fundamental frequency information of the source song, including a standard fundamental frequency and a converted fundamental frequency corresponding to the standard fundamental frequency. The standard fundamental frequency is extracted from the source song, and the converted fundamental frequency is obtained by modulating the standard fundamental frequency. The decoding module 130 in the song generation device is specifically used for:

[0168] For each base frequency in the base frequency information of the source song, decode the base frequency and the encoding information of the second song respectively to obtain the decoded song corresponding to each base frequency in the base frequency information of the source song;

[0169] Select the target song from the various decoded songs.

[0170] As an optional implementation, another embodiment of this application also discloses the timbre characteristics of the target singer, including the timbre characteristics of a virtual singer, which are obtained by fusing the timbre characteristics of real singers with different timbre features.

[0171] As an optional implementation, another embodiment of this application also discloses that the song generation apparatus further includes: a timbre feature acquisition module, used for:

[0172] Identify at least one intended vocal characteristic of the virtual singer;

[0173] Acquire the timbre characteristics of real singers that match each intended timbre characteristic;

[0174] According to preset weights, the timbre features of real singers that match each desired timbre feature are weighted and fused to obtain the timbre features of virtual singers.

[0175] As an optional implementation, another embodiment of this application also discloses that the timbre feature acquisition module performs weighted fusion of the timbre features of real singers that match each intended timbre feature according to preset weights to obtain the timbre features of virtual singers, including:

[0176] According to each weight combination, the timbre features of real singers that match each intention timbre feature are weighted and fused to obtain the candidate fused timbre features corresponding to each weight combination.

[0177] Based on each candidate fusion timbre feature, the song to be converted is timbre converted, and based on the song conversion result, the target fusion timbre feature is selected from the candidate fusion timbre features as the timbre feature of the virtual singer.

[0178] Corresponding to the above-described song generation method, this application also discloses a song generation apparatus, see [link to apparatus]. Figure 8 As shown, the device includes:

[0179] The feature input module 200 is used to input the acoustic features of the source song, the timbre features of the target singer, and the fundamental frequency information of the source song into a pre-trained song conversion model so that the song conversion model can remove the timbre information of the source singer from the source song, obtain the source song encoding features that do not contain the timbre information of the source singer, and use the source song encoding features and the timbre features of the target singer to generate a target song that contains the timbre of the target singer.

[0180] The training process of the song conversion model aims to minimize the sum of the first loss function and the second loss function. The first loss function is determined based on the difference between the result of the song conversion model recovering the sample source song using the first encoding information and the sample source song. The second loss function is determined based on the difference between the first encoding information and the second encoding information.

[0181] The first encoding information is obtained by encoding the acoustic features of the sample source song by the song conversion model. The second encoding information is obtained by fusing the third encoding information with the timbre information of the source singer corresponding to the sample source song by the song conversion model. The third encoding information is obtained by encoding the acoustic features of the sample source song that are unrelated to the speaker's timbre by a preset prior encoding network.

[0182] The song generation apparatus provided in the above embodiments belongs to the same concept as the song generation method provided in the above embodiments of this application. It can execute the song generation method provided in any of the above embodiments of this application and has the corresponding functional modules and beneficial effects for executing the song generation method. Technical details not described in detail in this embodiment can be found in the specific processing content of the song generation method provided in the above embodiments of this application, and will not be repeated here.

[0183] Exemplary electronic devices, storage media, and computing products

[0184] Corresponding to the above-described song production method, this application also discloses an electronic device, see [link to relevant documentation]. Figure 9 As shown, the electronic device includes:

[0185] Memory 300 and processor 310;

[0186] The memory 300 is connected to the processor 310 and is used to store programs;

[0187] The processor 310 is configured to implement the song generation method disclosed in any of the above embodiments by running a program stored in the memory 300.

[0188] Specifically, the aforementioned electronic device may further include: a bus, a communication interface 320, an input device 330, and an output device 340.

[0189] The processor 310, memory 300, communication interface 320, input device 330, and output device 340 are interconnected via a bus. Among them:

[0190] A bus can include a pathway for transmitting information between various components of a computer system.

[0191] The processor 310 can be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits used to control the execution of the program of the present application. It can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0192] Processor 310 may include a main processor, as well as a baseband chip, modem, etc.

[0193] The memory 300 stores a program for executing the technical solution of this application, and may also store an operating system and other critical business functions. Specifically, the program may include program code, which includes computer operation instructions. More specifically, the memory 300 may include read-only memory (ROM), other types of static storage devices capable of storing static information and instructions, random access memory (RAM), other types of dynamic storage devices capable of storing information and instructions, disk storage, flash memory, etc.

[0194] Input device 330 may include a device for receiving user input data and information, such as a keyboard, mouse, camera, scanner, light pen, voice input device, touch screen, pedometer, or gravity sensor.

[0195] Output device 340 may include devices that allow information to be output to a user, such as a display screen, printer, speaker, etc.

[0196] The communication interface 320 may include a device that uses any transceiver to communicate with other devices or communication networks, such as Ethernet, Radio Access Network (RAN), Wireless Local Area Network (WLAN), etc.

[0197] The processor 310 executes the program stored in the memory 300 and calls other devices, which can be used to implement the various steps of the song generation method provided in the above embodiments of this application.

[0198] Another embodiment of this application provides a storage medium storing a computer program, which, when executed by a processor, implements the various steps of the song generation method provided in any of the above embodiments.

[0199] For the foregoing method embodiments, in order to simplify the description, they are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to this application, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0200] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For apparatus embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

[0201] The steps in the methods of the various embodiments of this application can be adjusted, merged, or deleted in order according to actual needs, and the technical features described in each embodiment can be replaced or combined.

[0202] The modules and sub-modules in the various embodiments of the present application's devices and terminals can be merged, divided, and deleted according to actual needs.

[0203] It should be understood that the disclosed terminals, devices, and methods can be implemented in other ways, given the several embodiments provided in this application. For example, the terminal embodiments described above are merely illustrative. For instance, the division of modules or sub-modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple sub-modules or modules may be combined or integrated into another module, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or modules, and may be electrical, mechanical, or other forms.

[0204] The modules or submodules described as separate components may or may not be physically separate. The components that constitute a module or submodule may or may not be physical modules or submodules; that is, they may be located in one place or distributed across multiple network modules or submodules. Some or all of the modules or submodules can be selected to achieve the purpose of this embodiment's solution, depending on actual needs.

[0205] Furthermore, the functional modules or sub-modules in the various embodiments of this application can be integrated into one processing module, or each module or sub-module can exist physically separately, or two or more modules or sub-modules can be integrated into one module. The integrated modules or sub-modules described above can be implemented in hardware or in the form of software functional modules or sub-modules.

[0206] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0207] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software unit executed by a processor, or a combination of both. The software unit can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0208] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0209] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A song generation method, characterized in that, include: The acoustic features of the source song are encoded using a pre-trained song conversion model to obtain source song encoding information. Based on the timbre features of the source singer, the timbre information of the source singer is removed from the source song encoding information to obtain first song encoding information. The first song encoding information is then fused with the timbre features of the target singer to obtain second song encoding information. The target song is generated by decoding the fundamental frequency information of the source song and the second song encoding information. The source song encoding information includes the pronunciation content information of the source song, the timbre information of the source singer, and the singing details information. The singing details information includes at least one of emotional information, rhythmic information, and intonation information. The timbre features of the target singer include the timbre features of a virtual singer, which are obtained by fusing the timbre features of real singers with different timbre characteristics. The training process of the song conversion model aims to minimize the sum of the first loss function and the second loss function. The first loss function is determined based on the difference between the result of the song conversion model recovering the sample source song using the first encoding information and the sample source song. The second loss function is determined based on the difference between the first encoding information and the second encoding information. The first encoding information is obtained by encoding the acoustic features of the sample source song by the song conversion model. The second encoding information is obtained by fusing the third encoding information with the timbre information of the source singer corresponding to the sample source song by the song conversion model. The third encoding information is obtained by encoding the features of the sample source song that are unrelated to the speaker's timbre by a preset prior encoding network.

2. The method according to claim 1, characterized in that, The song conversion model includes: The coding network is used to encode the acoustic features of the source song to obtain the source song coding information; A reversible probability distribution model is used to remove the source singer's timbre information from the source song encoding information based on the source singer's timbre characteristics to obtain first song encoding information, and to fuse the first song encoding information with the target singer's timbre characteristics to obtain second song encoding information; A decoding network is used to generate a target song by decoding the baseband information of the source song and the encoding information of the second song.

3. The method according to claim 2, characterized in that, The song conversion model also includes: Timbre coding networks are used to encode the acoustic features of a singer's speech to obtain the singer's timbre characteristics.

4. The method according to claim 1, characterized in that, The song conversion model includes an encoding network, a reversible probability distribution model, a decoding network, and a timbre encoding network; The training process for the song conversion model includes: The acoustic features of the sample source song are encoded through the encoding network to obtain the first probability distribution corresponding to the sample source song, and the timbre features of the source singer are obtained by encoding the acoustic features of the sample source song through the timbre encoding network. The features of the sample source song that are independent of the speaker's timbre are input into a preset prior coding network to obtain the third probability distribution of the sample source song. The third probability distribution and the timbre features of the source singer are input into the reversible probability distribution model, so that the reversible probability distribution model fuses the third probability distribution with the timbre information of the source singer to obtain the second probability distribution corresponding to the sample source song. The first probability distribution and the fundamental frequency of the sample source song are input into the decoding network to obtain the decoded song. A first loss function is determined by comparing the decoded song with the sample source song, and a second loss function is determined by comparing the first probability distribution with the second probability distribution. The model parameters of the song conversion model are adjusted with the goal of minimizing the sum of the first loss function and the second loss function.

5. The method according to any one of claims 1 to 4, characterized in that, The fundamental frequency information of the source song includes a standard fundamental frequency and a converted fundamental frequency corresponding to the standard fundamental frequency, wherein the standard fundamental frequency is the fundamental frequency extracted from the source song, and the converted fundamental frequency is obtained by performing a modulation domain conversion on the standard fundamental frequency; The step of generating a target song by decoding the fundamental frequency information of the source song and the encoding information of the second song includes: For each base frequency in the base frequency information of the source song, the base frequency is decoded with the second song encoding information to obtain a decoded song corresponding to each base frequency in the base frequency information of the source song. Select the target song from the various decoded songs.

6. The method according to claim 1, characterized in that, The process of acquiring the vocal characteristics of the virtual singer includes: Identify at least one intended vocal characteristic of the virtual singer; Acquire the timbre characteristics of real singers that match each intended timbre characteristic; According to preset weights, the timbre features of real singers that match each desired timbre feature are weighted and fused to obtain the timbre features of virtual singers.

7. The method according to claim 6, characterized in that, The preset weights include multiple different weight combinations; The virtual singer's timbre features are obtained by weighted fusion of the timbre features of real singers that match each intended timbre feature according to preset weights, including: According to each weight combination, the timbre features of real singers that match each intention timbre feature are weighted and fused to obtain the candidate fused timbre features corresponding to each weight combination. Based on each candidate fusion timbre feature, the song to be converted is timbre converted, and based on the song conversion result, the target fusion timbre feature is selected from the candidate fusion timbre features as the timbre feature of the virtual singer.

8. A song generation method, characterized in that, include: The acoustic features of the source song, the timbre features of the target singer, and the fundamental frequency information of the source song are input into a pre-trained song conversion model. The song conversion model removes the timbre information of the source singer from the source song to obtain source song encoding features that do not contain the timbre information of the source singer. Then, using the source song encoding features and the timbre features of the target singer, a target song containing the timbre of the target singer is generated. The timbre features of the target singer include the timbre features of a virtual singer, which are obtained by fusing the timbre features of real singers with different timbre characteristics. The training process of the song conversion model aims to minimize the sum of the first loss function and the second loss function. The first loss function is determined based on the difference between the result of the song conversion model recovering the sample source song using the first encoding information and the sample source song. The second loss function is determined based on the difference between the first encoding information and the second encoding information. The first encoding information is obtained by encoding the acoustic features of the sample source song by the song conversion model. The second encoding information is obtained by fusing the third encoding information with the timbre information of the source singer corresponding to the sample source song by the song conversion model. The third encoding information is obtained by encoding the acoustic features of the sample source song that are unrelated to the speaker's timbre by a preset prior encoding network.

9. A song generation device, characterized in that, include: The model application module is used to encode the acoustic features of a source song using a pre-trained song conversion model to obtain source song encoding information. Based on the timbre features of the source singer, the timbre information of the source singer is removed from the source song encoding information to obtain first song encoding information. The first song encoding information is then fused with the timbre features of the target singer to obtain second song encoding information. Finally, the target song is generated by decoding the fundamental frequency information of the source song and the second song encoding information. The source song encoding information includes the pronunciation content information of the source song, the timbre information of the source singer, and singing detail information. The singing detail information includes at least one of emotional information, rhythmic information, and intonation information. The timbre features of the target singer include the timbre features of a virtual singer, which are obtained by fusing the timbre features of real singers with different timbre characteristics. The training process of the song conversion model aims to minimize the sum of the first loss function and the second loss function. The first loss function is determined based on the difference between the result of the song conversion model recovering the sample source song using the first encoding information and the sample source song. The second loss function is determined based on the difference between the first encoding information and the second encoding information. The first encoding information is obtained by encoding the acoustic features of the sample source song by the song conversion model. The second encoding information is obtained by fusing the third encoding information with the timbre information of the source singer corresponding to the sample source song by the song conversion model. The third encoding information is obtained by encoding the features of the sample source song that are unrelated to the speaker's timbre by a preset prior encoding network.

10. A song generation device, characterized in that, include: The feature input module is used to input the acoustic features of the source song, the timbre features of the target singer, and the fundamental frequency information of the source song into a pre-trained song conversion model. This allows the song conversion model to remove the source singer's timbre information from the source song, obtaining source song encoding features that do not contain the source singer's timbre information. The model then uses these source song encoding features and the target singer's timbre features to generate a target song that includes the target singer's timbre. The target singer's timbre features include those of a virtual singer, which are obtained by fusing the timbre features of real singers with different timbre characteristics. The training process of the song conversion model aims to minimize the sum of the first loss function and the second loss function. The first loss function is determined based on the difference between the result of the song conversion model recovering the sample source song using the first encoding information and the sample source song. The second loss function is determined based on the difference between the first encoding information and the second encoding information. The first encoding information is obtained by encoding the acoustic features of the sample source song by the song conversion model. The second encoding information is obtained by fusing the third encoding information with the timbre information of the source singer corresponding to the sample source song by the song conversion model. The third encoding information is obtained by encoding the acoustic features of the sample source song that are unrelated to the speaker's timbre by a preset prior encoding network.

11. An electronic device, characterized in that, include: Memory and processor; The memory is connected to the processor and is used to store programs; The processor is configured to implement the song generation method as described in any one of claims 1 to 7 or the song generation method as described in claim 8 by running a program in the memory.

12. A storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the song generation method as described in any one of claims 1 to 7 or the song generation method as described in claim 8.

Citation Information

Patent Citations

  • End-to-end voice conversion method and system, terminal and storage medium

    CN114974274A