Song conversion method and device, electronic equipment and storage medium
By constructing emotion triplets and timbre triplets to decouple the singer's timbre feature vector and emotion feature vector, the problem of the difficulty in decoupling the correlation between timbre and emotion in singing conversion is solved, generating songs with more realistic target singer style and emotional details.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- PING AN TECH (SHENZHEN) CO LTD
- Filing Date
- 2024-01-25
- Publication Date
- 2026-04-28
AI Technical Summary
In existing vocal conversion technologies, the correlation between timbre feature vectors and emotional feature vectors is difficult to decouple and control independently, resulting in insufficient emotional expression in the converted songs and an inability to fully capture the subtle emotional feature vectors of the target singer.
By constructing emotion triplets and timbre triplets, the singer's timbre feature vector and emotion feature vector are decoupled respectively. The waveform is reconstructed by splicing the feature vector and the linear spectrum feature vector, and a synthetic waveform is generated to realize the singing voice conversion.
It achieves effective decoupling and independent control of the singer's emotions and timbre, and the generated songs more accurately simulate the target singer's timbre style and emotional details, enhancing the realism of the vocal transitions and the emotional expressiveness.
Smart Images

Figure CN117912445B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of voice data processing, and more particularly to a method, apparatus, electronic device, and storage medium for converting singing voices. Background Technology
[0002] Vocal conversion technology, as an advanced sound processing technique, has the core function of accurately simulating and transferring the unique vocal style of the source singer to the timbre of the target singer while completely preserving the original content of the source song.
[0003] Despite significant progress in vocal conversion technology, the current method for converting vocal features based on the timbre and emotional feature vectors of song samples from the same singer is to encode and decode them. Since the timbre and emotional feature vectors are related in some dimensions, the direct encoding and decoding method makes it difficult to effectively decouple and independently control the correlation between the timbre and emotional feature vectors in the song.
[0004] Therefore, in terms of actual conversion results, the emotional expression of the converted song is often slightly insufficient, which invisibly limits the degree of authentic reproduction of the target singer's singing style and emotional details.
[0005] For example, financial and insurance companies are incorporating singing voice conversion apps into their mobile applications to provide diverse value-added or interactive entertainment services, thereby retaining existing users and attracting more new users.
[0006] For example, after completing an insurance transaction, users can upload their own recorded song snippets or select song samples provided by the platform to have their voices converted into the vocal styles of different well-known singers using a voice conversion model.
[0007] Due to the limitations of current technology, such as the difficulty in effectively decoupling and independently controlling the correlation between timbre feature vectors and emotional feature vectors in a song, the converted vocals may be slightly lacking in some complex emotional expressions and may not be able to fully capture all the subtle emotional feature vectors of the target singer. Summary of the Invention
[0008] In view of the above, it is necessary to provide a singing voice conversion method, the purpose of which is to effectively decouple the timbre feature vectors and emotional feature vectors of the source singer and the target singer in the song sample, so as to ensure that the converted song rendering retains the timbre characteristics and emotional details of the target singer.
[0009] The singing voice conversion method provided by this invention includes:
[0010] Obtain a first song sample from a first singer and a second song sample from a second singer. Extract the first timbre feature vector and the first emotion feature vector from the first song sample. Extract the second timbre feature vector and the second emotion feature vector from the second song sample.
[0011] An emotion triplet is constructed using the first emotion feature vector, the second emotion feature vector, and the emotion label feature vector of the emotion category to which the first song sample belongs; and a timbre triplet is constructed using the first timbre feature vector, the second timbre feature vector, and the timbre label feature vector of the timbre category to which the first song sample belongs.
[0012] Decouple the feature vectors of the emotion triplet and the timbre triplet, and then concatenate the decoupled feature vectors to obtain the concatenated feature vector.
[0013] Extract the linear spectrum feature vector of the first song sample, and reconstruct the waveform using the spliced feature vector and the linear spectrum feature vector to obtain the synthesized waveform;
[0014] Calculate the difference in probability distribution between the synthesized waveform and the real waveform of the first song sample. When the difference is less than a threshold, convert the first song sample from the voice of the first singer to the voice of the second singer to obtain the converted song.
[0015] Optionally, the step of extracting the first timbre feature vector and the first emotion feature vector of the first singer from the first song sample includes:
[0016] Calculate the first Mel frequency cepstral coefficient of the first song sample, and extract the first timbre feature vector of the first singer from the first Mel frequency cepstral coefficient;
[0017] The first emotional feature vector of the first singer is extracted from the audio signal of the first song sample, wherein the audio signal includes at least one of the following: spectral envelope signal, formant shift signal, and transient analysis signal.
[0018] Optionally, the step of extracting the second singer's second timbre feature vector and second emotional feature vector from the second song sample includes:
[0019] Calculate the second Mel frequency cepstral coefficients of the second song sample, and extract the second timbre feature vector of the second singer from the second Mel frequency cepstral coefficients;
[0020] The second emotional feature vector of the second singer is extracted from the audio signal of the second song sample, wherein the audio signal includes at least one of the following: spectral envelope signal, formant shift signal, and transient analysis signal.
[0021] Optionally, the sentiment triple includes a negative example feature vector, a positive example feature vector, and an anchor feature vector. The step of constructing the sentiment triple using the first sentiment feature vector, the second sentiment feature vector, and the sentiment label feature vector of the sentiment category to which the first song sample belongs includes:
[0022] The first emotion feature vector is used as the negative example feature vector of the emotion triple, the second emotion feature vector is used as the positive example feature vector of the emotion triple, and the emotion tag feature vector is used as the anchor feature vector of the emotion triple.
[0023] Optionally, the timbre triplet includes a negative example feature vector, a positive example feature vector, and an anchor point feature vector. The step of constructing the timbre triplet using the first timbre feature vector, the second timbre feature vector, and the timbre label feature vector of the timbre category to which the first song sample belongs includes:
[0024] The first timbre feature vector is used as the negative example feature vector of the timbre triplet, the second timbre feature vector is used as the positive example feature vector of the timbre triplet, and the timbre tag feature vector is used as the anchor point feature vector of the timbre triplet.
[0025] Optionally, the step of decoupling the feature vectors of the emotion triplet and the timbre triplet, and concatenating the decoupled feature vectors to obtain the concatenated feature vector, includes:
[0026] Using the anchor feature vector of the sentiment triple as the cluster center, the distance between the positive example feature vector and the anchor feature vector is reduced, and the distance between the negative example feature vector and the anchor feature vector is increased. The sentiment triple is iterated until the loss function value of the sentiment triple is less than a preset threshold, at which point the iteration stops, thus completing the decoupling of the latent distribution feature vectors between the feature vectors of the sentiment triple.
[0027] Using the anchor feature vector of the timbre triplet as the cluster center, the distance between the positive example feature vector and the anchor feature vector is reduced, and the distance between the negative example feature vector and the anchor feature vector is increased. The timbre triplet is iterated until the loss function value of the timbre triplet is less than a preset threshold, at which point the iteration stops, thus completing the decoupling of the latent distribution feature vectors between the feature vectors of the timbre triplet.
[0028] The concatenated feature vector is obtained by splicing the decoupled feature vectors together.
[0029] Optionally, the step of reconstructing the waveform from the spliced feature vector and the linear spectral feature vector to obtain the synthesized waveform includes:
[0030] The spliced feature vector is restored to a first time-domain signal, and the linear spectrum feature vector is restored to a second time-domain signal;
[0031] The first time-domain signal and the second time-domain signal are spliced together to obtain the synthesized waveform.
[0032] To address the above problems, the present invention also provides a singing voice conversion device, the device comprising:
[0033] The acquisition module is used to acquire a first song sample of a first singer and a second song sample of a second singer, extract a first timbre feature vector and a first emotion feature vector of the first singer from the first song sample, and extract a second timbre feature vector and a second emotion feature vector of the second singer from the second song sample;
[0034] The construction module is used to construct an emotion triplet using the first emotion feature vector, the second emotion feature vector, and the emotion label feature vector of the emotion category to which the first song sample belongs; and to construct a timbre triplet using the first timbre feature vector, the second timbre feature vector, and the timbre label feature vector of the timbre category to which the first song sample belongs.
[0035] The decoupling module is used to decouple each feature vector of the emotion triplet and the timbre triplet, and then splice the decoupled feature vectors to obtain the spliced feature vector.
[0036] The reconstruction module is used to extract the linear spectrum feature vector of the first song sample, and reconstruct the waveform from the spliced feature vector and the linear spectrum feature vector to obtain the synthesized waveform;
[0037] The conversion module is used to calculate the difference value of the probability distribution between the synthesized waveform and the real waveform of the first song sample. When the difference value is less than a threshold, the first song sample is converted from the singing voice of the first singer to the singing voice of the second singer to obtain the converted song.
[0038] To address the above problems, the present invention also provides an electronic device, the electronic device comprising:
[0039] At least one processor; and,
[0040] A memory communicatively connected to the at least one processor; wherein,
[0041] The memory stores a singing conversion program that can be executed by the at least one processor, the singing conversion program being executed by the at least one processor to enable the at least one processor to perform the singing conversion method described above.
[0042] To address the aforementioned problems, the present invention also provides a computer-readable storage medium storing a singing voice conversion program, which can be executed by one or more processors to implement the singing voice conversion method described above.
[0043] Compared with existing technologies, this invention extracts the first timbre feature vector and the first emotion feature vector of the first singer from the first song sample, and extracts the second timbre feature vector and the second emotion feature vector of the second singer from the second song sample. Using the extracted timbre feature vector and emotion feature vector, emotion triplet and timbre triplet are constructed respectively. By performing decoupling operations on each feature vector of the emotion triplet and timbre triplet, the effective decoupling and independent control of the emotion dimension and timbre dimension of the first and second singers can be achieved.
[0044] The waveform is reconstructed by decoupling each feature vector and the linear spectrum feature vector of the first song sample, and a synthetic waveform is generated. When the difference in probability distribution between the synthetic waveform and the real waveform of the first song sample is less than the threshold, it means that the singing voice conversion effect has reached the ideal state, and the singing voice of the first singer has been successfully converted into the voice style of the second singer, generating the converted song.
[0045] This invention can more accurately simulate and transfer the content of a source song to the vocal style of a target singer, while preserving and reproducing the rich emotional details of the target singer, effectively enhancing the expressiveness and realism of the vocal conversion technology in terms of emotional expression. In application scenarios such as mobile applications for financial and insurance companies, it can provide a more realistic, vivid, and emotionally rich vocal conversion experience, further improving user engagement and satisfaction. Attached Figure Description
[0046] Figure 1 This is a schematic flowchart of a singing voice conversion method provided in an embodiment of the present invention;
[0047] Figure 2 This is a schematic diagram of a singing voice conversion device provided in an embodiment of the present invention;
[0048] Figure 3 This is a schematic diagram of the structure of an electronic device for implementing a singing voice conversion method according to an embodiment of the present invention;
[0049] The objectives, features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0050] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without inventive effort are within the scope of protection of this invention.
[0051] It should be noted that the descriptions involving "first," "second," etc., in this invention are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of indicated technical feature vectors. Therefore, feature vectors defined with "first" and "second" may explicitly or implicitly include at least one of those feature vectors. Furthermore, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. If the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by this invention.
[0052] Reference Figure 1 The diagram shown is a flowchart illustrating a singing voice conversion method provided in an embodiment of the present invention.
[0053] This method is executed by an electronic device.
[0054] In this embodiment, the singing voice conversion method includes:
[0055] S1. Obtain the first song sample of the first singer and the second song sample of the second singer. Extract the first timbre feature vector and the first emotion feature vector of the first singer from the first song sample. Extract the second timbre feature vector and the second emotion feature vector of the second singer from the second song sample.
[0056] In this embodiment, in order to convert the vocal style of the source singer into the timbre and emotion of the target singer, the first singer is the source singer and the second singer is the target singer.
[0057] The first song sample of the first singer and the second song sample of the second singer are input into the singing conversion model. The digital audio files (such as .wav, .mp3, etc.) corresponding to the first and second song samples are read. The singing conversion model includes, but is not limited to, the VITS (Variational Inference with adversarial learning for end-to-end Text-to-Speech) model. The first song sample and the second song sample can be different songs or the same song. There is no limitation here.
[0058] The audio signals of the song samples are digitally sampled at a preset sampling rate (e.g., 44.1kHz or 48kHz) to obtain continuous audio signals. The encoder of the singing conversion model extracts the first timbre feature vector and the first emotional feature vector of the first singer from the audio signal of the first song sample, and extracts the second timbre feature vector and the second emotional feature vector of the second singer from the audio signal of the second song sample.
[0059] A timbre feature vector refers to a singer's specific quality or characteristic. It can distinguish different singers, even those singing the same lyrics. Timbre feature vectors can be obtained from the following sources: 1. Mel-frequency cepstral coefficients (MFCCs): A widely used feature vector parameter in speech processing, describing the basic structure of timbre by analyzing the frequency domain characteristics of audio signals. 2. Linear predictive coding (LPC) parameters: Used to represent the resonant modes of the speech waveform, reflecting the influence of the singer's vocal cords, oral cavity, nasal cavity, and other anatomical structures on the sound.
[0060] Songs are categorized into corresponding timbre types based on their timbre feature vectors, such as male / female soprano, male / female alto, male / female bass, lyric soprano, dramatic soprano, coloratura soprano, and other different timbre categories.
[0061] An emotional feature vector (EVV) refers to a singer's emotional state while performing a song, such as joy, sadness, excitement, or calmness. EVV is a set of feature vector values extracted from the audio signal of a song. Based on these EVVs, songs are categorized into corresponding emotional categories, such as joy, sadness, excitement, and lyrical.
[0062] This invention is illustrated by example H: For instance, a mobile application developed by a financial insurance company includes a built-in vocal conversion module (voice conversion model). After purchasing car insurance, user B becomes interested in the application's value-added services. He chooses to try converting his own rendition of "Warm Home" into the vocal style of a well-known singer, B.
[0063] User A recorded and uploaded his version of "Warm Home" through the application as the first song sample. The application has pre-stored several songs by well-known singer B as the second song sample set. The timbre feature vector and emotional feature vector of well-known singer B are extracted from them, and the timbre feature vector and emotional feature vector of user A are extracted from the songs uploaded by user A.
[0064] In one embodiment, extracting the first timbre feature vector and the first emotion feature vector of the first singer from the first song sample includes:
[0065] Calculate the first Mel frequency cepstral coefficient of the first song sample, and extract the first timbre feature vector of the first singer from the first Mel frequency cepstral coefficient;
[0066] The first emotional feature vector of the first singer is extracted from the audio signal of the first song sample, wherein the audio signal includes at least one of the following: spectral envelope signal, formant shift signal, and transient analysis signal.
[0067] The audio data of the first song sample is preprocessed, including framing and windowing. Framing involves dividing the continuous audio signal in the first speech sample into several short segments (one frame). Windowing involves applying a windowing function (e.g., Hamming window, Hanning window, etc.) to each frame to reduce frame edge artifacts.
[0068] Perform a Discrete Fourier Transform on the windowed signal of each frame to convert it from a time-domain signal to a frequency-domain signal, and obtain the spectrum of the signal for that frame. Extract the energy distribution of each frame signal at each Mel frequency in the spectrum using a filter bank.
[0069] The energy distribution at each Mel frequency is further compressed using Discrete Cosine Transform (DCT) to obtain the MFCC coefficient sequence. Since the low-frequency coefficients of the MFCC sequence contain more timbre information, while the high-frequency coefficients are relatively less important, the low-frequency coefficients are selected as the first timbre feature vector for the first singer.
[0070] For example, an MFCC coefficient sequence has 15 coefficients. Usually, the high-frequency coefficients are at the beginning and end of the MFCC coefficient sequence. The 2nd to 13th coefficients (low-frequency coefficients) are selected as the main timbre feature vectors to obtain the first timbre feature vector of the first singer.
[0071] The spectral envelope signal is extracted from the first Mel frequency cepstral coefficients. Various parameters of the spectral envelope signal are analyzed to obtain the first emotional feature vector. These parameters include peak value, slope, kurtosis, and fluctuation amplitude. The spectral envelope contains the energy distribution of the first song sample, and the energy distribution is closely related to emotional expression; for example, high-spirited emotions usually correspond to a higher energy concentration region.
[0072] Formant analysis of the signal detects the formant frequencies generated by the vocal tract during a singer's vocalization and their changes over time, yielding a first emotional feature vector. The speed, amplitude, and frequency of formant migration can reflect subtle feature vectors such as lip movements and breath control during singing, all of which are related to emotional expression. For example, anger or excitement may be accompanied by a faster formant migration speed.
[0073] Transient analysis is performed on the audio data to identify and separate transient components containing impact and suddenness, yielding the first emotional feature vector. Transient components include consonant bursts and the onset and ending phases of vowels. Feature vectors such as the intensity, duration, and shape of the transient signal reflect the singer's vocal power and rhythm, elements often closely linked to the song's emotion. Emotional feature vectors are constructed using transient analysis signals.
[0074] In one embodiment, extracting the second singer's second timbre feature vector and second emotional feature vector from the second song sample includes:
[0075] Calculate the second Mel frequency cepstral coefficients of the second song sample, and extract the second timbre feature vector of the second singer from the second Mel frequency cepstral coefficients;
[0076] The second emotional feature vector of the second singer is extracted from the audio signal of the second song sample, wherein the audio signal includes at least one of the following: spectral envelope signal, formant shift signal, and transient analysis signal.
[0077] Extract the second timbre feature vector and the second emotional feature vector of the second singer. Refer to the steps described above for extracting the first timbre feature vector and the first emotional feature vector of the first singer. The extraction steps are the same for both, so they will not be repeated here.
[0078] S2. Construct an emotion triplet using the first emotion feature vector, the second emotion feature vector, and the emotion label feature vector of the emotion category to which the first song sample belongs; and construct a timbre triplet using the first timbre feature vector, the second timbre feature vector, and the timbre label feature vector of the timbre category to which the first song sample belongs.
[0079] In this embodiment, in order to decouple and independently control the emotional feature vectors of the source singer and the target singer, so as to more accurately simulate and transfer emotional expression, an emotional triplet is constructed using the first emotional feature vector, the second emotional feature vector, and the emotional label feature vector of the emotional category to which the first song sample belongs.
[0080] In order to decouple and independently control the timbre feature vectors of the source singer and the target singer, so as to more accurately simulate and transfer timbre, a timbre triplet is constructed using the first timbre feature vector, the second timbre feature vector, and the timbre label feature vector of the timbre category to which the first song sample belongs.
[0081] Compared to existing technologies that encode and decode based on the timbre and emotion feature vectors of song samples from the same singer, this invention can effectively decouple the timbre and emotion feature vectors from multiple different singers, thereby achieving better sound conversion of emotion and timbre feature vectors, improving the subtlety of emotional expression and the matching degree of timbre feature vectors, and thus creating a more realistic synthesized voice that matches the characteristics of the target singer.
[0082] Continuing with example H: construct an emotional triplet using the emotional feature vector of user A, the emotional feature vector of famous singer B, and the emotional tag feature vector of the emotional category to which "Warm Home" belongs; and construct an emotional triplet using the timbre feature vector of user A, the timbre feature vector of famous singer B, and the timbre tag feature vector of the timbre category to which "Warm Home" belongs.
[0083] In one embodiment, the sentiment triplet includes a negative example feature vector, a positive example feature vector, and an anchor feature vector. The step of constructing the sentiment triplet using the first sentiment feature vector, the second sentiment feature vector, and the sentiment label feature vector of the sentiment category to which the first song sample belongs includes:
[0084] The first emotion feature vector is used as the negative example feature vector of the emotion triple, the second emotion feature vector is used as the positive example feature vector of the emotion triple, and the emotion tag feature vector is used as the anchor feature vector of the emotion triple.
[0085] Negative example feature vectors of the emotion triple: These represent emotion feature vectors that we do not want to retain after transformation, such as certain specific emotional elements of the source singer.
[0086] Positive feature vector of the emotional triple: represents the emotional feature vector expected to be reflected after transformation, that is, the emotional expression pattern unique to the target singer.
[0087] Anchor feature vector of the emotional triple: as a reference standard or center point, it is used to guide how to adjust the relationship between positive and negative feature vectors, so as to better approximate the emotional style of the target singer.
[0088] The pre-defined emotion label feature vector is obtained by labeling different song samples with emotion categories in advance, and then encoding the labeled emotion categories into feature vector form. For example, the labeled emotion categories are different emotion categories such as anger, joy, and sadness.
[0089] By constructing and decoupling emotional triples, the timbre and emotional information in the song samples are separated, and the emotional feature vectors are independently transferred from the source singer to the target singer in order to achieve a higher quality singing conversion effect.
[0090] In one embodiment, the timbre triplet includes a negative example feature vector, a positive example feature vector, and an anchor point feature vector. The step of constructing the timbre triplet using the first timbre feature vector, the second timbre feature vector, and the timbre label feature vector of the timbre category to which the first song sample belongs includes:
[0091] The first timbre feature vector is used as the negative example feature vector of the timbre triplet, the second timbre feature vector is used as the positive example feature vector of the timbre triplet, and the timbre tag feature vector is used as the anchor point feature vector of the timbre triplet.
[0092] Negative example feature vector of timbre triplet: represents the timbre feature vector of the source singer. This part of the feature vector needs to be replaced during the conversion process.
[0093] Positive feature vector of timbre triplet: Represents the timbre feature vector of the target singer, which is the target feature vector expected to be transferred from the source singer's vocal samples.
[0094] Anchor feature vector of timbre triplet: As a reference or center point, it helps to shorten the distance between positive feature vector and anchor point, while widening the distance between negative feature vector and anchor point, thereby mapping and transforming the timbre feature vector of source singer to the timbre feature vector of target singer.
[0095] The pre-defined timbre label feature vector is obtained by pre-labeling different song samples with timbre categories, encoding the labeled timbre categories into feature vector form. For example, the labeled timbre categories are different timbre categories such as male / female soprano, male / female alto, male / female bass, lyric soprano, dramatic soprano, and coloratura soprano.
[0096] By constructing timbre triples and applying specific distance adjustment algorithms, such as clustering methods, the originally closely related timbre feature vectors can be decoupled to more accurately simulate and reproduce the timbre characteristics of the target singer without being affected by the original timbre, ultimately achieving a high-quality timbre conversion effect.
[0097] S3. Decouple each feature vector of the emotion triplet and the timbre triplet, and splice the decoupled feature vectors to obtain the spliced feature vector.
[0098] In this embodiment, the feature vectors of the emotion triplet and the timbre triplet are decoupled, and the decoupled feature vectors are concatenated to obtain the concatenated feature vector.
[0099] In one embodiment, the step of decoupling the feature vectors of the emotion triplet and the timbre triplet, and concatenating the decoupled feature vectors to obtain a concatenated feature vector, includes:
[0100] Using the anchor feature vector of the sentiment triple as the cluster center, the distance between the positive example feature vector and the anchor feature vector is reduced, and the distance between the negative example feature vector and the anchor feature vector is increased. The sentiment triple is iterated until the loss function value of the sentiment triple is less than a preset threshold, at which point the iteration stops, thus completing the decoupling of the latent distribution feature vectors between the feature vectors of the sentiment triple.
[0101] Using the anchor feature vector of the timbre triplet as the cluster center, the distance between the positive example feature vector and the anchor feature vector is reduced, and the distance between the negative example feature vector and the anchor feature vector is increased. The timbre triplet is iterated until the loss function value of the timbre triplet is less than a preset threshold, at which point the iteration stops, thus completing the decoupling of the latent distribution feature vectors between the feature vectors of the timbre triplet.
[0102] The concatenated feature vector is obtained by splicing the decoupled feature vectors together.
[0103] The decoupling process of the emotional triad:
[0104] A1. Calculate the distance from the negative example feature vector to the anchor feature vector (e.g., Euclidean distance, cosine similarity, etc.), and calculate the distance from the positive example feature vector to the anchor feature vector.
[0105] A2. By using a preset function algorithm (e.g., triplet loss function, gradient descent, quasi-Newton method, etc.), update the feature vector parameters of the positive example feature vector, so that the distance between the positive example feature vector and the anchor feature vector gradually decreases, making their distribution in the feature vector space closer, thereby simulating and transferring to the emotional feature vector of the target singer.
[0106] A3. At the same time, perform the opposite operation on the negative example feature vector, update its feature vector parameters, increase the distance between it and the anchor feature vector, and reduce the influence of the source singer's emotional feature vector on the conversion result.
[0107] A4. Repeat the optimization process of A2-A3 above, iterating continuously until the preset convergence condition is reached, such as the loss function value being less than the preset threshold, or the distance change being small after several consecutive iterations, in order to decouple the latent distribution feature vectors between the feature vectors of the sentiment triple. The latent distribution feature vectors refer to the abstract, compressed or potential feature vector variables hidden behind the sentiment feature vectors.
[0108] Latent feature vectors contain high-order information that is not easily measured or perceived directly, but is crucial for vocal style and emotional expression. By decoupling the latent feature vectors, information in the timbre and emotion dimensions can be separated and processed independently, enabling better preservation and transfer of the target emotional expression during vocal conversion, while maintaining consistency in emotional style.
[0109] By analyzing the changes in the relationship between feature vectors before and after decoupling, it is possible to determine whether emotional information has been effectively separated and independently controlled.
[0110] The decoupling process of the timbre triplet:
[0111] B1. Calculate the distance from the negative example feature vector to the anchor feature vector (e.g., Euclidean distance, cosine similarity, etc.), and calculate the distance from the positive example feature vector to the anchor feature vector.
[0112] B2. By using a preset function algorithm (e.g., triplet loss function, gradient descent method, quasi-Newton method, etc.), update the feature vector parameters of the positive example feature vector, so that the distance between the positive example feature vector and the anchor feature vector gradually decreases, making their distribution in the feature vector space closer, thereby simulating and transferring to the timbre feature vector of the target singer.
[0113] B3. At the same time, perform the opposite operation on the negative example feature vector, update its feature vector parameters, increase the distance between it and the anchor feature vector, and reduce the influence of the source singer's timbre feature vector on the conversion result.
[0114] B4. Repeat the optimization process of B2-B3 above, iterating continuously until the preset convergence condition is reached, such as the loss function value being less than the preset threshold, or the distance change being small after several consecutive iterations, in order to decouple the latent distribution feature vectors between the feature vectors of the timbre triplet. The latent distribution feature vectors refer to the abstract, compressed or potential feature vector variables hidden behind the timbre feature vectors.
[0115] Latent feature vectors contain high-order information that is not easily measured or perceived directly, but is crucial for vocal style and timbre expression. By decoupling the latent feature vectors, the information of timbre and timbre dimensions can be separated and processed independently, enabling better preservation and transfer of the target timbre expression during vocal conversion, while maintaining consistency in timbre style.
[0116] By analyzing the changes in the relationship between feature vectors before and after decoupling, it is possible to determine whether the timbre information has been effectively separated and independently controlled.
[0117] After decoupling the emotion triplet and the timbre triplet, the decoupling of each emotion feature vector within the emotion triplet yields a new emotion feature vector representing the target emotion style, while the decoupling of each timbre feature vector within the timbre triplet yields a new timbre feature vector representing the target timbre style.
[0118] By splicing the new emotional feature vector and the new timbre feature vector in a certain order or dimension, a spliced feature vector that transforms both emotion and timbre into the style of the target singer is completed.
[0119] S4. Extract the linear spectrum feature vector of the first song sample, and reconstruct the waveform using the spliced feature vector and the linear spectrum feature vector to obtain the synthesized waveform;
[0120] In this embodiment, the spliced feature vector and the linear spectrum feature vector are used as input data for the first decoder in the singing conversion model. The spliced feature vector and the linear spectrum feature vector are used to reconstruct the waveform to obtain the synthesized waveform.
[0121] In one embodiment, extracting the linear spectral feature vector of the first song sample includes:
[0122] Calculate the similarity between the linear spectrum of the first song sample and all codewords in the preset codebook;
[0123] The linear spectrum is obtained by selecting codewords whose distance is less than a threshold from the similarity results and replacing them.
[0124] The linear spectrum of the first song sample is vectorized to reduce the data dimensionality of the linear spectrum while preserving key feature vectors, resulting in a linear spectrum feature vector, including:
[0125] 1. Treat the initial feature vector of the linear spectrum of each frame of the first song sample as a point in a high-dimensional space.
[0126] 2. Set a preset codebook, which contains a series of pre-trained representative feature vectors (codewords). In speech signal processing or audio coding, a codebook is a pre-trained set of representative timbre feature vectors or spectral feature vectors.
[0127] For example, in the process of vector quantization (VQ), the initial eigenvectors of the original high-dimensional linear spectrum are mapped to one or more of the closest codewords in the codebook, thereby compressing the data and preserving key information.
[0128] 3. Calculate the similarity (e.g., Euclidean distance or cosine distance) between the linear spectrum of the current frame and all codewords in the codebook.
[0129] 4. Select the codewords that are most similar to the initial feature vector of the linear spectrum of the current frame and whose distance is less than the threshold as the approximate representation of the frame.
[0130] 5. Replace the current frame's linear spectrum with the selected codewords to achieve vectorization of the linear spectrum and obtain the linear spectrum feature vector.
[0131] One of the innovations of this invention is that the linear spectrum extracted from the linear spectrum of the song is input into a vector quantizer for vector quantization (VQ), which can remove a small amount of the fundamental frequency signal of the first singer contained in the linear spectrum.
[0132] Meanwhile, the vectorized linear spectral feature vectors can better capture and preserve the main characteristics of the first song sample, such as melody and rhythm structure.
[0133] In one embodiment, before calculating the similarity between the linear spectrum of the first song sample and all codewords in a preset codebook, the method further includes:
[0134] The time-domain signal of each frame of the first song sample is converted into the frequency domain to obtain the time-spectrum.
[0135] The frequency domain feature vector of the time-spectrum is extracted to obtain the linear spectrum.
[0136] A time-spectrum graph is dynamic, showing the spectral changes of a fundamental frequency signal over different time periods. Therefore, a time-spectrum graph is a two-dimensional or three-dimensional image, where the horizontal axis represents time, the vertical axis usually represents frequency, and intensity information such as color or grayscale indicates the signal amplitude or energy density at the corresponding time and frequency point.
[0137] Read the digitized audio file (such as .wav, .mp3, etc.) corresponding to the first song sample, and digitize the audio signal of the first song sample according to the preset sampling rate (such as 44.1kHz or 48kHz) to obtain a continuous audio signal. Perform frame segmentation on the audio signal to obtain the fundamental frequency signal of each frame. Use a preset time-frequency analysis strategy (such as short-time Fourier transform (STFT), cepstral analysis, etc.) to convert the fundamental frequency signal of each frame from the time domain to the frequency domain to obtain the spectrum of each frame. Extract the frequency domain feature vector of the spectrum to obtain the linear spectrum.
[0138] In one embodiment, reconstructing the waveform from the spliced feature vector and the linear spectral feature vector to obtain the synthesized waveform includes:
[0139] The spliced feature vector is restored to a first time-domain signal, and the linear spectrum feature vector is restored to a second time-domain signal;
[0140] The first time-domain signal and the second time-domain signal are spliced together to obtain the synthesized waveform.
[0141] The decoder of the singing conversion model is used to perform decoding and inverse transformation operations on the spliced feature vector (e.g., the inverse transformation operation is the inverse Fourier transform (IFFT)). The spliced feature vector is gradually restored to a waveform that is close to the original singing style of the second singer (target singer) to ensure that the time-domain waveform signal is recovered from the frequency domain feature vector, and the first time-domain signal is obtained.
[0142] The decoder of the song conversion model performs decoding and inverse transformation operations on the linear spectrum feature vector (e.g., the inverse transformation operation is the inverse Fourier transform (IFFT)). The linear spectrum feature vector is gradually restored to a waveform that approximates the melody and rhythm of the first song sample, so as to ensure that the time domain waveform signal is recovered from the frequency domain feature vector, and the second time domain signal is obtained.
[0143] By splicing the first time-domain signal and the second time-domain signal, a composite waveform is obtained.
[0144] The decoder of a singing conversion model refers to a model that has been trained to decode audio signals in the time domain from spliced high-dimensional feature vectors.
[0145] S5. Calculate the difference value of the probability distribution between the synthesized waveform and the real waveform of the first song sample. When the difference value is less than the threshold, convert the first song sample from the singing voice of the first singer to the singing voice of the second singer to obtain the converted song.
[0146] In this embodiment, the synthesized waveform and the real waveform of the first song sample are input into the discriminator of the singing conversion model. The probability distribution of the synthesized waveform is calculated in the parameter space of the discriminator. The probability distribution value of the synthesized waveform is compared with the probability distribution value corresponding to the real waveform to obtain the difference value. If the difference value after comparison is lower than a preset threshold, it is considered that the converted singing voice renders the singing style and emotional details of the second singer (target singer).
[0147] The discriminator is a deep learning network used to distinguish between synthesized and real waveforms. In other words, one of the innovations of this invention is to treat the vocal conversion model as a generator, constructing a generative adversarial network (GAN) with it and the discriminator. Within this GAN framework, the synthesized and real waveforms output by the vocal conversion model are input into the discriminator for discrimination. When the difference value output by the discriminator is less than a threshold, the converted vocals are considered to closely resemble the original singer's singing style and emotional details.
[0148] Continuing with example H: the decoupled feature vectors are concatenated into a concatenated feature vector, and combined with the linear spectrum feature vector of user A's original song to generate a synthesized waveform. The difference in probability distribution between the synthesized waveform and the real waveform of user A's original singing is calculated. If the difference is less than a preset threshold, it indicates that the conversion effect is good. At this time, user A's singing has been successfully converted into a version that is close to the unique timbre and emotional characteristics of the famous singer B.
[0149] Through this emotionally decoupled vocal conversion technology, user A can enjoy novel and engaging interactive entertainment services while receiving financial services, enhancing user A's stickiness and satisfaction with the financial insurance application. Simultaneously, this technology effectively overcomes the limitations of traditional vocal conversion in emotional expression, enabling the converted song to more realistically and delicately showcase the target singer's singing style and emotional details.
[0150] In steps S1-S5, the present invention extracts the first timbre feature vector and the first emotion feature vector of the first singer from the first song sample, and extracts the second timbre feature vector and the second emotion feature vector of the second singer from the second song sample. Using the extracted timbre feature vector and emotion feature vector, emotion triplet and timbre triplet are constructed respectively. By performing decoupling operations on each feature vector of the emotion triplet and timbre triplet, the effective decoupling and independent control of the emotion dimension and timbre dimension of the first and second singers are achieved.
[0151] The waveform is reconstructed by decoupling each feature vector and the linear spectrum feature vector of the first song sample, and a synthetic waveform is generated. When the difference in probability distribution between the synthetic waveform and the real waveform of the first song sample is less than the threshold, it means that the singing voice conversion effect has reached the ideal state, and the singing voice of the first singer has been successfully converted into the voice style of the second singer, generating the converted song.
[0152] This invention can more accurately simulate and transfer the content of a source song to the vocal style of a target singer, while preserving and reproducing the rich emotional details of the target singer, effectively enhancing the expressiveness and realism of the vocal conversion technology in terms of emotional expression. In application scenarios such as mobile applications for financial and insurance companies, it can provide a more realistic, vivid, and emotionally rich vocal conversion experience, further improving user engagement and satisfaction.
[0153] like Figure 2 The diagram shown is a schematic diagram of a singing voice conversion device provided in an embodiment of the present invention.
[0154] The vocal conversion device 100 of this invention can be installed in an electronic device. Depending on the functions implemented, the vocal conversion device 100 may include an acquisition module 110, a construction module 120, a decoupling module 130, a reconstruction module 140, and a conversion module 150. The module described in this invention can also be called a unit, which refers to a series of computer program segments that can be executed by the processor of an electronic device and can perform a fixed function, and which are stored in the memory of the electronic device.
[0155] In this embodiment, the functions of each module / unit are as follows:
[0156] The acquisition module 110 is used to acquire a first song sample of a first singer and a second song sample of a second singer, extract a first timbre feature vector and a first emotion feature vector of the first singer from the first song sample, and extract a second timbre feature vector and a second emotion feature vector of the second singer from the second song sample;
[0157] The construction module 120 is used to construct an emotion triplet using the first emotion feature vector, the second emotion feature vector and the emotion label feature vector of the emotion category to which the first song sample belongs, and to construct a timbre triplet using the first timbre feature vector, the second timbre feature vector and the timbre label feature vector of the timbre category to which the first song sample belongs;
[0158] The decoupling module 130 is used to decouple each feature vector of the emotion triplet and the timbre triplet, and splice the decoupled feature vectors to obtain a spliced feature vector.
[0159] The reconstruction module 140 is used to extract the linear spectrum feature vector of the first song sample, and reconstruct the waveform from the spliced feature vector and the linear spectrum feature vector to obtain the synthesized waveform;
[0160] The conversion module 150 is used to calculate the difference value of the probability distribution between the synthesized waveform and the real waveform of the first song sample. When the difference value is less than a threshold, the first song sample is converted from the singing voice of the first singer to the singing voice of the second singer to obtain the converted song.
[0161] In one embodiment, extracting the first timbre feature vector and the first emotion feature vector of the first singer from the first song sample includes:
[0162] Calculate the first Mel frequency cepstral coefficient of the first song sample, and extract the first timbre feature vector of the first singer from the first Mel frequency cepstral coefficient;
[0163] The first emotional feature vector of the first singer is extracted from the audio signal of the first song sample, wherein the audio signal includes at least one of the following: spectral envelope signal, formant shift signal, and transient analysis signal.
[0164] In one embodiment, extracting the second singer's second timbre feature vector and second emotional feature vector from the second song sample includes:
[0165] Calculate the second Mel frequency cepstral coefficients of the second song sample, and extract the second timbre feature vector of the second singer from the second Mel frequency cepstral coefficients;
[0166] The second emotional feature vector of the second singer is extracted from the audio signal of the second song sample, wherein the audio signal includes at least one of the following: spectral envelope signal, formant shift signal, and transient analysis signal.
[0167] In one embodiment, the sentiment triplet includes a negative example feature vector, a positive example feature vector, and an anchor feature vector. The step of constructing the sentiment triplet using the first sentiment feature vector, the second sentiment feature vector, and the sentiment label feature vector of the sentiment category to which the first song sample belongs includes:
[0168] The first emotion feature vector is used as the negative example feature vector of the emotion triple, the second emotion feature vector is used as the positive example feature vector of the emotion triple, and the emotion tag feature vector is used as the anchor feature vector of the emotion triple.
[0169] In one embodiment, the timbre triplet includes a negative example feature vector, a positive example feature vector, and an anchor point feature vector. The step of constructing the timbre triplet using the first timbre feature vector, the second timbre feature vector, and the timbre label feature vector of the timbre category to which the first song sample belongs includes:
[0170] The first timbre feature vector is used as the negative example feature vector of the timbre triplet, the second timbre feature vector is used as the positive example feature vector of the timbre triplet, and the timbre tag feature vector is used as the anchor point feature vector of the timbre triplet.
[0171] In one embodiment, the step of decoupling the feature vectors of the emotion triplet and the timbre triplet, and concatenating the decoupled feature vectors to obtain a concatenated feature vector, includes:
[0172] Using the anchor feature vector of the sentiment triple as the cluster center, the distance between the positive example feature vector and the anchor feature vector is reduced, and the distance between the negative example feature vector and the anchor feature vector is increased. The sentiment triple is iterated until the loss function value of the sentiment triple is less than a preset threshold, at which point the iteration stops, thus completing the decoupling of the latent distribution feature vectors between the feature vectors of the sentiment triple.
[0173] Using the anchor feature vector of the timbre triplet as the cluster center, the distance between the positive example feature vector and the anchor feature vector is reduced, and the distance between the negative example feature vector and the anchor feature vector is increased. The timbre triplet is iterated until the loss function value of the timbre triplet is less than a preset threshold, at which point the iteration stops, thus completing the decoupling of the latent distribution feature vectors between the feature vectors of the timbre triplet.
[0174] The concatenated feature vector is obtained by splicing the decoupled feature vectors together.
[0175] In one embodiment, reconstructing the waveform from the spliced feature vector and the linear spectral feature vector to obtain the synthesized waveform includes:
[0176] The spliced feature vector is restored to a first time-domain signal, and the linear spectrum feature vector is restored to a second time-domain signal;
[0177] The first time-domain signal and the second time-domain signal are spliced together to obtain the synthesized waveform.
[0178] like Figure 3 The diagram shown is a structural schematic of an electronic device for implementing a singing voice conversion method according to an embodiment of the present invention.
[0179] In this embodiment, the electronic device 1 includes, but is not limited to, a memory 11, a processor 12, and a network interface 13 that can be interconnected via a system bus. The memory 11 stores a singing conversion program 10, which can be executed by the processor 12. Figure 3 Only the electronic device 1, which includes components 11-13 and the singing conversion program 10, is shown. Those skilled in the art will understand that... Figure 3 The structure shown does not constitute a limitation on the electronic device 1, and may include fewer or more components than shown, or combine certain components, or have different component arrangements.
[0180] The memory 11 includes RAM and at least one type of readable storage medium. The RAM provides a cache for the operation of the electronic device 1; the readable storage medium can be a non-volatile storage medium such as flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the readable storage medium can be an internal storage unit of the electronic device 1; in other embodiments, the non-volatile storage medium can also be an external storage device of the electronic device 1, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the electronic device 1. In this embodiment, the readable storage medium of the memory 11 is typically used to store the operating system and various application software installed on the electronic device 1, such as storing the code of the song conversion program 10 in one embodiment of the present invention. Furthermore, the memory 11 can also be used to temporarily store various types of data that have been output or will be output.
[0181] In some embodiments, processor 12 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. This processor 12 is typically used to control the overall operation of the electronic device 1, such as performing control and processing related to data interaction or communication with other devices. In this embodiment, the processor 12 is used to run program code stored in the memory 11 or process data, for example, running a vocal conversion program 10.
[0182] The network interface 13 may include a wireless network interface or a wired network interface, which is used to establish a communication connection between the electronic device 1 and the terminal (not shown in the figure).
[0183] Optionally, the electronic device 1 may further include a user interface, which may include a display, an input unit such as a keyboard, and optionally, a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen, etc. The display may also be appropriately referred to as a screen or display unit, used to display information processed in the electronic device 1 and to display a visual user interface.
[0184] It should be understood that the embodiments described are for illustrative purposes only and are not limited to this structure in the scope of the patent application.
[0185] The singing conversion program 10 stored in the memory 11 of the electronic device 1 is a combination of multiple instructions, which, when run in the processor 12, can achieve the following:
[0186] Obtain a first song sample from a first singer and a second song sample from a second singer. Extract the first timbre feature vector and the first emotion feature vector from the first song sample. Extract the second timbre feature vector and the second emotion feature vector from the second song sample.
[0187] An emotion triplet is constructed using the first emotion feature vector, the second emotion feature vector, and the emotion label feature vector of the emotion category to which the first song sample belongs; and a timbre triplet is constructed using the first timbre feature vector, the second timbre feature vector, and the timbre label feature vector of the timbre category to which the first song sample belongs.
[0188] Decouple the feature vectors of the emotion triplet and the timbre triplet, and then concatenate the decoupled feature vectors to obtain the concatenated feature vector.
[0189] Extract the linear spectrum feature vector of the first song sample, and reconstruct the waveform using the spliced feature vector and the linear spectrum feature vector to obtain the synthesized waveform;
[0190] Calculate the difference in probability distribution between the synthesized waveform and the real waveform of the first song sample. When the difference is less than a threshold, convert the first song sample from the voice of the first singer to the voice of the second singer to obtain the converted song.
[0191] Specifically, the processor 12's implementation method for the aforementioned singing voice conversion program 10 can be found in [reference needed]. Figure 1 The descriptions of the relevant steps in the corresponding embodiments are not repeated here.
[0192] Furthermore, if the modules / units integrated in the electronic device 1 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable medium can be non-volatile or non-combustible. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, or a read-only memory (ROM).
[0193] The computer-readable storage medium stores a singing voice conversion program 10, which can be executed by one or more processors. The specific implementation of the computer-readable storage medium of the present invention is basically the same as the embodiments of the singing voice conversion method described above, and will not be repeated here.
[0194] In the several embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.
[0195] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0196] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.
[0197] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0198] Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within the invention. No appended diagram markings in the claims should be construed as limiting the scope of the claims.
[0199] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in a system claim may also be implemented by a single unit or device through software or hardware. The term "second class" is used to indicate names and does not indicate any specific order.
[0200] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for converting singing voices, characterized in that, The method includes: Obtain a first song sample from a first singer and a second song sample from a second singer. Extract the first timbre feature vector and the first emotion feature vector from the first song sample. Extract the second timbre feature vector and the second emotion feature vector from the second song sample. An emotion triplet is constructed using the first emotion feature vector, the second emotion feature vector, and the emotion label feature vector of the emotion category to which the first song sample belongs; and a timbre triplet is constructed using the first timbre feature vector, the second timbre feature vector, and the timbre label feature vector of the timbre category to which the first song sample belongs. Decouple the feature vectors of the emotion triplet and the timbre triplet, and then concatenate the decoupled feature vectors to obtain the concatenated feature vector. Extract the linear spectrum feature vector of the first song sample, and reconstruct the waveform using the spliced feature vector and the linear spectrum feature vector to obtain the synthesized waveform; Calculate the difference in probability distribution between the synthesized waveform and the real waveform of the first song sample. When the difference is less than a threshold, convert the first song sample from the voice of the first singer to the voice of the second singer to obtain the converted song.
2. The singing voice conversion method as described in claim 1, characterized in that, The step of extracting the first timbre feature vector and the first emotion feature vector of the first singer from the first song sample includes: Calculate the first Mel frequency cepstral coefficient of the first song sample, and extract the first timbre feature vector of the first singer from the first Mel frequency cepstral coefficient; The first emotional feature vector of the first singer is extracted from the audio signal of the first song sample, wherein the audio signal includes at least one of the following: spectral envelope signal, formant shift signal, and transient analysis signal.
3. The singing voice conversion method as described in claim 1, characterized in that, The step of extracting the second singer's second timbre feature vector and second emotional feature vector from the second song sample includes: Calculate the second Mel frequency cepstral coefficients of the second song sample, and extract the second timbre feature vector of the second singer from the second Mel frequency cepstral coefficients; The second emotional feature vector of the second singer is extracted from the audio signal of the second song sample, wherein the audio signal includes at least one of the following: spectral envelope signal, formant shift signal, and transient analysis signal.
4. The singing voice conversion method as described in claim 1, characterized in that, The sentiment triplet includes a negative example feature vector, a positive example feature vector, and an anchor point feature vector. The step of constructing the sentiment triplet using the first sentiment feature vector, the second sentiment feature vector, and the sentiment label feature vector of the sentiment category to which the first song sample belongs includes: The first emotion feature vector is used as the negative example feature vector of the emotion triple, the second emotion feature vector is used as the positive example feature vector of the emotion triple, and the emotion tag feature vector is used as the anchor feature vector of the emotion triple.
5. The singing voice conversion method as described in claim 1, characterized in that, The timbre triplet includes a negative example feature vector, a positive example feature vector, and an anchor point feature vector. The step of constructing the timbre triplet using the first timbre feature vector, the second timbre feature vector, and the timbre label feature vector of the timbre category to which the first song sample belongs includes: The first timbre feature vector is used as the negative example feature vector of the timbre triplet, the second timbre feature vector is used as the positive example feature vector of the timbre triplet, and the timbre tag feature vector is used as the anchor point feature vector of the timbre triplet.
6. The singing voice conversion method as described in claim 1, characterized in that, The process of decoupling the feature vectors of the emotion triplet and the timbre triplet, and then concatenating the decoupled feature vectors to obtain the concatenated feature vector, includes: Using the anchor feature vector of the sentiment triple as the cluster center, the distance between the positive example feature vector and the anchor feature vector is reduced, and the distance between the negative example feature vector and the anchor feature vector is increased. The sentiment triple is iterated until the loss function value of the sentiment triple is less than a preset threshold, at which point the iteration stops, thus completing the decoupling of the latent distribution feature vectors between the feature vectors of the sentiment triple. Using the anchor feature vector of the timbre triplet as the cluster center, the distance between the positive example feature vector and the anchor feature vector is reduced, and the distance between the negative example feature vector and the anchor feature vector is increased. The timbre triplet is iterated until the loss function value of the timbre triplet is less than a preset threshold, at which point the iteration stops, thus completing the decoupling of the latent distribution feature vectors between the feature vectors of the timbre triplet. The concatenated feature vector is obtained by splicing the decoupled feature vectors together.
7. The singing voice conversion method as described in claim 1, characterized in that, The process of reconstructing the waveform from the spliced feature vector and the linear spectral feature vector to obtain the synthesized waveform includes: The spliced feature vector is restored to a first time-domain signal, and the linear spectrum feature vector is restored to a second time-domain signal; The first time-domain signal and the second time-domain signal are spliced together to obtain the synthesized waveform.
8. A singing voice conversion device, characterized in that, The device includes: The acquisition module is used to acquire a first song sample of a first singer and a second song sample of a second singer, extract a first timbre feature vector and a first emotion feature vector of the first singer from the first song sample, and extract a second timbre feature vector and a second emotion feature vector of the second singer from the second song sample; The construction module is used to construct an emotion triplet using the first emotion feature vector, the second emotion feature vector, and the emotion label feature vector of the emotion category to which the first song sample belongs; and to construct a timbre triplet using the first timbre feature vector, the second timbre feature vector, and the timbre label feature vector of the timbre category to which the first song sample belongs. The decoupling module is used to decouple each feature vector of the emotion triplet and the timbre triplet, and then splice the decoupled feature vectors to obtain the spliced feature vector. The reconstruction module is used to extract the linear spectrum feature vector of the first song sample, and reconstruct the waveform from the spliced feature vector and the linear spectrum feature vector to obtain the synthesized waveform; The conversion module is used to calculate the difference value of the probability distribution between the synthesized waveform and the real waveform of the first song sample. When the difference value is less than a threshold, the first song sample is converted from the singing voice of the first singer to the singing voice of the second singer to obtain the converted song.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores a vocal conversion program that can be executed by the at least one processor, the vocal conversion program being executed by the at least one processor to enable the at least one processor to perform the vocal conversion method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a singing voice conversion program, which can be executed by one or more processors to implement the singing voice conversion method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Singing synthesis method and device, computer device and storage medium
CN113555001A
Speech synthesis method, speech synthesis device, electronic equipment and storage medium
CN116665639A