Tone migration model training method, tone migration method and related product
By extracting and fusion of content features and tone features in audio data, and using the Unet network to train the tone migration model, the problem of separately training of the conversion model in the existing technology is solved, and the ability to transfer many tones is realized, reducing training costs.
Patent Information
- Application Number
- CN202510377462.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-06-10
AI Technical Summary
Existing tone migration algorithms require training of separate transformation models for the transfer of two specific tones, resulting in high training costs.
By obtaining audio data of various instruments, extracting content features and tone characteristics, and fusing them to generate Gaussian noise vectors and time step embedding vectors, the tone migration model is trained using the Unet network to achieve many-to-many tone migration.
A model is implemented to migrate multiple instruments' tones, reducing training costs and improving processing capabilities for complex tones.
Smart Images

Figure CN120126499A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of timbre transfer, and in particular to a training method for a timbre transfer model, a timbre transfer method, and related products. Background Art
[0002] Timbre is an attribute of sound and is the basis for a listener to distinguish two different sounds. It plays a decisive role in the process of humans distinguishing different musical instruments and sounds. Compared with other more objective characteristics of sound, timbre has more subjective colors and represents more complex regularities and dependence conditions. The characteristics it exhibits are often also related to human high-level cognitions such as sounding object recognition and emotion recognition. With the development of technology and the rise of interdisciplinary disciplines, timbre, which combines perception and acoustic laws, this unique musical composition feature has gradually become a hot topic in the field of music information retrieval.
[0003] The timbre transfer task is a challenging subtask of the timbre generation task and has gradually attracted more researchers' interest in recent years with the rise of voice conversion technology. The main purpose of the general research on timbre transfer (Timbre Transfer) is to apply the auditory characteristics of one musical instrument to another, and transfer as many timbre performance characteristics of the target instrument as possible to the source instrument while keeping the pitch content unchanged. For example, retaining the pitch content of a violin performance and converting it into a piano performance. As a special technology of timbre generation technology, timbre transfer technology can facilitate the selection and arrangement of musical instruments in the music production process, and can also serve fields such as music education and digital art.
[0004] Traditional timbre transfer technologies generally rely on the stacked use of some signal processing modules. This method not only requires a lot of professional knowledge but also is difficult to solve the problem of reconstructing complex timbres. With the significant progress of deep learning and neural network technologies in the timbre generation task, end-to-end timbre transfer algorithms based on deep learning have gradually become the mainstream method of timbre transfer technology due to their advantages such as convenience, high efficiency, and the ability to process complex timbres. Timbre transfer algorithms based on deep learning generally follow the architecture of deep generative models, and through the clustering analysis ability of neural networks, extract and reorganize the pitch information and timbre information in the audio and finally restore the data distribution of the target audio to achieve the purpose of timbre transfer. However, many existing timbre transfer algorithms based on balanced data need to train a separate conversion model for the transfer of specific two timbres, which greatly increases the training cost of timbre transfer. Summary of the Invention
[0005] The purpose of this application is to provide a training method for a timbre transfer model, a timbre transfer method, and related products, which can perform many-to-many timbre transfer and have the ability to transfer multiple musical instruments.
[0006] To achieve the above object, the present application provides the following solutions:
[0007] In a first aspect, the present application provides a method for training a timbre transfer model. The method for training the timbre transfer model includes:
[0008] Obtain audio data of various musical instruments;
[0009] Calculate the Mel spectrogram of the audio data;
[0010] Extract the content feature and the timbre feature from the audio data respectively; the content feature is extracted from the audio data providing the content; the timbre feature is extracted from the audio data providing the timbre;
[0011] Fuse the content feature and the timbre feature and generate corresponding Gaussian noise vectors and time step embedding vectors;
[0012] Taking the Gaussian noise vector and the time step embedding vector as inputs, the Mel spectrogram data as the output, and the Mel spectrogram as the label, train the Unet network to obtain a timbre transfer model; the timbre transfer model is used to transfer the timbre of various musical instruments.
[0013] In a second aspect, the present application provides a timbre transfer method. The timbre transfer method includes:
[0014] Obtain a waveform file of the content to be transferred and a waveform file of the target timbre;
[0015] Extract the content feature from the waveform file of the content to be transferred to obtain the content feature to be transferred;
[0016] Extract the timbre feature from the waveform file of the target timbre to obtain the timbre feature to be transferred;
[0017] Fuse the timbre feature to be transferred and the content feature to be transferred and generate corresponding Gaussian noise vectors and time step embedding vectors to obtain a transferred Gaussian noise vector and a transferred time step embedding vector;
[0018] Input the transferred Gaussian noise vector and the transferred time step embedding vector into the timbre transfer model to obtain transferred Mel spectrogram data; the timbre transfer model is trained by the method for training the timbre transfer model described in any one of the above.
[0019] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the computer program to implement the method for training the timbre transfer model or the timbre transfer method described in any one of the above.
[0020] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the training method or the timbre migration method of the timbre migration model described in any one of the above.
[0021] In a fifth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the training method or the timbre migration method of the timbre migration model described in any one of the above.
[0022] According to the specific embodiments provided by the present application, the following technical effects are disclosed in the present application:
[0023] The present application provides a training method for a timbre migration model, a timbre migration method and related products. The method includes: obtaining audio data of various musical instruments; extracting the Mel spectrogram of the audio data; respectively extracting the content feature and the timbre feature in the audio data; the content feature is extracted from the audio data providing the content; the timbre feature is extracted from the audio data providing the timbre; fusing the content feature and the timbre feature and generating a Gaussian noise vector and a time step embedding vector; using the Gaussian noise vector and the time step embedding vector as inputs, the Mel spectrogram data as outputs, and the Mel spectrogram as labels, training a Unet network to obtain a timbre migration model; the timbre migration model is used to migrate the timbres of various musical instruments. In the present application, through obtaining and migrating training on the audio data of various musical instruments, a model capable of performing many-to-many timbre migration is realized through the Unet network, that is, a model has the ability to migrate multiple musical instruments. Description of the Drawings
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0025] Figure 1 It is an application environment diagram of a training method for a timbre migration model in an embodiment of the present application;
[0026] Figure 2 It is a flowchart of a training method for a timbre migration model provided by an embodiment of the present application;
[0027] Figure 3 It is a schematic diagram of two sub-datasets Nsynth-Sub and Nsynth-MIni provided by an embodiment of the present application;
[0028] Figure 4 Schematic diagram of audio data acquisition provided by an embodiment of the present application;
[0029] Figure 5 Schematic diagram of the feature fusion method process when processing the content features of a single timbre provided by an embodiment of the present application;
[0030] Figure 6 Schematic diagram of the process when processing the content features of a melodic timbre provided by an embodiment of the present application;
[0031] Figure 7 Schematic diagram of the DCUnet network structure provided by an embodiment of the present application;
[0032] Figure 8 Schematic diagram of the ResUnet network structure provided by an embodiment of the present application;
[0033] Figure 9 Schematic diagram of the timbre transfer framework based on the denoising diffusion generative model provided by an embodiment of the present application;
[0034] Figure 10 Schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0035] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0036] In order to solve the problem that "existing timbre transfer algorithms need to train separate conversion models for the transfer training of specific two timbres", the present application designs a model capable of performing many-to-many timbre transfer, that is, a model has the ability to transfer multiple musical instruments.
[0037] Moreover, in order to better constrain the generation results, avoid fuzzy training results, and improve the utilization rate of features, the present application designs multiple sets of feature fusion schemes to fuse different features from multiple angles, and finally can respectively achieve two tasks of single timbre pitch transfer and melodic timbre transfer.
[0038] In the past, the method of timbre generation through adversarial networks was difficult to train and easily fell into the result of mode collapse. Therefore, the present application uses a diffusion generative network as the main architecture to solve the problems of difficult training process and poor training diversity.
[0039] To make the above objects, features, and advantages of the present application more apparent and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0040] The training method or timbre migration method of the timbre migration model provided by the embodiments of the present application can be applied to, for example, Figure 1 the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be set up separately, integrated on the server 104, placed on the cloud or other servers. Among them, the terminal 102 can be, but is not limited to, various desktop computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers, and can also be a cloud server.
[0041] In an exemplary embodiment, as Figure 2 shown, a training method of a timbre migration model is provided. This method is executed by a computer device, and can be specifically executed by a computer device such as a terminal or a server alone, or jointly executed by a terminal and a server. In the embodiments of the present application, taking this method applied to Figure 1 the server 104 therein as an example for illustration, it includes the following steps S1 to S6. Among them:
[0042] S1. Obtain audio data of various musical instruments.
[0043] In this embodiment, there are two types of data or sources. One is the original dataset Nsynth (Neural Audio Synthesis of Musical Notes with WaveNet Autoencoders), and the other is the POP909 dataset. The processing processes of the two datasets are similar, but their uses are different. If dealing with melodies, the POP909 dataset is used, and if dealing with single notes, the original dataset Nsynth is used.
[0044] This embodiment is based on the original dataset Nsynth. After manual selection and cleaning on the basis of the Nsynth dataset, most electronic timbres, significantly distorted timbres, timbres with a reverberation level that is extremely different from the average, and monophonic timbres with a pitch range outside 24 - 84 are removed. Then, according to the nature of the instruments, personalized pitch trimming is performed on different instruments. For example, the pitch range of the instrument Bass is between 45 - 64. Therefore, all data outside 45 - 64 is removed. Finally, two sub-datasets, Nsynth-Sub and Nsynth-MIni, are obtained. Nsynth-Mini is a small-scale dataset. For the generation experiment (which is a pre-experiment for timbre transfer) to verify the ability of the generation network, the small-scale dataset Nsynth-Mini is used for training and verification. In the formal experiment - the timbre transfer experiment to verify the timbre transfer method, the NsynthSub dataset is used for training and verification, as Figure 3 shown.
[0045] This embodiment builds a multi-melody timbre dataset: POPSUB based on the POP909 dataset. For the two hundred songs in POP909, three segments of suitable-range, independent, and ten-second-long melodies are selected from the melody midi (Musical Instrument Digital Interface) of each song as the audio content. For example, if a certain melody is to be converted into guitar audio data, a melody with a Pitch Range between 40 - 76, which conforms to the pitch range of the guitar instrument, needs to be selected. Since the midi needs to be converted into 6 timbres, the intersection of the pitch ranges of the six instruments is selected. Independent means that there is no overlapping part between the midi melodies. Then, in the digital audio workstation, the midi file is rendered through a commercial virtual instrument (rendering refers to a function button in the digital audio workstation, which can be directly clicked. This sentence can be understood as: create a project file in the digital audio workstation (similar to the project files in ps and pr), use a commercial virtual instrument (similar to an effect in ps) to convert the MIDI file into the sound effect of a real instrument, and finally select the wav format, that is, the waveform file, when exporting). Finally, the waveform data is exported as the dataset POPSUB for melody timbre transfer. After the above steps, a total of six instrument timbres (Guitar, E-Guitar, Flute, Piano, Percussion, String) are obtained. Each instrument timbre has 6000 sample audios with a length of 10s. All the audios have labels of the instrument name and content ID. For the audios with the same content ID, their audio content (referring to the melody) is exactly the same, that is, the data between the instruments is completely parallel data.
[0046] Therefore, the steps to obtain the timbres of these six musical instruments are as follows: Select three melodic segments with appropriate pitch ranges from each song in POP909 (a song midi dataset), each segment being 10 seconds long. For each segment, render it using six virtual instrument effects respectively and export it as a wav file. Six samples corresponding to each melody segment are obtained, which are the timbres of the six musical instruments, as Figure 4 shown.
[0047] S2. Calculate the Mel spectrogram of the audio data.
[0048] First, resample the audio data; frame it and apply a Hamming window, and convert the resampled data into spectral data using the Fourier transform method; take the modulus and square of the spectral data to obtain power spectral data; weight the power spectral data with a Mel filter bank to obtain Mel energy spectra; take the logarithm of the Mel energy spectra to obtain the final Mel spectrogram.
[0049] In this embodiment, each timbre audio file (audio data) in the dataset is resampled to 22050 Hz.
[0050] Convert the audio data (Wav waveform file) into a spectrogram using the short-time Fourier transform with a window length of 1024 and a sliding window step size of 256: The short-time Fourier transform (STFT) is a method for converting a signal from the time domain to the time-frequency domain. It divides the signal and performs Fourier transforms on each part to obtain the spectral characteristics of the signal at each moment, and is suitable for processing non-stationary signals such as audio. The following is the specific process of converting audio into a spectrogram using the short-time Fourier transform (STFT) with a window length of 1024 and a sliding window step size of 256:
[0051] 1. Signal preprocessing: Prepare the audio signal, that is, the audio dataset. Assume the length of the audio signal is N, that is, the total number of samples of the audio (total number of sampling points, for example, N = 22050 corresponds to 1 second, and the sampling rate is 22050 Hz).
[0052] 2. Select the window function:
[0053] The short-time Fourier transform needs to divide the audio signal into multiple short window functions (segments). Each window function is generally a fixed-size time period.
[0054] The window function can have different forms, and common ones include rectangular windows, Hamming windows, Hanning windows, etc. The choice of the window function will affect the spectral resolution. In this embodiment, the Hamming window is selected.
[0055] 3. The length and sliding step of the window function:
[0056] The window length is 1024, that is, the number of samples in each window is 1024.
[0057] The hop length is 256, which means that when the window function slides each time, the number of overlapping samples between adjacent windows is 1024 - 256 = 768.
[0058] The hop length determines the time points sampled for each Fourier transform. A smaller hop length means higher time-domain resolution but larger computational complexity.
[0059] 4. Framing:
[0060] The audio signal is divided into multiple frames according to a sliding window. The length of each frame is the window length (1024), and the position of each frame in the signal is determined by the hop length (256).
[0061] Assume the total length of the audio signal is N, and the total number of frames is:
[0062]
[0063] This is how the sliding window covers the entire audio signal.
[0064] 5. Fourier Transform
[0065] Perform a Fourier transform on each window function (each frame). The Fourier transform can convert the signal from the time domain to the frequency domain, represented in complex form, where the real part is the amplitude information and the imaginary part is the phase information.
[0066] Perform a discrete Fourier transform (DFT) on each frame to obtain the spectrum of that frame.
[0067] Mathematically, the formula for the short-time Fourier transform (STFT) is as follows:
[0068] X(f) = |X(f)|e jarg(X(f)) ;
[0069] where X(f) is the amplitude and arg(X(f)) is the phase.
[0070] 6. Obtaining the Power Spectrum
[0071] Take the modulus and square the obtained spectrum to get the power spectrum.
[0072] Taking the modulus of the spectrum is to calculate the amplitude part of the spectrum. The spectrum is in complex form, and taking the modulus is to calculate its amplitude, ignoring its phase information. The modulus of the spectrum is:
[0073]
[0074] where is the real part of the spectrum, is the imaginary part of the spectrum.
[0075] Squaring the amplitude of the spectrum gives the power information of the signal. The power spectrum represents the energy distribution of the signal at each frequency, which is equivalent to representing the energy intensity of the signal at a certain frequency. The result after squaring is: P(f) = |X(f)| 2 .
[0076] Weight the power spectrum with a Mel filter bank to obtain the Mel energy spectrum.
[0077] The Mel scale is a non - linear frequency scale, which is based on the perceptual characteristics of the human ear for frequency. The relationship between the frequency f (in Hertz) and the Mel frequency mmm (in Mel) is given by the following formula:
[0078]
[0079] where f is the traditional frequency (Hz) and m is the Mel frequency.
[0080] Construct the Mel filter bank: The key step in the Mel spectrum is to weight the traditional spectrum through the Mel filter bank. Each Mel filter covers a certain range of frequencies and is divided according to the Mel scale. In this embodiment, 80 Mel filters are selected. The following is the specific operation after determining the number of Mel filters:
[0081] Define the frequency range of the filter: Convert the frequency axis of the original spectrum so that it is in units of the Mel scale.
[0082] Select the upper and lower boundaries of the filter to cover the entire frequency range (usually from 0 to the Nyquist frequency).
[0083] Design the filter: Each Mel filter is usually a triangular filter, which applies weights to different frequency components of the spectrum, and the width is distributed according to the Mel scale (the width refers to the frequency range covered by each Mel filter, calculated according to the Mel frequency scale above).
[0084] Apply the Mel filter bank: Apply the Mel filter bank to the obtained power spectrum and calculate the weighted sum of each frequency band. Specifically, the filter bank maps the frequency components of the power spectrum to the Mel frequency axis and calculates the output of each Mel filter: For each Mel filter, calculate its overlapping part (weighted sum) with the power spectrum. Assuming the power spectrum is P(f) and the response of the Mel filter is Hm(f), then the power of each Mel band is:
[0085]
[0086] where M m represents the power at the Mel frequency m.
[0087] Obtain the Mel energy spectrum: After processing the power spectrum of each time frame through a Mel filter bank, the Mel energy spectrum is obtained. The shape of the Mel energy spectrum is a two-dimensional matrix, where: the horizontal axis represents time (number of frames), and the vertical axis represents Mel frequency (usually the number of filters).
[0088] Take the logarithm of the Mel energy spectrum to obtain the final Mel spectrogram.
[0089] This patent unifies the size of the Mel spectrogram output for a single timbre and expands it uniformly into an input of 80×1000. And to meet the input requirements of denoising diffusion generation, the absolute values of all inputs are compressed between -1 and 1.
[0090] S3. Extract the content feature and timbre feature from the audio data respectively; the content feature is extracted from the audio data providing content; the timbre feature is extracted from the audio data providing timbre.
[0091] In this embodiment, when processing the content feature of a single timbre, this specific method can be selected: map the content feature and the timbre feature to dimensions of 1×124 and 1×128 respectively, splice the mapped results to obtain a feature embedding vector of size 1×252; add the feature embedding vector to a sine time step embedding vector of the same size to obtain a time step embedding vector; generate a Gaussian noise with the same size as the Mel spectrogram of the audio data to obtain a Gaussian noise vector.
[0092] When processing the content feature of a melodic timbre, this specific method can be selected: fuse the content feature and the timbre feature and generate corresponding Gaussian noise vectors and time step embedding vectors, specifically including:
[0093] Input the content feature into a chroma encoder composed of two two-dimensional convolutional layers and a pooling layer to obtain content information regularized to an appropriate size (12*1000).
[0094] Splice the content information with the initial noise of size (92×1000) to obtain a Gaussian noise vector with the same size as the Mel spectrogram of the audio data;
[0095] Add the timbre feature to a sine time step embedding vector of the same size to obtain a time step embedding vector.
[0096] Specifically speaking, the method for extracting the content feature in this embodiment may include the following methods:
[0097] Monophonic timbre: The YIN algorithm is a classic audio signal processing algorithm for fundamental frequency estimation. It is used to analyze the fundamental frequency in sound, that is, the main pitch in the sound. The monophonic audio data of the provided content is input into the YIN algorithm to obtain a frequency prediction array with the length of the number of frames. The frequency band with the most occurrences is obtained and converted into pitch as the pitch label.
[0098] Melodic timbre: For melodic timbre, Chroma is used to extract content information (i.e., pitch information). The Chroma vector, that is, the chroma vector represents the energy characterization of the pitch class distribution within an octave based on the equal-tempered scale per unit time. Energy is accumulated along the same pitch class in different octaves. The pitch class (chromagram) is an overall expression composed of chroma vectors. Combining music theory knowledge, it can be known that the pitch class feature (Chroma) is an audio feature describing the content of musical tonality. It maps the pitch order information into the color depth on a two-dimensional plane, so that the tonality and melody information can be reflected from the pitch class with time on the horizontal axis and pitch on the vertical axis. The extraction process of the pitch class feature is as follows: 1. The audio file is Fourier-transformed from the time domain to the frequency domain and noise reduction processing is performed. 2. The target is fine-tuned according to the standard frequency, and frame-level information is generated by combining the time length with the window length. 3. The pitch spectrogram is obtained by recording the frame-level energy, and the energies of the notes in different octaves sounding simultaneously are represented by different colors and superimposed on this pitch class to obtain a pitch class spectrum with chord information. Comparing the pitch classes (chromagrams) of the electric guitar and violin under the same melody, the pitch class feature has good robustness to instruments and can represent the pitch sequence information to a certain extent. Therefore, the pitch class feature is widely used in genre classification, music recommendation, and chord recognition.
[0099] Ways to extract timbre features:
[0100] Monophonic timbre: VGGish is a pre-trained deep neural network based on the VGG (Visual Geometry Group) architecture. This model consists of stacked two-dimensional convolutional layers, max-pooling layers, and fully connected layers. VGGish is pre-trained using the large-scale dataset AudioSet from YouTube, and its output features can be used for downstream tasks such as audio analysis and audio representation learning. This patent uses the pre-trained Viggish to extract a timbre feature vector of size (Duration×128) dimensions as the timbre feature for subsequent processing.
[0101] Melodic timbre: For timbre information, Vggish is not used for modeling, but the one-hot encoding of its instrument label is used to represent the difference in timbre. (For example, when inputting a waveform file that is a piano piece, the one-hot encoding of the piano is taken.)
[0102] S4. Fuse the content feature and the timbre feature and generate corresponding Gaussian noise vectors and time step embedding vectors.
[0103] Feature processing module: Input the audio providing content (Content Wav) and the audio providing timbre (StyleWav) into the feature processing module respectively for feature extraction and fusion. After being processed by the feature processing module, two special tensors fusing content and timbre features are output: the initial Gaussian noise vector (N-input) and the time step embedding vector (T-input).
[0104] When processing the content feature of a single timbre:
[0105] Concatenate the content feature and the timbre feature to obtain a feature embedding vector; fuse and add the feature embedding vector and the sinusoidal time step embedding vector to obtain a time step embedding vector; generate a random Gaussian noise with the same size as the mel spectrogram of the audio data to obtain a Gaussian noise vector.
[0106] Please refer to Figure 5 , in this method, the content feature (YIN-pitch label) extracted by the content feature extractor (extracting the pitch of the audio data providing content) is mapped to the dimension (1×124) through an embedding layer. Then, the timbre feature (with a size of 4×128) extracted by the VGGish model is adjusted to the corresponding timbre feature (1×128) after passing through a one-dimensional convolution to adjust its number of channels. Finally, the regularized timbre feature and the content feature are concatenated to form a feature embedding vector. The size of this timbre feature embedding vector is equal to the size of the time step embedding vector (1×252). Finally, add the timbre feature embedding vector and the time step embedding vector, that is, fuse all features in the time step vector. After feature fusion processing, the output of the downstream denoising generation module is a time embedding vector fusing timbre and pitch features and a pure Gaussian noise vector with the same size as the mel spectrogram.
[0107] Among them, the time step embedding vector = Sinusoidal Timestep Embedding is a time step embedding method commonly used in deep learning models. It was initially proposed by Vaswani et al. in the Transformer paper "Attention is All You Need" to represent the position information (Positional Encoding) of each time step in a sequence. In time series generation tasks, this method is also used to embed the timestep information, enabling the model to perceive the order of time or the features of time steps.
[0108] When processing the content features of the melody timbre:
[0109] Fuse and transform the content features and the timbre features into a Gaussian noise vector and a time step embedding vector, specifically including: passing the content features into a chroma encoder composed of two two-dimensional convolutional layers and a pooling layer to obtain content information regularized to a suitable size (12 * 1000); passing the content features into a chroma encoder composed of two two-dimensional convolutional layers and a pooling layer to obtain content information regularized to a suitable size (12 * 1000); adding the timbre features to the sine time step embedding vector to obtain a time step embedding vector.
[0110] Please refer to Figure 6 , the melody timbre feature fusion method: Use the one-hot encoding of its instrument label to represent the difference in timbre (where "its" refers to the input wav file. For example, if an input waveform file is a piano piece, then the one-hot encoding of the piano is adopted). For the extracted chroma (chromagram) features (12 * 435), a chroma encoder is designed to regularize the chroma features to a suitable size (12 * 1000) through two two-dimensional convolutional layers and a pooling layer. For the concatenation of the content information (equivalent to Contentwav in the above text) and the initial noise, the method of concatenation in dimension two is used, that is, Figure 6 as shown in
[0111] S5. Using the Gaussian noise vector and the time step embedding vector as inputs, the Mel spectrogram data as outputs, and the Mel spectrogram as labels, train the Unet network to obtain a timbre transfer model; the timbre transfer model is used to transfer the timbre of various instruments.
[0112] In this embodiment, the Unet network is selected as the denoising generation module: The two output tensors of the noise (N-input) and the time step embedding vector (T-input) will be used as the inputs of the denoising generation module for denoising generation, and the generated Mel spectrogram will be used as the target of the real data Mel spectrogram to optimize the network. This embodiment designs two Unet architecture networks for Mel spectrogram generation as Mel generators. The Unet structure network is usually used in the denoising generation architecture because the Unet structure is good at processing images with noise and can retain both detailed information and structural information at the same time. The basic framework of the Unet structure is to capture the global features and local features of the image through several upsampling and downsampling processes and use skip connections to retain the subtle features and edge information in the data.
[0113] Please refer to Figure 7 , DCUnet (Double Convolution Unet) is improved from the denoising network in the classical image generation task based on the denoising diffusion model. The design of two convolutions can make an intermediate output with a dimensional size to buffer the instability caused by sudden dimensional changes. As Figure 7 shown, in the DCUnet architecture, the input data undergoes three upsampling processes and downsampling processes, and time embeddings with feature conditions are involved each time the resolution changes. After fusing the specified conditions, the network will combine the conditional content for a purposeful conditional generation process. In the upsampling module, the data passes through two convolutional layers and a self-attention layer to expand the dimension, and finally passes through a max pooling layer to reduce the resolution. In the downsampling module, after the data fuses the output of the skip connection before, it also passes through two convolutional layers and a self-attention layer to reduce the dimension, and then the resolution is enlarged by the method of interpolation upsampling. In the bottleneck layer (Bottleneck), three stacked two-dimensional convolutions are used for deep feature extraction and recombination.
[0114] Please refer to Figure 8 , the module of ResUnet (Residual U-Net) adds a connection composed of a single two-dimensional convolution for residual connection on the basis of DCUnet, thus combining the characteristics of Unet and Resnet. The ResUnet architecture allows the underlying features to flow to the deeper layers, and while solving the vanishing gradient problem, it enables the network to learn higher-quality feature representations. ResUnet++ is a Unet network further improved based on ResUnet. It was initially designed for medical image segmentation tasks. It adds more residual connections and incorporates the ASPP (Atrous Spatial Pyramid Pooling) structure, which can capture features at multiple different spatial scales, while the ES (Excite&Seeze) model uses channel attention to strengthen the feature connections within the channel. As Figure 8 shown, in the downsampling process, using the residual connection, the Res_block module cooperates with the ES module of the channel attention mechanism to reduce the resolution through the max pooling layer, and further feature fusion is performed and conditional feature time embeddings are added. The ASPP layer replaces the bottleneck layer in DCUnet. It obtains three-scale outputs through three two-dimensional convolutions of different sizes, and all of them are concatenated and output through the output layer. In the upsampling part, the mutual attention module is used for the fusion of the residual connection, and the fused result passes through a module with a residual connection and a transposed convolution layer (Transposed Conv2d) to increase the resolution.
[0115] Finally, in the inference stage, the generated Mel spectrogram will be converted into a waveform file through a vocoder.
[0116] Map the result with an absolute value between -1 and 1 generated by the denoising generation network back to the magnitude characteristics of the Mel spectrogram. Then, input the Mel spectrogram after feature mapping into the WaveGlow vocoder for inference to restore the audio.
[0117] Please refer to Figure 9 , in this embodiment, a single-timbre dataset Nsynth-Sub, Nsynth-Min, and a melodic-timbre dataset POPSUB are first established based on the Nsynth dataset. The wav files of the audio are cleaned and standardized as the input of the model, and the corresponding Mel spectrogram of the audio is generated. The Mel spectrogram generated by the model is used as the target to optimize the network.
[0118] Input the pitch content waveform file and the style timbre waveform file into the feature extraction module to extract the corresponding content features and timbre features, and then input these two features into the feature processing module. For the single-timbre transfer task, use the single-timbre feature fusion method to fuse the extracted features into the initial noise and time step embedding vectors. For melodic timbre transfer, use the melodic timbre transfer feature fusion method to extract the initial noise and time step embedding vectors.
[0119] Input the initial Gaussian noise vector and time step embedding vector obtained by the feature processing module into the denoising generation module for denoising generation. Two networks based on the Unet architecture are proposed in this paper: DCUnet and ResUnet++. The Mel spectrogram generated by the network is optimized with the Mel spectrogram of the real data as the target.
[0120] Finally, convert the generated Mel spectrogram into a waveform file through the WaveGlow vocoder. Through the above process, two tasks of single-timbre pitch transfer and melodic timbre transfer are achieved, and arbitrary transfer between different instruments can be performed.
[0121] In an exemplary embodiment, a timbre transfer method is provided, and the method includes:
[0122] A1. Obtain a waveform file of the content to be transferred and a waveform file of the target timbre.
[0123] A2. Extract the content features in the waveform file of the content to be transferred to obtain the content features to be transferred.
[0124] A3. Extract the timbre features in the waveform file of the target timbre to obtain the timbre features to be transferred.
[0125] A4. Fuse the to-be-migrated timbre features and the to-be-migrated content features and generate corresponding Gaussian noise vectors and time step embedding vectors to obtain a migrated Gaussian noise vector and a migrated time step embedding vector.
[0126] A5. Input the migrated Gaussian noise vector and the migrated time step embedding vector into the timbre migration model to obtain migrated Mel spectrogram data; the timbre migration model is trained by the training method of the timbre migration model described above.
[0127] After inputting the migrated Gaussian noise vector and the migrated time step embedding vector into the timbre migration model to obtain migrated Mel spectrogram data, it further includes:
[0128] Use the WaveGlow vocoder to generate migrated audio data based on the migrated Mel spectrogram data.
[0129] This embodiment is specifically described below by taking the example of obtaining a piano timbre audio:
[0130] Prepare input files: Prepare an audio with a piano performance as a style timbre waveform file (style wav) (to-be-migrated timbre waveform file), and a pitch content waveform file (content wav) (target content waveform file) with any timbre but the target pitch (monophonic timbre) / melody (melodic timbre).
[0131] Feature extraction module: Extract the content features in the audio data of the provided content and the timbre features of the audio data of the provided timbre respectively.
[0132] If the waveform file is a monophonic timbre, use the YIN algorithm to extract the main pitch in the sound as the content feature of the audio data. Use the pre-trained VGGish network to extract a timbre feature vector of size (Duration×128) as the timbre feature for subsequent processing.
[0133] If the waveform file is a melodic timbre, use Chroma to extract the content information (i.e., pitch information). The Chroma vector, that is, the chroma vector represents the energy characterization of the pitch class distribution within an octave based on the equal temperament within the unit time. For the timbre feature, use the one-hot encoding of its instrument label to represent the difference in timbre. (For example, input a waveform file, and if the waveform file is a piano piece, then take the one-hot encoding of the piano.)
[0134] Feature processing module: Input the extracted content features and timbre features into the feature processing module, fuse them and generate corresponding Gaussian noise vectors and time step embedding vectors.
[0135] If the waveform file is a single timbre, the extracted content features are mapped to an appropriate dimension through an embedding layer to form a 1×124 vector. The timbre features are adjusted in terms of the number of channels through a one-dimensional convolution to obtain corresponding timbre feature (1×128) vectors. Then, they are concatenated in the way of Figure 5 .
[0136] If the waveform file is a melodic timbre, as shown in Figure 6 , the content features are fed into a chroma encoder composed of two two-dimensional convolutional layers and a pooling layer to obtain content information regularized to an appropriate size (12*1000); the content information is concatenated with initial noise of size (92×1000) to obtain a Gaussian noise vector equal in size to the Mel spectrogram of the audio data; the timbre features are added to a sine time step embedding vector of the same size to obtain a time step embedding vector.
[0137] Denoising generation module: The output of the feature processing module is input into the denoising generation module, and one of the following two methods can be selected.
[0138] Method 1: DCUnet (Double Convolution Unet). Method 2: ResUnet (Residual U-Net).
[0139] Vocoder: The output of the denoising generation module is input into the vocoder waveglow and converted into a waveform file. This waveform file is the target audio.
[0140] In an exemplary embodiment, a computer device is provided. This computer device can be a server or a terminal, and its internal structure diagram can be as shown in Figure 10 . The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of this computer device is used to provide computing and control capabilities. The memory of this computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of this computer device is used to exchange information between the processor and external devices. The communication interface of this computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a training method for a timbre migration model or a timbre migration method.
[0141] Those skilled in the art can understand that Figure 10The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0142] In an exemplary embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0143] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0144] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0145] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0146] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random-access memories (ReRAMs), magnetoresistive random-access memories (MRAMs), ferroelectric random-access memories (FRAMs), phase change memories (PCMs), graphene memories, etc. Volatile memories can include random access memories (RAMs) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0147] The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0148] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0149] Specific examples are used in this article to elaborate on the principles and implementation manners of the present application. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A method for training a timbre migration model, characterized in that: The training method of the timbre migration model includes: Get audio data of various musical instruments; Calculate a Mel-spectrogram of the audio data; Extracting content features and timbre features from the audio data respectively; the content features are extracted from the audio data providing the content; the timbre features are extracted from the audio data providing the timbre; The content feature and the timbre feature are merged to generate a corresponding Gaussian noise vector and a time step embedding vector; The Gaussian noise vector and the time step embedding vector are used as input, the Mel spectrum data is used as output, and the Mel spectrum graph is used as a label to train a Unet network to obtain a timbre migration model; the timbre migration model is used to migrate the timbre of various musical instruments.
2. The training method of the timbre migration model according to claim 1, characterized in that: When processing the content feature of a single timbre: fusing the content feature and the timbre feature to generate a corresponding Gaussian noise vector and a time step embedding vector, specifically including: Mapping and concatenating the content feature and the timbre feature to obtain a feature embedding vector; Adding the feature embedding vector to the sinusoidal time step embedding vector to obtain a time step embedding vector; the sinusoidal time step embedding vector has the same size as the feature embedding vector; Generate a Gaussian noise with the same size as the Mel-spectrogram of the audio data to obtain a Gaussian noise vector.
3. The training method of the timbre migration model according to claim 1, characterized in that: When processing the content feature of the melody timbre: fusing the content feature and the timbre feature to generate a corresponding Gaussian noise vector and a time step embedding vector, specifically including: Passing the content features into a chroma encoder composed of two two-dimensional convolutional layers and a pooling layer to obtain content information; The content information is concatenated with the initial noise to obtain a Gaussian noise vector; the Gaussian noise vector has the same size as the Mel-spectrogram of the audio data; The timbre feature is added to a sinusoidal time step embedding vector of the same size to obtain a time step embedding vector.
4. The training method of the timbre migration model according to claim 3, characterized in that: The content information is concatenated with the initial noise to obtain a Gaussian noise vector; the Gaussian noise vector has the same size as the Mel-spectrogram of the audio data, specifically comprising: Regularize the content information into 12*1000 content data; The content data is concatenated with initial noise of size 92*1000 to obtain a Gaussian noise vector.
5. The training method of the timbre migration model according to claim 1, characterized in that: Calculating the Mel-spectrogram of the audio data specifically includes: Resampling the audio data to obtain resampled data; The resampled data is framed and Hamming windowed, and then converted into spectrum data using a Fourier transform method; Taking the modulus and squaring the frequency spectrum data to obtain power spectrum data; Using a Mel filter bank to perform weighted calculation on the power spectrum data to obtain a Mel energy spectrum; The logarithm of the Mel energy spectrum is taken to obtain a Mel frequency spectrum graph.
6. A timbre migration method, characterized in that: The timbre migration method comprises: Obtaining the content waveform file to be migrated and the target timbre waveform file; Extracting content features from the waveform file of the content to be migrated to obtain content features to be migrated; Extracting the timbre features in the target timbre waveform file to obtain the timbre features to be migrated; The timbre feature to be transferred and the content feature to be transferred are merged to generate corresponding Gaussian noise vectors and time step embedding vectors, so as to obtain transfer Gaussian noise vectors and transfer time step embedding vectors; The migration Gaussian noise vector and the migration time step embedding vector are input into a timbre migration model to obtain migration Mel spectrum data; the timbre migration model is trained by the training method of the timbre migration model described in any one of claims 1-5.
7. The timbre migration method according to claim 6, characterized in that: After the migration Gaussian noise vector and the migration time step embedding vector are input into the timbre migration model to obtain the migration Mel spectrum data, the method further includes: The Waveglow vocoder is used to generate migrated audio data based on the migrated Mel-spectrogram data.
8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the training method of the timbre migration model described in any one of claims 1-5 or the timbre migration method described in any one of claims 6-7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the training method of the timbre migration model described in any one of claims 1-5 or the timbre migration method described in any one of claims 6-7.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, it implements the training method of the timbre migration model described in any one of claims 1-5 or the timbre migration method described in any one of claims 6-7.