Method and apparatus for audio processing
By using a large-scale self-supervised music understanding model to extract audio representations, the problem of fundamental frequency extraction error in singing conversion is solved, achieving higher accuracy and effect in singing conversion.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
- Filing Date
- 2025-05-19
- Publication Date
- 2026-07-21
Smart Images

Figure CN120510853B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio technology, and in particular to an audio processing method and apparatus. Background Technology
[0002] Vocal conversion can transform audio into a specified timbre, that is, it can synthesize audio with a specified timbre, and is often used in the pitch correction function of singing applications.
[0003] Currently, vocal conversion relies on the fundamental frequency of the audio to be converted. However, when the song contains frying sounds or harmonies, fundamental frequency extraction is very prone to errors, resulting in poor vocal conversion quality. Summary of the Invention
[0004] This application provides an audio processing method and apparatus that can solve the problem of poor vocal conversion quality caused by fundamental frequency extraction errors. The technical solution is as follows:
[0005] In a first aspect, an audio processing method is provided, the method comprising:
[0006] Acquire the target timbre audio and extract the target timbre features from the target timbre audio;
[0007] The target representation model is used to extract features from the audio to be converted, thereby obtaining the audio representation of the audio to be converted; the target representation model is a music understanding model based on large-scale self-supervised training.
[0008] The target timbre features and the audio representation are input into the trained singing conversion model to obtain the converted target audio, wherein the timbre of the target audio is the target timbre, and the content and melody of the target audio are the content and melody of the audio to be converted.
[0009] In one possible implementation, the method further includes:
[0010] Obtain target pitch shift information, which indicates the target number of pitches offset from the pitch of the audio to be converted;
[0011] The step of inputting the target timbre features and the audio representation into the trained singing voice conversion model to obtain the converted target audio includes:
[0012] The target timbre features, the audio representation, and the target pitch shifting information are input into the trained singing conversion model to obtain the converted target audio, wherein the pitch of the target audio is shifted by the target number of pitches based on the pitch of the audio to be converted.
[0013] In one possible implementation, obtaining the target shift information includes:
[0014] Receive target relocation information input by the user; or,
[0015] Target pitch information is calculated based on the pitch information of the target timbre audio and the pitch information of the audio to be converted.
[0016] In one possible implementation, inputting the target timbre features and the audio representation into the trained singing voice conversion model includes:
[0017] The target timbre features, the audio representation, and the target spectrum are input into the trained singing conversion model, wherein the target spectrum is a spectrum with zeros set.
[0018] Secondly, a training method for a singing voice conversion model is provided, the method comprising:
[0019] Obtain sample audio, wherein the timbre of the sample audio is the timbre of the first audio, and the content and melody of the sample audio are the content and melody of the second audio;
[0020] The audio representation of the sample audio is obtained by extracting features from the sample audio using a target representation model.
[0021] Obtain the timbre characteristics and first spectrum of the second audio;
[0022] The first spectrum is randomly masked to obtain the second spectrum;
[0023] The audio representation, the timbre feature, and the second spectrum are input into the singing conversion model to be trained to obtain the converted predicted spectrum.
[0024] The singing conversion model to be trained is trained based on the first spectrum and the predicted spectrum.
[0025] In one possible implementation, the method further includes:
[0026] Obtain sample shift information;
[0027] The step of inputting the audio representation, the timbre features, and the second spectrum into the singing voice conversion model to be trained includes:
[0028] The audio representation, the timbre features, the second spectrum, and the sample modulation information are input into the singing conversion model to be trained.
[0029] In one possible implementation, obtaining the sample shift information includes:
[0030] Based on the pitch information of the first audio and the pitch information of the second audio, the sample pitch shifting information is calculated.
[0031] Thirdly, an audio processing apparatus is provided, the apparatus comprising:
[0032] The acquisition module is used to acquire the target timbre audio and extract the target timbre features of the target timbre audio; and to extract features of the audio to be converted through a target representation model to obtain the audio representation of the audio to be converted; the target representation model is a music understanding model based on large-scale self-supervised training.
[0033] The conversion module is used to input the target timbre features and the audio representation into the trained singing conversion model to obtain the converted target audio, wherein the timbre of the target audio is the target timbre, and the content and melody of the target audio are the content and melody of the audio to be converted.
[0034] In one possible implementation, the acquisition module is further configured to:
[0035] Obtain target pitch shift information, which indicates the target number of pitches offset from the pitch of the audio to be converted;
[0036] The conversion module is used for:
[0037] The target timbre features, the audio representation, and the target pitch shifting information are input into the trained singing conversion model to obtain the converted target audio, wherein the pitch of the target audio is shifted by the target number of pitches based on the pitch of the audio to be converted.
[0038] In one possible implementation, the acquisition module is configured to:
[0039] Receive target relocation information input by the user; or,
[0040] Target pitch information is calculated based on the pitch information of the target timbre audio and the pitch information of the audio to be converted.
[0041] In one possible implementation, the conversion module is used to:
[0042] The target timbre features, the audio representation, and the target spectrum are input into the trained singing conversion model, wherein the target spectrum is a spectrum with zeros set.
[0043] Fourthly, a training device for a singing voice conversion model is provided, the device comprising:
[0044] An acquisition module is used to acquire sample audio, wherein the timbre of the sample audio is the timbre of a first audio, and the content and melody of the sample audio are the content and melody of a second audio; feature extraction is performed on the sample audio through a target representation model to obtain the audio representation of the sample audio; the timbre features and a first spectrum of the second audio are acquired; and the first spectrum is randomly masked to obtain the second spectrum;
[0045] The training module is used to input the audio representation, the timbre feature, and the second spectrum into the singing conversion model to be trained to obtain the converted predicted spectrum; and to train the singing conversion model to be trained based on the first spectrum and the predicted spectrum.
[0046] In one possible implementation, the acquisition module is further configured to:
[0047] Obtain sample shift information;
[0048] The training module is used for:
[0049] The audio representation, the timbre features, the second spectrum, and the sample modulation information are input into the singing conversion model to be trained.
[0050] In one possible implementation, the acquisition module is configured to:
[0051] Based on the pitch information of the first audio and the pitch information of the second audio, the sample pitch shifting information is calculated.
[0052] Fifthly, a computing device is provided, the computing device including a processor and a memory, the memory storing at least one instruction, the instruction being loaded and executed by the processor to perform the operations performed as described in the first aspect and any possible method of audio processing described in the first aspect.
[0053] In a sixth aspect, a computing device is provided, the computing device including a processor and a memory, the memory storing at least one instruction, the instruction being loaded and executed by the processor to perform the operations performed as described in the second aspect above and any possible training method of the singing voice conversion model described in the second aspect.
[0054] In a seventh aspect, a computer-readable storage medium is provided, the storage medium storing at least one instruction, the instruction being loaded and executed by a processor to perform the operations performed as described in the first aspect and any possible implementation of the audio processing method described in the first aspect.
[0055] Eighthly, a computer-readable storage medium is provided, the storage medium storing at least one instruction, the instruction being loaded and executed by a processor to perform the operations performed by the training method of the singing voice conversion model described in the second aspect above and any possible implementation of the second aspect.
[0056] In a ninth aspect, a computer program product is provided, the computer program product storing at least one instruction, the instruction being loaded and executed by a processor to perform the operations performed as described in the first aspect and any possible implementation of the audio processing method described in the first aspect.
[0057] In a tenth aspect, a computer program product is provided, the computer program product storing at least one instruction, the instruction being loaded and executed by a processor to perform the operations performed by the training method of the singing voice conversion model described in the second aspect above and any possible implementation of the second aspect.
[0058] The beneficial effects of the technical solution provided in this application are:
[0059] In this method, for the audio to be converted, an audio representation is extracted using a music understanding model based on large-scale self-supervised training. For the target timbre audio, its timbre features are extracted. Then, the audio representation and timbre features are input into a trained singing conversion model to obtain an audio with the target timbre, the content of the audio to be converted, and the melody of the audio to be converted. It is evident that this method uses the audio representation extracted by the music understanding model based on large-scale self-supervised training instead of the fundamental frequency. This audio representation already contains the melody information of the audio, eliminating the need for separate fundamental frequency extraction. The extraction of the audio representation is more accurate and less prone to errors compared to fundamental frequency extraction, avoiding the problem of poor singing conversion results caused by errors in fundamental frequency extraction. Attached Figure Description
[0060] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0061] Figure 1 This is a flowchart of an audio processing method provided in an embodiment of this application;
[0062] Figure 2 This is a flowchart of an audio processing method provided in an embodiment of this application;
[0063] Figure 3 This is a flowchart of a training method for a singing voice conversion model provided in an embodiment of this application;
[0064] Figure 4 This is a schematic diagram of a data perturbation provided in an embodiment of this application;
[0065] Figure 5 This is a flowchart of a training method for a singing voice conversion model provided in an embodiment of this application;
[0066] Figure 6 This is a schematic diagram of an audio processing device provided in an embodiment of this application;
[0067] Figure 7 This is a schematic diagram of the training device structure for the singing voice conversion model provided in this application embodiment;
[0068] Figure 8 This is a schematic diagram of the structure of a computing device provided in an embodiment of this application;
[0069] Figure 9 This is a schematic diagram of the structure of a computing device provided in an embodiment of this application. Detailed Implementation
[0070] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0071] This application provides an audio processing method, which can be implemented by a computing device, such as a terminal or a server. The terminal can be a mobile phone, computer, tablet, etc., and the server can be a single server or a server cluster. The terminal may have a music application installed, which may include pitch correction and music synthesis functions, such as the ability to synthesize songs with a specified timbre. When a user needs to convert a song to a target timbre, they simply select the audio to be converted and the target timbre in the aforementioned pitch correction and music synthesis function interfaces. The terminal then processes the audio to be converted and the target timbre to generate an audio with the target timbre, the content of the audio to be converted, and the melody of the audio to be converted. In some implementations, the terminal can upload the audio to be converted (or its identifier) and the target timbre (or its identifier) to a server. The server then processes the audio to be converted and the target timbre to generate an audio with the target timbre, the content of the audio to be converted, and the melody of the audio to be converted, and sends it to the terminal. The terminal and the server perform the same operations when processing the audio to be converted and the audio of the target timbre. This application embodiment takes the terminal performing the above operations as an example for explanation.
[0072] Currently, the aforementioned technologies rely on the fundamental frequency of the audio to be converted to achieve the above functions. However, when the song contains vocal fry or harmony, the fundamental frequency extraction is very prone to errors, resulting in poor vocal conversion.
[0073] This application provides an audio processing method. In this method, for the audio to be converted, an audio representation is extracted using a music understanding model based on large-scale self-supervised training. For the target timbre audio, its timbre features are extracted. Then, the audio representation and timbre features are input into a trained singing conversion model to obtain an audio with the target timbre, the content of the audio to be converted, and the melody of the audio to be converted. It is evident that in this method, the audio representation extracted by the music understanding model based on large-scale self-supervised training replaces the fundamental frequency. This audio representation already contains the melody information of the audio, eliminating the need for separate fundamental frequency extraction. The extraction of the audio representation is more accurate and less prone to errors compared to fundamental frequency extraction, avoiding the problem of poor singing conversion results caused by errors in fundamental frequency extraction.
[0074] The audio processing method provided in the embodiments of this application is described below with reference to the accompanying drawings. This method can be implemented by a terminal. See [link to documentation]. Figure 1 The method may include the following steps:
[0075] Step 101: Obtain the target timbre audio and extract the target timbre features from the target timbre audio.
[0076] Among them, the target timbre feature is used to characterize the target timbre of the target timbre audio.
[0077] In practice, if a user wants to convert the timbre of the audio to be converted to the target timbre, they can obtain the target timbre audio through the terminal. The timbre of the target timbre audio is the target timbre.
[0078] For example, the terminal may have a music application installed, which may have a pitch correction function option. The user can select this pitch correction function option to enter the pitch correction function interface. In the pitch correction function interface, the user selects a target audio tone. The target audio tone can be audio that the user recorded and stored on the terminal. Then, the terminal can read the target audio tone from its local storage space.
[0079] In some implementations, the terminal can also obtain the target timbre audio via the Internet or automatically generate the target timbre audio. This application embodiment does not limit the method of obtaining the target timbre audio.
[0080] After acquiring the target timbre audio, the terminal can extract the target timbre features from the target timbre audio, wherein the target timbre features can be in the form of vectors. The extraction of target timbre features can employ methods such as MFCC (Mel-Frequency Cepstral Coefficients) and deep learning models. This application embodiment does not limit the method used for timbre feature extraction.
[0081] Step 102: Extract features from the audio to be converted using the target representation model to obtain the audio representation of the audio to be converted.
[0082] The target representation model is mert (Music Understanding Model with Large-Scale Self-supervised Training).
[0083] In practice, users can obtain the audio to be converted through their terminals. For example, in the aforementioned audio editing function interface, users can select the audio to be converted, and then the terminal can obtain the audio locally or via the Internet.
[0084] After acquiring the audio to be converted, it is input into mert to obtain the audio representation of the audio to be converted, which is called the mert representation. The mert representation contains the content, timbre, and melody information of the audio to be converted.
[0085] Step 103: Input the target timbre features and audio representation into the trained singing conversion model to obtain the converted target audio.
[0086] The singing voice conversion model can be an SVC (Singing-voice-conversion) model based on conditional flow matching. The internal structure of this model can be stacked DiT (diffusion transformer) blocks. The transformer module is a deep learning model based on a self-attention mechanism. The timbre of the target audio is the target timbre, and the content and melody of the target audio are the content and melody of the audio to be converted.
[0087] In implementation, after obtaining the target timbre features of the target audio and the audio representation of the audio to be converted, the target timbre features and audio representation are input into the trained singing voice conversion model. The singing voice conversion model can directly output the converted target audio. Alternatively, the singing voice conversion model can output the converted spectrum, which is then input into a vocoder, and the vocoder outputs the converted target audio. The spectrum can be a Mel spectrum.
[0088] In one possible implementation, before inputting the audio representation into the trained vocal conversion model, the audio representation can first be input into a first encoder to obtain an audio encoding, and then the audio encoding can be input into the trained vocal conversion model. The first encoder can be a transformer module, a conformer module, etc., where the conformer module is a convolution-enhanced transformer model.
[0089] In one possible implementation, step 103 above can also be processed as follows:
[0090] The target timbre features, audio representation, and target spectrum are input into the trained singing conversion model, where the target spectrum is a spectrum with zeros set. Here, the spectrum can be a Mel spectrum.
[0091] In one possible implementation, the method provided in this application embodiment can also support a pitch shifting function. Accordingly, the terminal can obtain target pitch shifting information, which indicates the target number of keys (pitches) to be shifted based on the pitch of the audio to be converted. Based on this, the processing in step 103 above can also be as follows:
[0092] The target timbre features, audio representation, and target pitch shifting information are input into the trained vocal conversion model. The vocal conversion model can directly output the converted target audio. Alternatively, the vocal conversion model can output the converted spectrum, which is then input into a vocoder, and the vocoder outputs the converted target audio. In this case, the pitch of the target audio is shifted by the target number of keys from the pitch of the audio to be converted.
[0093] There are multiple methods for obtaining target shift information, and two examples are given below:
[0094] Method 1:
[0095] In the pitch correction interface of music applications, there may be a pitch shift input box. Users can enter the target pitch shift information in the input box to indicate the key to be shifted. The terminal can then obtain the target pitch shift information entered by the user. For example, entering "+1" means raising the pitch of the audio to be converted by 1 key. Entering "-1" means lowering the pitch of the audio to be converted by 1 key.
[0096] Method 2:
[0097] The terminal calculates the target pitch shift information based on the pitch information of the target timbre audio and the pitch information of the audio to be converted. The pitch information can be the MIDI (Musical Instrument Digital Interface) average. Accordingly, the processing in Method Two is as follows:
[0098] Calculate the target pitch shift information using the MIDI mean of the target audio and the MIDI mean of the audio to be converted.
[0099] Specifically, the terminal calculates the fundamental frequency mean of the target audio timbre and converts it into a MIDI value to obtain the MIDI mean of the target audio timbre. The terminal also calculates the fundamental frequency mean of the audio to be converted and converts it into a MIDI value to obtain the MIDI mean of the audio to be converted. Finally, the terminal subtracts the MIDI mean of the audio to be converted from the MIDI mean of the target audio timbre to obtain the target pitch shift information.
[0100] In one possible implementation, before inputting the target pitch shifting information into the trained vocal conversion model, the target pitch shifting information can first be input into a second encoder. The second encoder encodes the one-dimensional target pitch shifting information into a multi-dimensional vector, which serves as the pitch shifting code. This pitch shifting code is then input into the trained vocal conversion model. The second encoder can be an embedding module, also known as an embedding layer, embedding model, etc.
[0101] In one possible implementation, such as Figure 2 As shown in the embodiment of this application, feature extraction can be performed on the target timbre to obtain target timbre features. Feature extraction can be performed on the audio to be converted to obtain a Mert representation. The Mert representation is input into the first encoder to obtain the Mert encoding. The target transposition information is input into the second encoder to obtain the transposition encoding. Then, the target timbre features, Mert representation, transposition encoding and the zeroed-out Mel spectrum are concatenated and input into the trained singing conversion model to obtain the converted Mel spectrum. The converted Mel spectrum is then input into the vocoder to obtain the converted target audio.
[0102] In the method provided in this application embodiment, the mert representation of the audio to be converted is used instead of the fundamental frequency. The mert representation already contains the melody information of the audio to be converted, so there is no need to extract the fundamental frequency separately. The extraction of mert features is more accurate and less prone to errors than the extraction of the fundamental frequency, which can avoid the problem of poor singing conversion effect caused by the error in the extraction of the fundamental frequency.
[0103] The training of the singing voice conversion model is explained below with reference to the accompanying diagram. Training can be performed using computing devices, such as servers or server clusters. See also... Figure 3 The training method for the singing voice conversion model provided in this application embodiment may include the following steps:
[0104] Step 201: Obtain sample audio.
[0105] The timbre of the sample audio is the same as that of the first audio, and the content and melody of the sample audio are the same as those of the second audio.
[0106] In implementation, such as Figure 4 As shown, the computing device acquires a first audio and a second audio, and then inputs the first audio and the second audio into a singing conversion system in the related technology to obtain a sample audio. The timbre of the sample audio is the timbre of the first audio, and the content and melody of the sample audio are the content and melody of the second audio.
[0107] The first audio file can be an audio file with a specified timbre, which can be uploaded to the computing device by relevant personnel, or obtained by the computing device via the internet or from a music library. The second audio file can be obtained from the music library.
[0108] Step 202: Extract features from the sample audio using the target representation model to obtain the audio representation of the sample audio. The target representation model is Mert.
[0109] In practice, the computing device inputs the sample audio into mert to obtain the audio representation of the sample audio. This audio representation, also known as the mert representation, contains the content, timbre, and melodic information of the sample audio.
[0110] Step 203: Obtain the timbre characteristics and first spectrum of the second audio.
[0111] Among them, the timbre features can be in the form of vectors, and the first spectrum can be the Mel spectrum.
[0112] In implementation, the computing device can extract the timbre features of the second audio and extract the first spectrum of the first audio. The extraction of the target timbre features can be performed using the MFCC (Mel-Frequency Cepstral Coefficients) method, deep learning models, etc. In this application embodiment, the method used for timbre feature extraction is not limited, as long as the same algorithm is used to extract the timbre features in steps 101 and 203 above.
[0113] Step 204: Perform a random mask on the first spectrum to obtain the second spectrum.
[0114] In practice, in order for the singing conversion model to learn the converted melody in the audio representation, the computing device needs to perturb the first spectrum. Specifically, the first spectrum can be randomly masked to obtain the second spectrum.
[0115] Step 205: Input the audio representation of the sample audio, the timbre features of the second audio, and the second spectrum into the singing conversion model to be trained to obtain the converted predicted spectrum.
[0116] In practice, the audio representation of the sample audio, the timbre features of the second audio, and the second spectrum are input into the singing conversion model to be trained. The singing conversion model to be trained outputs the prediction vector field corresponding to each time step according to the specified time step. Then, the ordinary differential equation is solved through the prediction vector field to obtain the converted prediction spectrum.
[0117] Step 206: Train the singing conversion model to be trained based on the first spectrum and the predicted spectrum.
[0118] In implementation, the loss value between the predicted vector field at each time step and the target vector field at that time step in the first spectrum is calculated. The loss value can be L1 loss, also known as MAE (Mean Absolute Error). Then, the singing conversion model to be trained is trained based on the obtained loss value, so that the predicted vector field approaches the target vector field.
[0119] In one possible implementation, to enable the above singing conversion model to have a pitch-shifting function, the processing in step 201 above can be as follows:
[0120] The first audio, the second audio, and the sample pitch shifting information are input into a singing conversion system in related technologies to obtain a sample audio. The timbre of the sample audio is the same as that of the first audio, and the content and melody of the sample audio are the same as those of the second audio. Furthermore, the pitch of the sample audio is shifted from the second audio by the key indicated by the sample pitch shifting information. The method for calculating the sample pitch shifting information can be as follows:
[0121] Based on the pitch information of the first audio and the pitch information of the second audio, the pitch shift information of the samples is calculated. The pitch information can be the MIDI mean, and the calculation method can be as follows:
[0122] Calculate the fundamental frequency mean of the first audio audio and convert it to a MIDI value to obtain the MIDI mean of the first audio audio. Calculate the fundamental frequency mean of the second audio audio and convert it to a MIDI value to obtain the MIDI mean of the second audio audio. Then, subtract the MIDI mean of the second audio audio from the MIDI mean of the first audio audio to obtain the sample pitch shift information.
[0123] Accordingly, the processing in step 205 above can be as follows:
[0124] The audio representation of the sample audio, the timbre features of the second audio, the second spectrum, and the pitch shifting information of the sample are input into the singing conversion model to be trained.
[0125] In one possible implementation, the audio representation of the sample audio and the pitch shifting information of the sample audio can be encoded before being input into the singing conversion model to be trained. For example... Figure 5 As shown, the Mert representation (i.e., audio representation) of the sample audio is input into the first encoder to obtain the sample Mert encoding, and the sample pitch shifting information is input into the second encoder to obtain the sample pitch shifting encoding. Then, the timbre features of the second audio, the second spectrum, the sample Mert encoding, and the sample pitch shifting encoding are concatenated and input into the singing conversion model to be trained.
[0126] All of the above-mentioned optional technical solutions can be combined in any way to form the optional embodiments of this application, and will not be described in detail here.
[0127] Based on the same technical concept, embodiments of this application also provide an audio processing apparatus, which can be applied to computing devices, such as... Figure 6 As shown, the device includes an acquisition module 610 and a conversion module 620, wherein:
[0128] The acquisition module 610 is used to acquire the target timbre audio, extract the target timbre features of the target timbre audio; and extract features of the audio to be converted through a target representation model to obtain the audio representation of the audio to be converted; the target representation model is a music understanding model based on large-scale self-supervised training.
[0129] The conversion module 620 is used to input the target timbre features and the audio representation into the trained singing conversion model to obtain the converted target audio, wherein the timbre of the target audio is the target timbre, and the content and melody of the target audio are the content and melody of the audio to be converted.
[0130] In one possible implementation, the acquisition module 610 is further configured to:
[0131] Obtain target pitch shift information, which indicates the target number of pitches offset from the pitch of the audio to be converted;
[0132] The conversion module 620 is used for:
[0133] The target timbre features, the audio representation, and the target pitch shifting information are input into the trained singing conversion model to obtain the converted target audio, wherein the pitch of the target audio is shifted by the target number of pitches based on the pitch of the audio to be converted.
[0134] In one possible implementation, the acquisition module 610 is used to:
[0135] Receive target relocation information input by the user; or,
[0136] Target pitch information is calculated based on the pitch information of the target timbre audio and the pitch information of the audio to be converted.
[0137] In one possible implementation, the conversion module 620 is configured to:
[0138] The target timbre features, the audio representation, and the target spectrum are input into the trained singing conversion model, wherein the target spectrum is a spectrum with zeros set.
[0139] In the solution provided in this application embodiment, the mert representation of the audio to be converted is used instead of the fundamental frequency. The mert representation already contains the melody information of the audio to be converted, so there is no need to extract the fundamental frequency separately. The extraction of mert features is more accurate and less prone to errors than the extraction of the fundamental frequency, which can avoid the problem of poor singing conversion effect caused by the error in the extraction of the fundamental frequency.
[0140] It should be noted that the audio processing apparatus provided in the above embodiments is only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computing device can be divided into different functional modules to complete all or part of the functions described above. In addition, the audio processing apparatus provided in the above embodiments and the method embodiments for audio processing belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0141] Based on the same technical concept, embodiments of this application also provide a training device for a singing voice conversion model. This method can be applied to computing devices, such as... Figure 7 As shown, the device includes an acquisition module 710 and a training module 720, wherein:
[0142] The acquisition module 710 is used to acquire sample audio, wherein the timbre of the sample audio is the timbre of a first audio, and the content and melody of the sample audio are the content and melody of a second audio; the sample audio is used to extract features through a target representation model to obtain the audio representation of the sample audio; the timbre features and a first spectrum of the second audio are acquired; and the first spectrum is randomly masked to obtain the second spectrum;
[0143] The training module 720 is used for the audio representation, the timbre features and the second spectrum, inputting the singing conversion model to be trained to obtain the converted predicted spectrum; and training the singing conversion model to be trained based on the first spectrum and the predicted spectrum.
[0144] In one possible implementation, the acquisition module 710 is further configured to:
[0145] Obtain sample shift information;
[0146] The training module 720 is used for:
[0147] The audio representation, the timbre features, the second spectrum, and the sample modulation information are input into the singing conversion model to be trained.
[0148] In one possible implementation, the acquisition module 710 is used for:
[0149] Based on the pitch information of the first audio and the pitch information of the second audio, the sample pitch shifting information is calculated.
[0150] It should be noted that the training device for the singing voice conversion model provided in the above embodiments is only illustrated by the division of the above functional modules during the training of the singing voice conversion model. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computing device can be divided into different functional modules to complete all or part of the functions described above. In addition, the training device for the singing voice conversion model provided in the above embodiments and the training method embodiments for the singing voice conversion model belong to the same concept, and the specific implementation process is detailed in the method embodiments, which will not be repeated here.
[0151] Figure 8This illustration shows a structural block diagram of an electronic device 800 provided in an exemplary embodiment of this application. The electronic device 800 may be a portable mobile terminal, such as a smartphone, tablet computer, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 player (Moving Picture Experts Group Audio Layer IV), laptop computer, or desktop computer. The electronic device 800 may also be referred to as a user device, portable terminal, laptop terminal, desktop terminal, or other names.
[0152] Typically, electronic device 800 includes a processor 801 and a memory 802.
[0153] Processor 801 may include one or more processing cores, such as a quad-core processor or an octa-core processor. Processor 801 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 801 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 801 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 801 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0154] The memory 802 may include one or more computer-readable storage media, which may be non-transitory. The memory 802 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 802 are used to store at least one instruction, which is executed by the processor 801 to implement the audio processing method provided in the method embodiments of this application.
[0155] In some embodiments, the electronic device 800 may optionally include a peripheral device interface 803 and at least one peripheral device. The processor 801, memory 802, and peripheral device interface 803 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 803 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of the following: a radio frequency circuit 804, a display screen 805, a camera assembly 806, an audio circuit 807, a positioning assembly 808, and a power supply 809.
[0156] Peripheral device interface 803 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 801 and memory 802. In some embodiments, processor 801, memory 802 and peripheral device interface 803 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 801, memory 802 and peripheral device interface 803 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.
[0157] The radio frequency (RF) circuit 804 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 804 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 804 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 804 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 804 can communicate with other terminals through at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 804 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.
[0158] Display screen 805 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 805 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 801 for processing. In this case, display screen 805 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 805, disposed on the front panel of electronic device 800; in other embodiments, there may be at least two display screens, disposed on different surfaces of electronic device 800 or in a folded design; in still other embodiments, display screen 805 may be a flexible display screen, disposed on a curved or folded surface of electronic device 800. Furthermore, display screen 805 may be configured as a non-rectangular irregular shape, i.e., a non-rectangular screen. Display screen 805 may be made of materials such as LCD (Liquid Crystal Display) or OLED (Organic Light-Emitting Diode).
[0159] The camera assembly 806 is used to acquire images or videos. Optionally, the camera assembly 806 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the terminal, and the rear-facing camera is located on the back of the terminal. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 806 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm-light flash and a cool-light flash, which can be used for light compensation at different color temperatures.
[0160] The audio circuit 807 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting the sound waves into electrical signals that are input to the processor 801 for processing, or input to the radio frequency circuit 804 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each located in a different part of the electronic device 800. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert the electrical signals from the processor 801 or the radio frequency circuit 804 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 807 may also include a headphone jack.
[0161] The positioning component 808 is used to locate the current geographical location of the electronic device 800 in order to enable navigation or LBS (Location Based Service). The positioning component 808 can be a positioning component based on the US GPS (Global Positioning System), China's BeiDou system, or Russia's Galileo system.
[0162] Power supply 809 is used to supply power to various components in electronic device 800. Power supply 809 can be alternating current, direct current, a disposable battery, or a rechargeable battery. When power supply 809 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, while a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.
[0163] In some embodiments, the electronic device 800 further includes one or more sensors 810. The one or more sensors 810 include, but are not limited to: an accelerometer 811, a gyroscope 812, a pressure sensor 813, a fingerprint sensor 814, an optical sensor 815, and a proximity sensor 816.
[0164] Accelerometer 811 can detect the magnitude of acceleration on the three coordinate axes of a coordinate system established by electronic device 800. For example, accelerometer 811 can be used to detect the components of gravitational acceleration on the three coordinate axes. Processor 801 can control display screen 805 to display the user interface in either a landscape or portrait view based on the gravitational acceleration signal acquired by accelerometer 811. Accelerometer 811 can also be used for games or for acquiring user motion data.
[0165] The gyroscope sensor 812 can detect the orientation and rotation angle of the electronic device 800. The gyroscope sensor 812, in conjunction with the accelerometer sensor 811, can collect 3D motion data from the user on the electronic device 800. Based on the data collected by the gyroscope sensor 812, the processor 801 can perform the following functions: motion sensing (e.g., changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.
[0166] The pressure sensor 813 can be disposed on the side bezel of the electronic device 800 and / or on the lower layer of the display screen 805. When the pressure sensor 813 is disposed on the side bezel of the electronic device 800, it can detect the user's grip signal on the electronic device 800, and the processor 801 can perform left / right hand recognition or quick operation based on the grip signal collected by the pressure sensor 813. When the pressure sensor 813 is disposed on the lower layer of the display screen 805, the processor 801 can control the operable controls on the UI interface based on the user's pressure operation on the display screen 805. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.
[0167] The fingerprint sensor 814 is used to collect a user's fingerprint. The processor 801 identifies the user based on the fingerprint collected by the fingerprint sensor 814, or vice versa. When the user's identity is verified as trusted, the processor 801 authorizes the user to perform relevant sensitive operations, including unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings. The fingerprint sensor 814 can be located on the front, back, or side of the electronic device 800. When the electronic device 800 has a physical button or manufacturer logo, the fingerprint sensor 814 can be integrated with the physical button or manufacturer logo.
[0168] An optical sensor 815 is used to collect ambient light intensity. In one embodiment, the processor 801 can control the display brightness of the display screen 805 based on the ambient light intensity collected by the optical sensor 815. Specifically, when the ambient light intensity is high, the display brightness of the display screen 805 is increased; when the ambient light intensity is low, the display brightness of the display screen 805 is decreased. In another embodiment, the processor 801 can also dynamically adjust the shooting parameters of the camera assembly 806 based on the ambient light intensity collected by the optical sensor 815.
[0169] A proximity sensor 816, also known as a distance sensor, is typically located on the front panel of an electronic device 800. The proximity sensor 816 is used to detect the distance between the user and the front of the electronic device 800. In one embodiment, when the proximity sensor 816 detects that the distance between the user and the front of the electronic device 800 is gradually decreasing, the processor 801 controls the display screen 805 to switch from a screen-on state to a screen-off state; when the proximity sensor 816 detects that the distance between the user and the front of the electronic device 800 is gradually increasing, the processor 801 controls the display screen 805 to switch from a screen-off state to a screen-on state.
[0170] Those skilled in the art will understand that Figure 8 The structure shown does not constitute a limitation on the electronic device 800, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0171] Figure 9 This is a schematic diagram of a server structure provided in an embodiment of this application. The server 1000 can vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) 1001 and one or more memories 1002. The memories 1002 store at least one instruction, which is loaded and executed by the processors 1001 to implement the methods provided in the various method embodiments described above. Of course, the server may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The server may also include other components for implementing device functions, which will not be elaborated upon here.
[0172] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including instructions that can be executed by a processor in a terminal to perform the audio processing method or the singing voice conversion model training method in the above embodiments. The computer-readable storage medium may be non-transitory. For example, the computer-readable storage medium may be ROM (Read-Only Memory), RAM (Random Access Memory), CD-ROM (Compact Disc Read-Only Memory), magnetic tape, floppy disk, and optical data storage device, etc.
[0173] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals (including but not limited to signals transmitted between the user terminal and other devices) involved in this application have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, all audio (such as first audio, second audio, target timbre audio, audio to be converted, etc.) and modulation information involved in this application were obtained with full authorization.
[0174] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0175] It should be understood that "multiple" as used herein refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. Furthermore, the step numbers described herein are merely illustrative of one possible execution order. In some other embodiments, the steps may not be executed in numerical order, such as two steps with different numbers being executed simultaneously, or two steps with different numbers being executed in the reverse order of the illustration. This application does not limit this.
[0176] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. An audio processing method, characterized in that, The method includes: Acquire the target timbre audio and extract the target timbre features from the target timbre audio; Obtain target pitch shift information, which indicates the target number of pitches to be shifted from the pitch of the audio to be converted; The target representation model is used to extract features from the audio to be converted to obtain the audio representation of the audio to be converted; the target representation model is the music understanding model MERT based on large-scale self-supervised training. The target timbre features, the target pitch shifting information, and the audio representation are input into the trained singing conversion model to obtain the converted target audio. The timbre of the target audio is the target timbre represented by the target timbre features, the content and melody of the target audio are the content and melody of the audio to be converted, and the pitch of the target audio is shifted from the pitch of the audio to be converted by the target number of pitches.
2. The method according to claim 1, characterized in that, The acquisition of target relocation information includes: Receive target relocation information input by the user; or, Target pitch information is calculated based on the pitch information of the target timbre audio and the pitch information of the audio to be converted.
3. The method according to any one of claims 1-2, characterized in that, The step of inputting the target timbre features, the target modulation information, and the audio representation into the trained singing voice conversion model includes: The target timbre features, the audio representation, the target modulation information, and the target spectrum are input into the trained singing conversion model, wherein the target spectrum is a spectrum with zeros set.
4. A training method for a singing voice conversion model, characterized in that, The method includes: Obtain sample audio, wherein the timbre of the sample audio is the timbre of the first audio, and the content and melody of the sample audio are the content and melody of the second audio; Obtain sample shift information; The audio representation of the sample audio is obtained by extracting features from the sample audio using a target representation model; the target representation model is a music understanding model based on large-scale self-supervised training. Obtain the timbre characteristics and first spectrum of the second audio; The first spectrum is randomly masked to obtain the second spectrum; The audio representation, the timbre features, the sample modulation information, and the second spectrum are input into the singing conversion model to be trained to obtain the converted predicted spectrum. The singing conversion model to be trained is trained based on the first spectrum and the predicted spectrum.
5. The method according to claim 4, characterized in that, The acquisition of sample shift information includes: Based on the pitch information of the first audio and the pitch information of the second audio, the sample pitch shifting information is calculated.
6. A computing device, characterized in that, The computing device includes a processor and a memory, the memory storing at least one instruction that is loaded and executed by the processor to perform the operation performed by the audio processing method as described in any one of claims 1-3, or the operation performed by the training method of the singing voice conversion model as described in any one of claims 4-5.
7. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction, which is loaded and executed by a processor to perform the operation performed by the audio processing method as described in any one of claims 1-3, or the operation performed by the training method of the singing voice conversion model as described in any one of claims 4-5.
8. A computer program product, characterized in that, The computer program product stores at least one instruction, which is loaded and executed by a processor to perform the operation performed by the audio processing method as described in any one of claims 1-3, or the operation performed by the training method of the singing voice conversion model as described in any one of claims 4-5.