Audio processing method and device
By using CQT features instead of fundamental frequency and jointly training the CQT feature encoder and MIDI encoder, the problem of poor singing conversion caused by fundamental frequency extraction errors is solved, and converted audio with more accurate target timbre and melody is generated.
Patent Information
- Application Number
- CN202510850745.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-09-19
AI Technical Summary
When a song contains bubbling sounds and harmonies, fundamental frequency extraction is prone to errors, resulting in poor singing voice conversion effects.
CQT features are used to replace the fundamental frequency, and the converted audio is generated by combining CQT features, content features and timbre features. The timbre information interference in the CQT features is eliminated by jointly training the CQT feature encoder and the MIDI encoder.
The accuracy of singing voice conversion is improved, the problem of poor effect caused by incorrect fundamental frequency extraction is avoided, and the generated audio is closer to the target timbre and melody.
Smart Images

Figure CN120673729A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of audio technology, and in particular to a method and device for audio processing. Background Art
[0002] Singing voice conversion can convert the audio to be converted into a specified timbre, that is, it can synthesize audio with a specified timbre, and is often used for pitch correction in singing applications.
[0003] Currently, singing voice conversion relies on the fundamental frequency of the audio to be converted. However, when the song contains bubbling sounds and harmonies, fundamental frequency extraction is very prone to errors, resulting in poor singing voice conversion results. Summary of the Invention
[0004] The present invention provides an audio processing method and device that can solve the problem of poor singing voice conversion caused by fundamental frequency extraction errors. The technical solution is as follows:
[0005] In a first aspect, a training method for a singing voice conversion model is provided, wherein the singing voice conversion audio includes a CQT (Constant-Q Transform) feature encoder and decoder, and the method includes:
[0006] Acquire a first audio sample and a second audio sample, wherein the first audio sample and the second audio sample have the same timbre;
[0007] Extracting a target timbre feature of the first audio sample, wherein the target timbre feature is used to characterize the timbre of the first audio sample;
[0008] extracting a target MIDI (Musical Instrument Digital Interface) feature, a target CQT feature, and a target content feature of the second audio sample, wherein the target MIDI feature and the target CQT feature are used to characterize the melody of the second audio sample, and the target content feature is used to characterize the content of the second audio sample;
[0009] Inputting the target MIDI feature into a MIDI encoder to obtain a target MIDI encoding feature, and inputting the CQT feature into a CQT feature encoder to obtain a target CQTI encoding feature;
[0010] calculating a first loss value between the target MIDI code signature and the target CQT code signature based on the first loss value;
[0011] Selecting the target MIDI coding feature or the target coding CQT feature as a melody feature, and inputting the melody feature, the target content feature, and the target timbre feature into the decoder to obtain a converted spectrum feature;
[0012] Extract the target spectral features of the second sample audio, calculate a second loss value between the converted spectral features and the target spectral features, train the decoder based on the second loss value, and train the CQT feature encoder and the MIDI encoder based on the first loss value and the second loss value.
[0013] In a possible implementation, the first loss value is a mean square error loss value.
[0014] In a possible implementation, inputting the melody feature, the target content feature, and the target timbre feature into the decoder includes:
[0015] The melody feature, the target content feature and the target timbre feature are subjected to feature fusion to obtain a target fusion feature, and the target fusion feature is input into the decoder.
[0016] In a possible implementation, selecting the target MIDI feature or the target CQT feature as the melody feature includes:
[0017] The target MIDI feature or the target CQT feature is randomly selected as a melody feature.
[0018] In a possible implementation, obtaining the first audio sample and the second audio sample includes:
[0019] Get the target song audio;
[0020] Cutting out a continuous audio of a first duration from the target song audio;
[0021] The audio of the second duration before the audio of the first duration is used as the first sample audio;
[0022] The audio of the second duration after the audio of the first duration is used as the second sample audio.
[0023] In a second aspect, a method for audio processing is provided, the method comprising:
[0024] Acquire target timbre audio, and extract timbre features of the target timbre audio, wherein the timbre features are used to characterize the timbre of the target timbre audio;
[0025] Acquire audio to be converted, and extract CQT features and content features of the audio to be converted, wherein the CQT features are used to characterize the melody of the audio to be converted, and the content features are used to characterize the content of the audio to be converted;
[0026] Inputting the CQT feature into a CQT feature encoder to obtain a CQT encoding feature, wherein the CQT feature encoder is a CQT feature encoder trained according to any one of claims 1 to 5;
[0027] The timbre feature, the CQT coding feature and the content feature are input into a decoder to obtain converted audio, wherein the decoder is a decoder trained according to any one of claims 1 to 5.
[0028] In one possible implementation, the method further includes:
[0029] Get the rising and falling pitch values;
[0030] Based on the up-down adjustment value, performing up-down adjustment processing on the CQT feature;
[0031] The step of inputting the CQT feature into a CQT feature encoder comprises:
[0032] The CQT features after up- and down-tuning processing are input into the CQT feature encoder.
[0033] In a third aspect, a training device for a singing voice conversion model is provided, wherein the singing voice conversion audio includes a CQT feature encoder and a decoder, and the device includes:
[0034] an acquisition module, configured to acquire a first audio sample and a second audio sample, wherein the first audio sample and the second audio sample have the same timbre; extract a target timbre feature of the first audio sample, wherein the target timbre feature is used to characterize the timbre of the first audio sample; and extract a target MIDI feature, a target CQT feature, and a target content feature of the second audio sample, wherein the target MIDI feature and the target CQT feature are used to characterize the melody of the second audio sample, and the target content feature is used to characterize the content of the second audio sample;
[0035] A training module is configured to input the target MIDI feature into a MIDI encoder to obtain a target MIDI encoding feature, input the CQT feature into the CQT feature encoder to obtain a target CQT encoding feature; calculate a first loss value between the target MIDI encoding feature and the target CQT encoding feature; select the target MIDI encoding feature or the target encoding CQT feature as a melody feature, and input the melody feature, the target content feature, and the target timbre feature into the decoder to obtain a converted spectrum feature; extract the target spectrum feature of the second sample audio, calculate a second loss value between the converted spectrum feature and the target spectrum feature, train the decoder based on the second loss value, and train the CQT feature encoder and the MIDI encoder based on the first loss value and the second loss value.
[0036] In a possible implementation, the first loss value is a mean square error loss value.
[0037] In a possible implementation, the training module is used to:
[0038] The melody feature, the target content feature and the target timbre feature are subjected to feature fusion to obtain a target fusion feature, and the target fusion feature is input into the decoder.
[0039] In a possible implementation, the training module is used to:
[0040] The target MIDI feature or the target CQT feature is randomly selected as a melody feature.
[0041] In a possible implementation, the acquisition module is configured to:
[0042] Get the target song audio;
[0043] Cutting out a continuous audio of a first duration from the target song audio;
[0044] The audio of the second duration before the audio of the first duration is used as the first sample audio;
[0045] The audio of the second duration after the audio of the first duration is used as the second sample audio.
[0046] In a fourth aspect, an audio processing device is provided, the device comprising:
[0047] An acquisition module is configured to acquire target timbre audio and extract timbre features of the target timbre audio, wherein the timbre features are used to characterize the timbre of the target timbre audio; acquire audio to be converted and extract CQT features and content features of the audio to be converted, wherein the CQT features are used to characterize the melody of the audio to be converted and the content features are used to characterize the content of the audio to be converted;
[0048] A conversion module is configured to input the CQT feature into a CQT feature encoder to obtain a CQT coding feature, wherein the CQT feature encoder is a CQT feature encoder trained according to any one of claims 1-5; and input the timbre feature, the CQT coding feature, and the content feature into a decoder to obtain converted audio, wherein the decoder is a decoder trained according to any one of claims 1-5.
[0049] In a possible implementation, the acquisition module is further configured to:
[0050] Get the rising and falling pitch values;
[0051] Based on the up-down adjustment value, performing up-down adjustment processing on the CQT feature;
[0052] The conversion module is used to:
[0053] The CQT features after up- and down-tuning processing are input into the CQT feature encoder.
[0054] In a fifth aspect, a computing device is provided, comprising a processor and a memory, wherein the memory stores at least one instruction, and the instruction is loaded and executed by the processor to implement the operations performed by the training method of the singing conversion model as described in the first aspect and any possible implementation of the first aspect.
[0055] In a sixth aspect, a computing device is provided, comprising a processor and a memory, wherein the memory stores at least one instruction, and the instruction is loaded and executed by the processor to implement the operations performed by the audio processing method as described in the second aspect and any possible implementation of the second aspect.
[0056] In the seventh aspect, a computer-readable storage medium is provided, wherein the storage medium stores at least one instruction, and the instruction is loaded and executed by a processor to implement the operations performed by the training method of the singing voice conversion model as described in the first aspect and any possible implementation of the first aspect.
[0057] In an eighth aspect, a computer-readable storage medium is provided, wherein the storage medium stores at least one instruction, which is loaded and executed by a processor to implement the operations performed by the audio processing method as described in the first aspect and any possible implementation of the first aspect.
[0058] In the ninth aspect, a computer program product is provided, wherein at least one instruction is stored in the computer program product, and the instruction is loaded and executed by a processor to implement the operations performed by the training method of the singing voice conversion model as described in the first aspect and any possible implementation of the first aspect.
[0059] In a tenth aspect, a computer program product is provided, wherein at least one instruction is stored in the computer program product, and the instruction is loaded and executed by a processor to implement the operations performed by the audio processing method as described in the first aspect and any possible implementation of the first aspect.
[0060] The beneficial effects of the technical solution provided by this application are:
[0061] In this method, the CQT features and content features of the audio to be converted are extracted, and the timbre features of the target timbre audio are extracted. Then, by combining the CQT features, content features, and timbre features, a converted audio is generated that has the target timbre, the content of the audio to be converted, and the melody of the audio to be converted. This method uses CQT features instead of fundamental frequency to characterize the melody of the audio to be converted, eliminating the need to extract the fundamental frequency separately. This CQT feature extraction is more accurate and less prone to errors than fundamental frequency extraction, thus avoiding the problem of poor vocal conversion due to incorrect fundamental frequency extraction. In addition, considering that the CQT feature contains the timbre information of the song, in order to avoid the timbre information in the CQT feature from interfering with the audio conversion, the CQT feature encoder used in the embodiment of the present application is jointly trained with the MIDI encoder, so that the encoded CQT feature output by the CQT feature encoder is close to the encoded MIDI feature output by the MIDI encoder. The MIDI feature does not contain timbre information, but it is difficult to obtain and is not suitable for direct use in audio conversion. Through this method, the encoded CQT feature can be made close to the MIDI feature, and the timbre information contained therein before encoding can be eliminated as much as possible. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0063] Figure 1 This is a flow chart of an audio processing method provided by an embodiment of the present application;
[0064] Figure 2 This is a flow chart of an audio processing method provided by an embodiment of the present application;
[0065] Figure 3 This is a flow chart of a training method for a singing voice conversion model provided in an embodiment of the present application;
[0066] Figure 4 This is a flow chart of a training method for a singing voice conversion model provided in an embodiment of the present application;
[0067] Figure 5 This is a schematic diagram of the structure of an audio processing device provided in an embodiment of the present application;
[0068] Figure 6 Schematic diagram of the structure of the training device for the singing voice conversion model provided in an embodiment of the present application;
[0069] Figure 7 is a structural diagram of a computing device provided in an embodiment of the present application;
[0070] Figure 8 It is a structural diagram of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0071] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0072] The present invention provides an audio processing method that can be implemented by a computing device. The computing device can be a terminal or a server. The terminal can be a mobile phone, computer, tablet computer, etc., and the server can be a single server, a server cluster, etc. The terminal can be installed with a music application, which can include functions such as tuning and music synthesis, which can include synthesizing songs with a specified timbre. When a user needs to convert a song to a target timbre, they simply select the audio to be converted and the audio of the target timbre in the above-mentioned tuning and music synthesis function interfaces. The terminal can then process the audio to be converted and the audio of the target timbre to generate an audio file having the target timbre, the content of the audio to be converted, and the melody of the audio to be converted. In some implementations, the terminal can upload the audio to be converted and the audio of the target timbre to a server, which then processes the audio to be converted and the audio of the target timbre to generate an audio file having the target timbre, the content of the audio to be converted, and the melody of the audio to be converted, and transmits the file to the terminal. The terminal and server perform the same operations when processing the audio to be converted and the audio of the target timbre. The present invention uses the terminal performing the above operations as an example for explanation.
[0073] Currently, the implementation of the above functions in related technologies relies on the fundamental frequency of the audio to be converted. However, when a song contains bubbling sounds and harmonies, errors in fundamental frequency extraction are very likely to occur, resulting in poor singing conversion effects.
[0074] An embodiment of the present application provides a method for audio processing, in which, for the audio to be converted, its CQT (Constant-Q Transform) features and content features are extracted, and for the target timbre audio, its timbre features are extracted. Then, by combining the CQT features, content features, and timbre features, a converted audio having the target timbre, the content of the audio to be converted, and the melody of the audio to be converted is generated. As can be seen, in this method, the CQT feature is used instead of the fundamental frequency to characterize the melody of the audio to be converted, and there is no need to extract the fundamental frequency separately. Compared with the extraction of the fundamental frequency, the extraction of the CQT feature is more accurate and less prone to errors, which can avoid the problem of poor singing voice conversion effect caused by errors in the extraction of the fundamental frequency. In addition, considering that the CQT feature contains the timbre information of the song, in order to avoid the timbre information in the CQT feature from interfering with the audio conversion, the CQT feature encoder used in the embodiment of the present application is jointly trained with the MIDI (Musical Instrument Digital Interface) encoder, so that the encoded CQT feature output by the CQT feature encoder is close to the encoded MIDI feature output by the MIDI encoder. The MIDI feature does not contain timbre information, but it is difficult to obtain and is not suitable for direct use in audio conversion. Through this method, the encoded CQT feature can be made close to the MIDI feature, and the timbre information contained therein before encoding can be eliminated as much as possible.
[0075] The following describes the audio processing method provided by the embodiment of the present application in conjunction with the accompanying drawings. The method can be implemented by a terminal. Figure 1 The method may include the following steps:
[0076] Step 101: Obtain target timbre audio and extract timbre features of the target timbre audio.
[0077] The timbre feature is used to characterize the timbre of the target timbre audio.
[0078] In implementation, if the user wants to convert the timbre of the audio to be converted into a target timbre, the user can obtain the target timbre audio through the terminal, and the timbre of the target timbre audio is the target timbre.
[0079] For example, the terminal may be installed with a music application, which may include a tuning function option. The user may select the tuning function option to enter the tuning function interface. In the tuning function interface, the user may select a target timbre audio, which may be an audio recorded by the user and stored in the terminal. The terminal may then read the target timbre audio from the local storage space.
[0080] In some implementations, the terminal may also obtain the target timbre audio through the Internet, or automatically generate the target timbre audio. The embodiment of the present application does not limit the method for obtaining the target timbre audio.
[0081] After acquiring the target timbre audio, the terminal may extract the target timbre features of the target timbre audio, wherein the target timbre features may be in the form of a vector. The target timbre features may be extracted using MFCC (Mel-Frequency Cepstral Coefficients) methods, deep learning models, etc. The present embodiment of the application does not limit the method used for timbre feature extraction.
[0082] Step 102: Acquire the audio to be converted, and extract the CQT features and content features of the audio to be converted.
[0083] The CQT feature is used to characterize the melody of the audio to be converted, and the content feature is used to characterize the content of the audio to be converted.
[0084] In practice, the user can obtain the audio to be converted through the terminal. For example, in the above-mentioned audio modification function interface, the user can select the audio to be converted, and then the terminal can obtain the audio to be converted locally or through the Internet.
[0085] After obtaining the audio to be converted, the audio to be converted is input into a content feature extraction model to obtain the content features of the audio to be converted. The content feature extraction model may be Whisper, Paraformer, etc.
[0086] In addition, the CQT algorithm is used to extract the CQT features of the audio to be converted.
[0087] Step 103: Input the CQT feature into the CQT feature encoder to obtain the CQT encoding feature.
[0088] In implementation, Figure 2 As shown in the figure, after obtaining the CQT features, the CQT features can be input into a pre-trained CQT feature encoder, which then outputs the CQT encoded features. The CQT feature encoder can use a transformer model or a conformer model. The transformer model is a deep learning model based on the self-attention mechanism, while the conformer model is a hybrid deep learning model that combines the transformer model with the CNN (Convolutional Neural Network) model.
[0089] Step 104: Input the timbre features, CQT coding features, and content features into a decoder to obtain converted audio.
[0090] In implementation, after obtaining the timbre features of the target audio, the CQT coding features of the audio to be converted, and the content features, these features are input into a pre-trained decoder. The decoder outputs a converted spectrum, which is then input into a vocoder, which outputs the converted audio. The spectrum can be a mel-spectrogram.
[0091] In one possible implementation, Figure 2 As shown, the timbre features of the target timbre audio, the CQT coding features and the content features of the audio to be converted can be mapped to the same dimension through a linear layer and then fused. Then, the fused features are input into a pre-trained decoder.
[0092] In one possible implementation, Figure 2 As shown, the method provided in the embodiment of the present application can also support the pitch shifting function (pitch shifting is also called pitch shifting). Accordingly, the terminal can obtain the pitch shifting value, which is used to indicate the target number of keys (tones) offset based on the pitch of the audio to be converted. On this basis, the processing of the above step 103 can also be as follows:
[0093] Obtaining the pitch value, performing pitch processing on the CQT feature based on the pitch value, and inputting the pitch-processed CQT feature into the CQT feature encoder. Specifically, the CQT feature is in matrix form, i.e., the CQT feature is a CQT feature matrix. The CQT feature matrix can be shifted according to the pitch value to implement pitch processing on the CQT feature.
[0094] There are many methods for obtaining the above-mentioned ascending and descending pitch values. Two of them are exemplified below for illustration:
[0095] Method 1:
[0096] In the pitch-correction interface of a music application, there may be a pitch-up / down value input box. The user can enter a pitch-up / down value in the pitch-up / down value input box to indicate the key to be offset. The terminal can then receive the pitch-up / down value entered by the user. For example, entering "+1" means raising the pitch of the audio to be converted by one key. For another example, entering "-1" means lowering the pitch of the audio to be converted by one key.
[0097] Method 2:
[0098] The terminal calculates the pitch value based on the MIDI average of the target timbre audio and the MIDI average of the audio to be converted.
[0099] Specifically, the terminal calculates the mean fundamental frequency of the target timbre audio and converts the mean fundamental frequency into a MIDI value to obtain the MIDI mean of the target timbre audio. The terminal also calculates the mean fundamental frequency of the audio to be converted and converts the mean fundamental frequency into a MIDI value to obtain the MIDI mean of the audio to be converted. Then, the terminal subtracts the MIDI mean of the audio to be converted from the MIDI mean of the target timbre audio to obtain the pitch value.
[0100] An embodiment of the present application provides a method for audio processing, in which the CQT features and content features of the audio to be converted are extracted, and the timbre features of the target timbre audio are extracted. Then, by combining the CQT features, content features, and timbre features, a converted audio having the target timbre, the content of the audio to be converted, and the melody of the audio to be converted is generated. It can be seen that in this method, the CQT features are used instead of the fundamental frequency to characterize the melody of the audio to be converted, eliminating the need to extract the fundamental frequency separately. The extraction of CQT features is more accurate and less prone to errors than the extraction of the fundamental frequency, and can avoid the problem of poor singing voice conversion due to errors in fundamental frequency extraction.
[0101] Since the CQT feature contains the timbre information of the audio to be converted, it may interfere with the timbre characteristics of the target timbre audio during decoding. Based on this, an embodiment of the present application also provides a training method for a singing conversion model. In this method, the CQT feature encoder is jointly trained with the MIDI encoder, so that the CQT coding feature output by the CQT feature encoder is close to the MIDI coding feature output by the MIDI encoder. The MIDI coding feature does not contain timbre information. In this way, the CQT coding feature is close to the MIDI coding feature, and the timbre information contained therein before encoding can be eliminated as much as possible.
[0102] The following describes the training of the singing voice conversion model with reference to the accompanying drawings. The singing voice conversion model may include the above-mentioned CQT feature encoder and decoder. The training may be implemented by a computing device, which may be a server or a server cluster. Figure 3 and Figure 4 The training method of the singing voice conversion model provided in the embodiment of the present application may include the following steps:
[0103] Step 301: Obtain a first audio sample and a second audio sample.
[0104] The first and second audio samples have the same timbre. The first audio sample is used to extract the timbre, while the second audio sample is used to extract the melody and content. The timbre of the first audio sample, the melody, and content of the second audio sample are fused together using a singing voice conversion model to generate a converted audio. The converted audio is the predicted value of the singing voice conversion model, and its corresponding true value should be an audio with the timbre of the first audio sample and the melody and content of the second audio sample, that is, the second audio sample.
[0105] In implementation, the computing device may first obtain the target song audio, wherein the target song audio may be a solo singer's song audio obtained through a music library, the Internet, recording, or other means.
[0106] Then, a continuous audio of the first duration is intercepted from the target song audio, wherein the first duration can be configured by relevant technical personnel according to actual needs, for example, the first duration is 40 seconds. Then, an audio of the second duration is obtained from the audio of the first duration as the first sample audio, and another audio of the second duration is obtained from the audio of the first duration as the second sample audio, and the first sample audio and the second sample audio are overlapped. The second duration can be configured according to actual needs. For example, if the first duration is 40 seconds, the second duration needs to be no more than 20 seconds, such as the second duration is 20 seconds.
[0107] Step 302: Extract target timbre features of the first audio sample.
[0108] The target timbre feature is used to characterize the timbre of the first sample audio.
[0109] In implementation, the target timbre features can be extracted using the MFCC (Mel-Frequency Cepstral Coefficients) method, a deep learning model, etc. The embodiment of the present application does not limit the method used for timbre feature extraction, and it only needs to be the same as the method used in the above step 101.
[0110] Step 303: Extract target MIDI features, target CQT features, and target content features of the second sample audio.
[0111] The target MIDI feature and the target CQT feature are used to characterize the melody of the second sample audio, and the target content feature is used to characterize the content of the second sample audio.
[0112] In implementation, after acquiring the second audio sample, the second audio sample is input into a content feature extraction model to obtain target content features for the second audio sample. The content feature extraction model may be Whisper, Paraformer, or similar. The target CQT features of the second audio sample are extracted using a CQT algorithm. The second audio sample is then input into a MIDI feature prediction model to obtain target MIDI features for the second audio sample. The MIDI prediction model is a neural network model.
[0113] Step 304: Input the target MIDI feature into the MIDI encoder to obtain the target MIDI coding feature, and input the target CQT feature into the CQT feature encoder to obtain the target CQT coding feature.
[0114] The CQT feature encoder can use a transformer model, a conformer model, or other models. The transformer model is a deep learning model based on the self-attention mechanism, while the conformer model is a hybrid deep learning model that integrates the transformer model with the CNN model. Similarly, the MIDI encoder can also use a transformer model or a conformer model.
[0115] In implementation, after obtaining the target MIDI feature, the MIDI feature is input into a MIDI encoder, which outputs the MIDI encoding feature. After obtaining the target CQT feature, the target CQT feature is input into a CQT feature encoder, which outputs the target CQT encoding feature.
[0116] Step 305: Calculate a first loss value between the target MIDI coding feature and the target CQT coding feature.
[0117] The first loss value may be an MSE (mean-square error) loss value.
[0118] In implementation, after obtaining the target MIDI coding features and the target CQT coding features, the MSE loss value between the target MIDI coding features and the target CQT coding features is calculated, and the CQT feature encoder and the MIDI encoder are trained according to the MSE loss value, and the model parameters of the two are adjusted to improve the similarity between the CQT coding features output by the CQT feature encoder and the MIDI coding features output by the MIDI encoder. In this way, the CQT coding features can have the advantages of MIDI coding features at the same time, while maintaining the ability to capture information such as melody and harmony, reducing timbre leakage.
[0119] Step 306: Select the target MIDI coding feature or the target coding CQT feature as the melody feature, and input the melody feature, the target content feature and the target timbre feature into the decoder to obtain the converted spectrum feature.
[0120] In implementation, after obtaining the target MIDI coding feature and the target CQT coding feature, a random selector can be used to randomly select one of the target MIDI coding feature and the target CQT coding feature as a melody feature, and the selected melody feature, the above-mentioned target content feature and the above-mentioned target timbre feature are input into the decoder to obtain the converted spectral feature, wherein the converted spectral feature can be a Mel spectrum.
[0121] In a possible implementation, the melody features, target content features, and target timbre features may be mapped to the same dimension through a linear layer, and then feature fusion may be performed. The fused features may then be input into a decoder.
[0122] Step 307: Extract the target spectral features of the second sample audio, calculate a second loss value between the converted spectral features and the target spectral features, train the decoder based on the second loss value, and train the CQT feature encoder and the MIDI encoder based on the first loss value.
[0123] In implementation, for the second sample audio obtained in the above step 301, the target spectrum features of the second sample audio are extracted, wherein the target spectrum features and the above converted spectrum features need to be features extracted using the same algorithm, for example, the target spectrum features and the above converted spectrum features are both Mel spectrum.
[0124] After obtaining the target spectral features and converted spectral features of the second sample audio, a second loss value of the target spectral features and the converted spectral features is calculated, where the second loss value can be an L1 loss, which is also called a MAE (Mean Absolute Error) loss. Furthermore, a gradient descent method is used to train the CQT feature encoder, MIDI encoder, and decoder based on the second loss value, and adjust the model parameters of the three. Then, a gradient descent method is used again to additionally train the CQT feature encoder and MIDI encoder based on the first loss value.
[0125] In addition, the above process is only one round of training. In actual implementation, a large amount of sample audio can be collected and multiple rounds of training can be performed until the training stop condition is met, and then the training is stopped to obtain the trained CQT feature encoder and decoder.
[0126] In the training method provided in the embodiment of the present application, the CQT feature encoder is jointly trained with the MIDI encoder, so that the CQT coding features output by the CQT feature encoder are close to the MIDI coding features output by the MIDI encoder. The MIDI coding features do not contain timbre information. Through this method, the CQT coding features can be made close to the MIDI coding features to eliminate the timbre information contained therein before encoding as much as possible.
[0127] All of the above optional technical solutions can be combined in any way to form optional embodiments of the present application, and will not be described in detail here.
[0128] Based on the same technical concept, the embodiment of the present application also provides an audio processing device, which can be applied to a computing device, such as Figure 5 As shown, the apparatus includes an acquisition module 510 and a conversion module 520, wherein:
[0129] The acquisition module 510 is configured to acquire target timbre audio and extract timbre features of the target timbre audio, wherein the timbre features are used to characterize the timbre of the target timbre audio; acquire audio to be converted and extract CQT features and content features of the audio to be converted, wherein the CQT features are used to characterize the melody of the audio to be converted and the content features are used to characterize the content of the audio to be converted;
[0130] The conversion module 520 is configured to input the CQT feature into a CQT feature encoder to obtain a CQT encoding feature, wherein the CQT feature encoder is a CQT feature encoder trained according to any one of claims 1-5; and input the timbre feature, the CQT encoding feature, and the content feature into a decoder to obtain converted audio, wherein the decoder is a decoder trained according to any one of claims 1-5.
[0131] In a possible implementation, the acquisition module 510 is further configured to:
[0132] Get the rising and falling pitch values;
[0133] Based on the up-down adjustment value, performing up-down adjustment processing on the CQT feature;
[0134] The conversion module 520 is configured to:
[0135] The CQT features after up- and down-tuning processing are input into the CQT feature encoder.
[0136] In the solution provided in the embodiments of the present application, the CQT features and content features of the audio to be converted are extracted, and the timbre features of the target timbre audio are extracted. Then, by combining the CQT features, content features, and timbre features, a converted audio is generated that has the target timbre, the content of the audio to be converted, and the melody of the audio to be converted. As can be seen, in this method, the CQT features are used instead of the fundamental frequency to characterize the melody of the audio to be converted, eliminating the need to extract the fundamental frequency separately. Compared to extracting the fundamental frequency, extracting the CQT features is more accurate and less prone to errors, thus avoiding the problem of poor singing conversion effects caused by errors in fundamental frequency extraction.
[0137] It should be noted that the audio processing device provided in the above embodiment only uses the division of the above functional modules as an example to illustrate audio processing. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computing device can be divided into different functional modules to complete all or part of the functions described above. In addition, the audio processing device provided in the above embodiment and the method embodiment of touch audio processing are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0138] Based on the same technical concept, the embodiment of the present application also provides a training device for a singing voice conversion model, which can be applied to computing devices such as Figure 6 As shown, the apparatus includes an acquisition module 710 and a training module 720, wherein:
[0139] An acquisition module 710 is configured to acquire a first audio sample and a second audio sample, wherein the first audio sample and the second audio sample have the same timbre; extract a target timbre feature of the first audio sample, wherein the target timbre feature is used to characterize the timbre of the first audio sample; and extract a target Musical Instrument Digital Interface (MIDI) feature, a target CQT feature, and a target content feature of the second audio sample, wherein the target MIDI feature and the target CQT feature are used to characterize the melody of the second audio sample, and the target content feature is used to characterize the content of the second audio sample.
[0140] The training module 720 is used to input the target MIDI feature into the MIDI encoder to obtain the target MIDI coding feature, input the CQT feature into the CQT feature encoder to obtain the target CQT encoding feature; calculate the first loss value between the target MIDI coding feature and the target CQT coding feature; select the target MIDI coding feature or the target coding CQT feature as the melody feature, and input the melody feature, the target content feature and the target timbre feature into the decoder to obtain the converted spectrum feature; extract the target spectrum feature of the second sample audio, calculate the second loss value between the converted spectrum feature and the target spectrum feature, train the decoder based on the second loss value, and train the CQT feature encoder and the MIDI encoder based on the first loss value.
[0141] In a possible implementation, the first loss value is a mean square error loss value.
[0142] In a possible implementation, the training module 720 is configured to:
[0143] The melody feature, the target content feature and the target timbre feature are subjected to feature fusion to obtain a target fusion feature, and the target fusion feature is input into the decoder.
[0144] In a possible implementation, the training module 720 is configured to:
[0145] The target MIDI feature or the target CQT feature is randomly selected as a melody feature.
[0146] In a possible implementation, the obtaining module 710 is configured to:
[0147] Get the target song audio;
[0148] Cutting out a continuous audio of a first duration from the target song audio;
[0149] The audio of the second duration before the audio of the first duration is used as the first sample audio;
[0150] The audio of the second duration after the audio of the first duration is used as the second sample audio.
[0151] In the technical solution provided in the embodiment of the present application, the CQT feature encoder is jointly trained with the MIDI encoder, so that the CQT coding features output by the CQT feature encoder are close to the MIDI coding features output by the MIDI encoder. The MIDI coding features do not contain timbre information. Through this method, the CQT coding features can be made close to the MIDI coding features to eliminate the timbre information contained therein before encoding as much as possible.
[0152] It should be noted that the training device for the singing voice conversion model provided in the above embodiment only uses the division of the above functional modules as an example to illustrate the training of the singing voice conversion model. In actual applications, the above functions can be distributed and completed by different functional modules as needed, that is, the internal structure of the computing device can be divided into different functional modules to complete all or part of the functions described above. In addition, the training device for the singing voice conversion model provided in the above embodiment and the training method embodiment of the singing voice conversion model are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0153] Figure 7 The following is a block diagram of an electronic device 800 according to an exemplary embodiment of the present application. The electronic device 800 may be a portable mobile terminal, such as a smartphone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer, or a desktop computer. The electronic device 800 may also be referred to as a user device, a portable terminal, a laptop terminal, a desktop terminal, or other similar names.
[0154] Typically, the electronic device 800 includes a processor 801 and a memory 802 .
[0155] The processor 801 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 801 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 801 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 801 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 801 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.
[0156] The memory 802 may include one or more computer-readable storage media, which may be non-transitory. The memory 802 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 802 is used to store at least one instruction, which is executed by the processor 801 to implement the audio processing method provided in the method embodiment of the present application.
[0157] In some embodiments, electronic device 800 may optionally include a peripheral device interface 803 and at least one peripheral device. The processor 801, memory 802, and peripheral device interface 803 may be connected via a bus or signal lines. Each peripheral device may be connected to peripheral device interface 803 via a bus, signal lines, or circuit boards. Specifically, the peripheral device may include at least one of a radio frequency circuit 804, a display screen 805, a camera assembly 806, an audio circuit 807, a positioning assembly 808, and a power supply 809.
[0158] The peripheral device interface 803 can be used to connect at least one I / O (Input / Output)-related peripheral device to the processor 801 and the memory 802. In some embodiments, the processor 801, the memory 802, and the peripheral device interface 803 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 801, the memory 802, and the peripheral device interface 803 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.
[0159] The radio frequency circuit 804 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 804 communicates with communication networks and other communication devices via electromagnetic signals. The radio frequency circuit 804 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. Optionally, the radio frequency circuit 804 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The radio frequency circuit 804 can communicate with other terminals via at least one wireless communication protocol. Such wireless communication protocols include, but are not limited to, the World Wide Web, a metropolitan area network, an intranet, various generations of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 804 may also include circuits related to NFC (Near Field Communication), which is not limited in this application.
[0160] The display screen 805 is used to display a user interface (UI). This UI may include graphics, text, icons, videos, or any combination thereof. When the display screen 805 is a touch screen, it is also capable of collecting touch signals on or above the surface of the display screen 805. These touch signals can be input as control signals to the processor 801 for processing. In this case, the display screen 805 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there can be one display screen 805, located on the front panel of the electronic device 800. In other embodiments, there can be at least two display screens 805, located on different surfaces of the electronic device 800 or in a foldable design. In other embodiments, the display screen 805 can be a flexible display, located on a curved or foldable surface of the electronic device 800. Furthermore, the display screen 805 can be configured as a non-rectangular irregular shape, i.e., a special-shaped screen. The display screen 805 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0161] The camera assembly 806 is used to capture images or videos. Optionally, the camera assembly 806 includes a front camera and a rear camera. Typically, the front camera is arranged on the front panel of the terminal, and the rear camera is arranged on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth of field camera, a wide-angle camera, and a telephoto camera, so as to realize the fusion of the main camera and the depth of field camera to realize the background blur function, the fusion of the main camera and the wide-angle camera to realize panoramic shooting and VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera assembly 806 may also include a flash. The flash can be a monochrome temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation at different color temperatures.
[0162] The audio circuit 807 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, and convert the sound waves into electrical signals to be input into the processor 801 for processing, or input into the radio frequency circuit 804 to achieve voice communication. For the purpose of stereo sound collection or noise reduction, there can be multiple microphones, which are respectively arranged in different parts of the electronic device 800. The microphone can also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signals from the processor 801 or the radio frequency circuit 804 into sound waves. The speaker can be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signals into sound waves audible to humans, but also convert the electrical signals into sound waves inaudible to humans for purposes such as ranging. In some embodiments, the audio circuit 807 may also include a headphone jack.
[0163] Positioning component 808 is used to locate the current geographic location of electronic device 800 to implement navigation or LBS (Location Based Service). Positioning component 808 can be a positioning component based on the US GPS (Global Positioning System), China's Beidou system, or Russia's Galileo system.
[0164] Power supply 809 is used to power the various components of electronic device 800. Power supply 809 can be AC power, DC power, a disposable battery, or a rechargeable battery. When power supply 809 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, while a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.
[0165] In some embodiments, the electronic device 800 further includes one or more sensors 810 , including but not limited to: an acceleration sensor 811 , a gyroscope sensor 812 , a pressure sensor 813 , a fingerprint sensor 814 , an optical sensor 815 , and a proximity sensor 816 .
[0166] The accelerometer 811 can detect the magnitude of acceleration along the three coordinate axes of the coordinate system established by the electronic device 800. For example, the accelerometer 811 can be used to detect the components of gravity acceleration along the three coordinate axes. The processor 801 can control the display screen 805 to display the user interface in a landscape or portrait view based on the gravity acceleration signal collected by the accelerometer 811. The accelerometer 811 can also be used to collect game or user motion data.
[0167] The gyroscope sensor 812 can detect the orientation and rotation angle of the electronic device 800. It can work in conjunction with the accelerometer 811 to capture the user's 3D movements of the electronic device 800. Based on the data collected by the gyroscope sensor 812, the processor 801 can implement the following functions: motion sensing (for example, changing the UI based on the user's tilt operation), image stabilization during shooting, game control, and inertial navigation.
[0168] The pressure sensor 813 can be set on the side frame of the electronic device 800 and / or the lower layer of the display screen 805. When the pressure sensor 813 is set on the side frame of the electronic device 800, it can detect the user's grip signal of the electronic device 800, and the processor 801 performs left and right hand recognition or shortcut operations based on the grip signal collected by the pressure sensor 813. When the pressure sensor 813 is set on the lower layer of the display screen 805, the processor 801 controls the operable controls on the UI interface based on the user's pressure operation on the display screen 805. The operable controls include at least one of a button control, a scroll bar control, an icon control, and a menu control.
[0169] The fingerprint sensor 814 is used to collect the user's fingerprint. The processor 801 identifies the user's identity based on the fingerprint collected by the fingerprint sensor 814, or the fingerprint sensor 814 identifies the user's identity based on the collected fingerprint. When the user's identity is identified as a trusted identity, the processor 801 authorizes the user to perform relevant sensitive operations, such as unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings. The fingerprint sensor 814 can be set on the front, back, or side of the electronic device 800. When a physical button or manufacturer logo is set on the electronic device 800, the fingerprint sensor 814 can be integrated with the physical button or manufacturer logo.
[0170] The optical sensor 815 is used to detect ambient light intensity. In one embodiment, the processor 801 can control the display brightness of the display screen 805 based on the ambient light intensity detected by the optical sensor 815. Specifically, when the ambient light intensity is high, the display brightness of the display screen 805 is increased; when the ambient light intensity is low, the display brightness of the display screen 805 is decreased. In another embodiment, the processor 801 can also dynamically adjust the shooting parameters of the camera assembly 806 based on the ambient light intensity detected by the optical sensor 815.
[0171] Proximity sensor 816, also known as a distance sensor, is typically located on the front panel of electronic device 800. Proximity sensor 816 is used to detect the distance between the user and the front of electronic device 800. In one embodiment, when proximity sensor 816 detects that the distance between the user and the front of electronic device 800 is gradually decreasing, processor 801 controls display screen 805 to switch from the screen-on state to the screen-off state. When proximity sensor 816 detects that the distance between the user and the front of electronic device 800 is gradually increasing, processor 801 controls display screen 805 to switch from the screen-off state to the screen-on state.
[0172] Those skilled in the art will understand that Figure 7 The structure shown in the figure does not constitute a limitation on the electronic device 800, and the electronic device 800 may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.
[0173] Figure 8 1 is a schematic diagram of the structure of a server provided in an embodiment of the present application. The server 1000 may vary significantly due to different configurations or performance, and may include one or more processors (central processing units, CPUs) 1001 and one or more memories 1002. The memories 1002 store at least one instruction, which is loaded and executed by the processor 1001 to implement the methods provided in the above-mentioned various method embodiments. Of course, the server may also have components such as a wired or wireless network interface, a keyboard, and input / output interfaces for input and output. The server may also include other components for implementing device functions, which are not described in detail here.
[0174] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including instructions, which can be executed by a processor in a terminal to complete the audio processing method or the training method of the singing voice conversion model in the above embodiment. The computer-readable storage medium can be non-transitory. For example, the computer-readable storage medium can be a ROM (Read-Only Memory), a RAM (Random Access Memory), a CD-ROM (Compact Disc Read-Only Memory), a magnetic tape, a floppy disk, an optical data storage device, etc.
[0175] It should be noted that the information involved in this application (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals (including but not limited to signals transmitted between user terminals and other devices, etc.) are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions. For example, all audio (such as the first sample audio, the second sample audio, the target timbre audio, the audio to be converted, etc.) and pitch values involved in this application are obtained with full authorization.
[0176] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.
[0177] It should be understood that the "multiple" mentioned in this article refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent three situations: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the previous and subsequent associated objects are in an "or" relationship. In addition, the step numbers described in this article only illustrate a possible execution sequence between the steps. In some other embodiments, the above steps may not be executed in the order of the numbers, such as two steps with different numbers are executed at the same time, or two steps with different numbers are executed in the opposite order to the diagram. The embodiments of the present application do not limit this.
[0178] The above description is merely an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A training method for a singing voice conversion model, characterized in that: The singing voice conversion model includes a constant Q transform feature encoder and a decoder, and the method includes: Acquire a first audio sample and a second audio sample, wherein the first audio sample and the second audio sample have the same timbre; Extracting a target timbre feature of the first audio sample, wherein the target timbre feature is used to characterize the timbre of the first audio sample; extracting a target Musical Instrument Digital Interface feature, a target Constant Q transform feature, and a target content feature of the second audio sample, wherein the target Musical Instrument Digital Interface feature and the target Constant Q transform feature are used to characterize the melody of the second audio sample, and the target content feature is used to characterize the content of the second audio sample; Inputting the target Musical Instrument Digital Interface feature into a Musical Instrument Digital Interface encoder to obtain a target Musical Instrument Digital Interface coding feature, and inputting the constant Q transform feature into a constant Q transform feature encoder to obtain a target constant Q transform coding feature; calculating a first loss value between the target Musical Instrument Digital Interface code feature and the target Constant Q Transform code feature based on the first loss value; selecting the target Musical Instrument Digital Interface coding feature or the target coding constant Q transform feature as a melody feature, and inputting the melody feature, the target content feature, and the target timbre feature into the decoder to obtain a converted spectrum feature; Extract target spectral features of the second sample audio, calculate a second loss value between the converted spectral features and the target spectral features, train the decoder based on the second loss value, and train the constant Q transform feature encoder and the Musical Instrument Digital Interface encoder based on the first loss value and the second loss value.
2. The method according to claim 1, characterized in that The first loss value is a mean square error loss value.
3. The method according to claim 1, characterized in that The step of inputting the melody feature, the target content feature, and the target timbre feature into the decoder comprises: The melody feature, the target content feature and the target timbre feature are subjected to feature fusion to obtain a target fusion feature, and the target fusion feature is input into the decoder.
4. The method according to claim 1, wherein The selecting the target MIDI feature or the target constant-Q transform feature as the melody feature includes: The target MIDI feature or the target constant-Q transform feature is randomly selected as a melody feature.
5. The method according to any one of claims 1 to 4, characterized in that The obtaining of the first audio sample and the second audio sample includes: Get the target song audio; Cutting out a continuous audio of a first duration from the target song audio; The audio of the second duration before the audio of the first duration is used as the first sample audio; The audio of the second duration after the audio of the first duration is used as the second sample audio.
6. A method for audio processing, characterized in that: The method comprises: Acquire target timbre audio, and extract timbre features of the target timbre audio, wherein the timbre features are used to characterize the timbre of the target timbre audio; Acquiring audio to be converted, and extracting a constant Q transform feature and a content feature of the audio to be converted, wherein the constant Q transform feature is used to characterize the melody of the audio to be converted, and the content feature is used to characterize the content of the audio to be converted; Inputting the constant Q transform feature into a constant Q transform feature encoder to obtain a constant Q transform coded feature, wherein the constant Q transform feature encoder is a constant Q transform feature encoder trained according to any one of claims 1 to 5; The timbre feature, the constant Q transform coding feature and the content feature are input into a decoder to obtain converted audio, wherein the decoder is a decoder trained according to any one of claims 1-5.
7. The method according to claim 6, characterized in that The method further comprises: Get the pitch value; Based on the up-down tuning value, performing up-down tuning processing on the constant Q transform feature; The step of inputting the constant Q transform feature into a constant Q transform feature encoder comprises: The constant Q transform feature after up- and down-scaling is input into the constant Q transform feature encoder.
8. A computing device, characterized in that The computing device includes a processor and a memory, wherein the memory stores at least one instruction, and the instruction is loaded and executed by the processor to implement the operations performed by the training method of the singing voice conversion model as described in any one of claims 1-5, or the operations performed by the audio processing method as described in any one of claims 6-7.
9. A computer-readable storage medium, characterized in that The storage medium stores at least one instruction, which is loaded and executed by the processor to implement the operations performed by the training method for a singing voice conversion model as described in any one of claims 1-5, or the operations performed by the audio processing method as described in any one of claims 6-7.
10. A computer program product, characterized in that The computer program product stores at least one instruction, which is loaded and executed by the processor to implement the operations performed by the training method for a singing voice conversion model as described in any one of claims 1-5, or the operations performed by the audio processing method as described in any one of claims 6-7.