Method and apparatus for mixing audio, and device, medium and program product
Patent Information
- Application Number
- PCT/CN2024/135316
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-08
- Filing Date
- 2024-11-28
- Publication Date
- 2025-10-02
AI Technical Summary
Existing technologies have difficulty capturing complex audio signal characteristics during the audio mixing process, resulting in limited mixing effects. They also require professional knowledge and experience and are difficult to adapt to changing music styles and creative needs.
By obtaining the features of target vocals and background music from multiple audio tracks, the target mixed audio is generated using normalization processing, including equalization normalization, dynamic range control, audio image normalization, and loudness normalization, combined with a neural network model for audio mixing.
It improves the automation and adaptability of audio mixing, realizes more intelligent and adaptive audio processing, and enhances the user experience.
Smart Images

Figure CN2024135316_02102025_PF_FP_ABST
Abstract
Description
Method, apparatus, device, medium, and program product for mixing audio
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to Chinese patent application No. 202410268781.6, filed on March 8, 2024, with application number 202410268781.6 and invention name “Method, device, apparatus, medium and program product for mixing audio”. The entire contents of this application are incorporated by reference into this application. Technical Field
[0003] Embodiments of the present disclosure generally relate to the field of audio processing, and more particularly to methods, devices, apparatuses, media, and program products for mixing audio. Background Art
[0004] At present, with the development of deep neural networks, deep learning is receiving more and more attention, and it involves many application scenarios in music creation, such as lyrics and music writing, accompaniment arrangement generation, song synthesis, automatic tuning, etc.
[0005] The traditional manual mastering process requires extensive experience and a keen auditory sense of sound quality and spatial perception, as well as patience and meticulous adjustments. With the rapid development of deep neural networks and deep learning, and the widespread application of various deep learning architectures, the use of machine learning models in automated audio mixing is becoming increasingly common. However, many challenges remain in the mixing process. Summary of the Invention
[0006] Embodiments of the present disclosure provide a method, apparatus, device, medium, and program product for mixing audio.
[0007] According to a first aspect of the present disclosure, a method for mixing audio is provided. The method includes obtaining a target vocal for a first track among a plurality of audio tracks and a target background music for a second track among the plurality of audio tracks. The method also includes determining a first set of sound features for a first group of audio tracks related to the vocals and a second set of sound features for a second group of audio tracks related to the background music. The method also includes normalizing the target vocal and target background music based on the first set of sound features and the second set of sound features. The method also includes generating a target mixed audio based on the processed target vocal and the processed target background music.
[0008] In a second aspect of the present disclosure, a device for mixing audio is provided. The device includes a track acquisition module configured to acquire a target vocal for a first track among a plurality of tracks and a target background music for a second track among the plurality of tracks; a feature determination module configured to determine a first set of sound features for a first group of tracks related to vocals and a second set of sound features for a second group of tracks related to background music; a normalization module configured to normalize the target vocal and target background music based on the first set of sound features and the second set of sound features; and a mixed audio generation module configured to generate a target mixed audio based on the processed target vocal and the processed target background music.
[0009] In a third aspect of the present disclosure, an electronic device is provided, comprising at least one processor; and a storage device for storing at least one program, wherein when the at least one program is executed by the at least one processor, the at least one processor implements the method according to the first aspect of the present disclosure.
[0010] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the method according to the first aspect of the present disclosure is implemented.
[0011] In a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program, which implements the method according to the first aspect of the present disclosure when executed by a processor.
[0012] It should be understood that the content described in this content section is not intended to limit the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easy to understand through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The above and other objects, features and advantages of the present disclosure will become more apparent through a more detailed description of exemplary embodiments of the present disclosure with reference to the accompanying drawings, wherein like reference numerals generally represent like components throughout the exemplary embodiments of the present disclosure.
[0014] FIG1 illustrates a schematic diagram of an example environment in which devices and / or methods according to embodiments of the present disclosure may be implemented;
[0015] FIG2 illustrates a schematic diagram of an example of a flowchart for mixing audio according to an embodiment of the present disclosure;
[0016] FIG3 illustrates a flow chart of an example method for mixing audio according to an embodiment of the present disclosure;
[0017] FIG4 illustrates a schematic diagram of an example of a mixing model architecture for mixing audio according to an embodiment of the present disclosure;
[0018] FIG5 illustrates a schematic block diagram of an apparatus for mixing audio according to an embodiment of the present disclosure;
[0019] FIG6 illustrates a schematic block diagram of an example device suitable for implementing embodiments of the present disclosure.
[0020] In the various drawings, the same or corresponding reference numerals denote the same or corresponding parts. DETAILED DESCRIPTION
[0021] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions. In response to receiving the user's active request, a prompt message is sent to the user to clearly remind the user that the operation requested to be performed will require the acquisition and use of the user's personal information. Thus, the user can independently choose whether to provide personal information to the electronic device, application, server or storage medium and other software or hardware that performs the operation of the technical solution of this disclosure based on the prompt message.
[0022] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0023] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to." The term "based on" should be understood as "based at least in part on." The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment." The terms "first," "second," etc. may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0024] As mentioned above, many challenges remain to be addressed in the audio mixing process. For example, deep neural networks and deep learning still face numerous technical challenges in analyzing and applying information such as vocals, background music, instrumental accompaniment, and finished music. To achieve more accurate and efficient audio mixing, traditional rule-based audio mixing methods have been proposed. For example, rule-based automatic mixing algorithms typically require mixing engineers or experts to design a rule-based mixing strategy and rule set, accompanied by a series of operations such as feature extraction, rule application, mixing decision-making using expert systems, and parameter adjustment. This process is time-consuming and labor-intensive, and rule-based algorithms are limited by pre-defined rule sets, making them difficult to adapt to complex and ever-changing musical styles and creative requirements. Furthermore, rule-based methods struggle to capture and express high-level features and complex relationships in audio signals, resulting in limited mixing results. Furthermore, they require professional mixing engineers to design and adjust the rule system, requiring a high level of expertise and experience. Furthermore, rule-based systems struggle to fully simulate the creativity and intuition of human mixing engineers and may, in some cases, fall short of achieving human-level mixing performance. Therefore, there are certain limitations when dealing with complex and changing musical situations and creative needs.
[0025] To at least address the above and other potential issues, embodiments of the present disclosure provide a method for audio mixing. In this method, a computing device may first obtain a target vocal for a first track among multiple audio tracks and a target background music for a second track among the multiple audio tracks. Furthermore, the computing device may obtain a first set of sound features for a first group of audio tracks related to the vocals and a second set of sound features for a second group of audio tracks related to the background music. The computing device then uses the determined first and second sets of sound features to normalize the target vocal and target background music. Finally, the computing device combines the processed target vocal and target background music to generate a target mixed audio. This method utilizes the target vocal for the first track among the multiple audio tracks and the target background music for the second track among the multiple audio tracks to obtain a first set of sound features for the first group of audio tracks related to the vocals and a second set of sound features for the second group of audio tracks related to the background music. Based on these features, the target vocal and target background music are normalized. This method better captures the complex characteristics of audio signals, improves automation, and provides greater adaptability, enabling more intelligent and adaptive audio mixing, thereby improving the user experience.
[0026] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. FIG1 illustrates an example environment in which the devices and / or methods of the embodiments of the present disclosure may be implemented. In environment 100 , a computing device 108 is configured to process a plurality of audio tracks 102 to determine a target audio mix 116 .
[0027] Examples of computing device 108 include, but are not limited to, personal computers, server computers, handheld or laptop devices, mobile devices (such as mobile phones, personal digital assistants (PDAs), media players, etc.), multi-processor systems, consumer electronics, minicomputers, mainframe computers, distributed computing environments that include any of the above systems or devices, and the like.
[0028] As shown in FIG1 , multiple audio tracks 102 are information input by a user into a computing device 108 for audio mixing. Each of the multiple audio tracks can correspond to collected, unprocessed audio. The multiple audio tracks include a first audio track 104 and a second audio track 106. The first audio track corresponds to the audio of the target vocals, and the second audio track corresponds to the audio of background music, such as the sound of a particular instrument or background music such as a popular song. In some embodiments, the multiple audio tracks may only include the vocals and background music. In some embodiments, the multiple audio tracks may include audio information of the vocals, background music, and other instruments. FIG1 shows that the computing device 108 receives the multiple audio tracks 102, for example, from other computing devices or remote storage devices in a network. This is merely an example and not a specific limitation of the present disclosure. The computing device 108 may also directly obtain the multiple audio tracks 102 from a local computer.
[0029] The computing device 108 also obtains a first set of sound features 110 corresponding to a first set of audio tracks 118 for human voices and a second set of sound features corresponding to a second set of audio tracks 120. The first set of sound features 110 or the second set of sound features 112 obtained by the computing device includes one or more of the amplitude spectrum mean, the amplitude mean of the transient value, the variance of the transient value, the average similarity measure for the sound image, or the loudness mean for the corresponding set of audio tracks.
[0030] In some embodiments, computing device 108 may directly obtain first set of sound features 110 and second set of sound features 112 that have been pre-processed from first set of audio tracks 118 and second set of audio tracks. This pre-processing may be performed by computing device 108 or by another computing device and then transmitted to computing device 108.
[0031] In some embodiments, the first set of tracks 118 and the second set of tracks 120 are derived from a track dataset obtained by separating a multi-track music source. The track dataset may be a track dataset obtained by separating a song and may include multiple sets of tracks of different types. The first set of tracks 118 is a subset of tracks for vocals, and the second set of tracks 120 is a subset of tracks for background music.
[0032] During the process that the computing device 108 can process the first group of audio tracks and the second group of audio tracks, the computing device 108 can calculate the amplitude spectrum mean, the amplitude mean of the transient value, the variance of the transient value, the average similarity measure for the sound image and / or the loudness mean of a group of audio tracks by processing each audio track in the first group of audio tracks or each audio track in the second group of audio tracks as the first group of sound features 110 corresponding to the first group of audio tracks 118 and the second group of sound features 112 corresponding to the second group of audio tracks 120.
[0033] The computing device 108 also includes a normalization module 114, which is used to obtain the first audio track 104 and the second audio track 106. In addition, the normalization module 114 can also obtain the first group of sound features 110 and the second group of sound features 112. The normalization module then uses the first group of sound features of the first group of audio tracks and the second group of sound features of the second group of audio tracks to normalize the first audio track 104 and the second audio track 106 to obtain the target mixed audio. For example, the computing device performs equalization normalization, dynamic range control normalization, sound image normalization, and loudness normalization on the first group of sound features and the second group of sound features. Finally, the computing device inputs the normalized target vocals or target background music into the mixing model for processing, and performs automatic mixing of the audio to obtain the target mixed audio 116.
[0034] Through this method, since the target human voice of the first audio track among the multiple audio tracks and the target background music of the second audio track among the multiple audio tracks are utilized, a first group of sound features of the first group of audio tracks related to the human voice and a second group of sound features of the second group of audio tracks related to the background music are obtained, and the target human voice and the target background music are normalized based on this, the complex characteristics of the audio signal can be better captured, the degree of automation is improved, the method has strong adaptability, and more intelligent and adaptive audio mixing processing can be achieved, thereby improving the user experience.
[0035] 1 , a schematic diagram of an example environment in which the apparatus and / or method of an embodiment of the present disclosure may be implemented is described above. Below, a schematic diagram of an example of a flowchart for mixing audio according to an embodiment of the present disclosure is described in conjunction with FIG. 2 .
[0036] As shown in Figure 2, in example 200, a computing device obtains a multi-track audio source separation dataset 202, wherein the multi-track audio source separation dataset 202 is a dataset of multiple tracks obtained by separating multiple processed finished mixed audio. In one example, after the computing device obtains the dataset of the separated multiple tracks, the sound features of the audio in the different types of tracks in the dataset are calculated. For example, the different types may include human voice and background music. The sound features of each type of track in the dataset of the separated multiple tracks can include at least one of the following: amplitude spectrum mean, transient value amplitude mean, transient value variance, average similarity measure for sound image, or loudness mean.
[0037] The example in Figure 2 includes two processes: one is the training process of the mixing model, and the other is the inference process of the mixing model. During the training process, the tracks in the multi-track sound source separation dataset 202 are used as sample tracks. After obtaining the wet sound tracks 204 in the multi-track sound source separation dataset 202, multiple tracks in the wet sound tracks 204 are synthesized to obtain sample mixed audio 208. Wet sound refers to sound that has undergone post-processing, particularly sound that has been adjusted in audio editing software, such as reverberation, delay, equalization, compression, and pitch. Each track in the wet sound tracks 204 is normalized using sound features corresponding to the track type, such as equalization normalization, dynamic range control normalization, sound image normalization, and loudness normalization. The normalized wet sound tracks are then obtained. Next, the normalized wet sound tracks are input into the mixing model to predict mixed audio 210. Then, a loss function is calculated for the predicted mixed audio 210 and the sample mixed audio 208, so as to further adjust the parameters of the mixing model. Finally, the training of the mixing model is achieved.
[0038] To further improve the accuracy of the mixing model, the loss function uses not only the amplitude spectrum of the audio signal but also the phase spectrum of the audio signal. In some embodiments, the computing device obtains the sample amplitude spectrum and sample phase spectrum of the channels of the sample mixed audio. The computing device also obtains the predicted amplitude spectrum and predicted phase spectrum of the channels of the mixed audio predicted by the mixing model, and then uses this information to calculate the loss function to adjust the parameters of the mixing model. In this way, when the loss function is used to adjust the mixing model, the mixing model can be trained faster and more accurately, and the mixed audio quality is higher.
[0039] During the inference process of the mixing model, the computing device obtains an unprocessed dry sound track 206. Dry sound refers to the original recording without any post-processing (such as reverberation, equalization, compression, noise elimination, etc.), which usually refers to the direct sound signal of human voice or musical instrument, that is, the sound directly captured and recorded by the microphone, retaining the most original and authentic sound quality characteristics. The computing device then uses the sound features corresponding to the types of tracks in the dry sound track 206 in the multi-track sound source separation data set 202 to normalize the dry sound track. For example, the dry sound human voice and dry sound background music are normalized. Further, the computing device inputs the processed dry sound human voice and dry sound background music into the trained mixing model to obtain mixed audio.
[0040] A schematic diagram of an example of a flowchart for audio mixing according to an embodiment of the present disclosure is described above in conjunction with Figure 2. A flowchart of an example method 300 for audio mixing according to an embodiment of the present disclosure is described below in conjunction with Figure 3.
[0041] At block 302, a target vocal track for a first track among multiple tracks and a target background music track for a second track among the multiple tracks are obtained. The computing device 108 may perform a mixing process on the vocal track and the background music track. During this process, the computing device 108 needs to obtain the vocal track and the background music track.
[0042] In some embodiments, when obtaining a target vocal for a first track and target background music for a second track among multiple tracks, the computing device may determine the dry sound of the vocal as the target vocal for the first track, and determine the dry sound of the background music as the target background music for the second track. In some embodiments, the computing device obtains a wet vocal and wet background music as the target vocal and target background music. The above examples are merely illustrative of the present disclosure and are not intended to limit the present disclosure.
[0043] At block 304, a first set of sound features is determined for a first set of audio tracks associated with vocals and a second set of sound features is determined for a second set of audio tracks associated with background music. For example, computing device 108 obtains a first set of sound features corresponding to first set of audio tracks 118 and a second set of sound features corresponding to second set of audio tracks 120. In some embodiments, the first set of audio tracks and the second set of audio tracks are from a multi-track audio source separation dataset. In some embodiments, the computing device may obtain the first set of sound features based on audio signals in the first set of audio tracks and the second set of sound features based on audio signals in the second set of audio tracks.
[0044] At block 306, the target vocals and the target background music are normalized based on the first set of sound features and the second set of sound features. For example, the computing device normalizes the target vocals using the first set of sound features and normalizes the target background music using the second set of sound features.
[0045] In some embodiments, the sound features of a set of audio tracks may include the mean amplitude spectrum, the mean amplitude of transient values, the variance of transient values, the average similarity measure for the sound image, or the mean loudness. The process of obtaining these features and the process of using these features for normalization are described in detail below.
[0046] In some embodiments, the computing device may determine the amplitude spectrum mean of a group of audio tracks of the same type using the following formulas (1) and (2):
[0047] Where k is the type of the track, for example, human voice or background music. For example, if the type of a track is human voice (k=0 means human voice), it is first divided into M frames, and the short-time Fourier transform is performed to obtain the amplitude spectrum of the M frames. ω is the frequency index, and then the mean of the amplitude spectrum of the song M frames is calculated, which is Γ(ω) (k) , then for the N tracks in a group of tracks of a type, calculate the amplitude spectrum mean of these N tracks, recorded as
[0048] When normalizing the amplitude spectrum of the target vocal or background music, the computing device may first obtain the mean amplitude spectrum of the corresponding type. After obtaining the mean amplitude spectrum of the corresponding type of track, the computing device may calculate the difference in the spectrum, for example, the difference in the spectrum is calculated by the mean amplitude spectrum Γ(ω) of the vocal or background music track. (k) and the mean of the amplitude spectrum for the first or second set of tracks The difference is indicated by , and is calculated by the following formula (3):
[0049] The computing device can then design and apply filters based on the differences in the frequency spectra to find the optimal filter settings, thereby achieving equalization and normalization. During the training process of the mixing model, both the first and second sets of audio tracks are used to train the mixing model, and equalization and normalization are also performed on each sample track in the first or second set of audio tracks using the above method.
[0050] In some embodiments, the computing device may calculate the positions of transients of a set of tracks using an onset detection algorithm, and then calculate the amplitude mean of these transient values. and variance where μi k is the amplitude value of the transient value of the i-th track in the k-th type of track, σ i k is the variance of the transient value of the i-th track in the k-th type of track, and N is the number of tracks.
[0051] When the computing device normalizes the dynamic range control of the target human voice or target background music, it is necessary to limit the audio peak. (k) >(P μ (k) +P σ (k) ), the conventional Dynamic Range Control (DRC) algorithm is used to compress the peak value according to the attack and release parameters until μ is satisfied. (k) <(P μ (k) +P σ (k) ), thereby achieving normalization of dynamic range control.
[0052] In some embodiments, the computing device may also calculate an average similarity measure for the sound image. In this process, the computing device first calculates a similarity measure Ψ(ω) for the left and right channels to approximate the sound image shift added during mixing. Since only the amplitude shift is of interest, Ψ(ω) can be calculated using only the frequency amplitude using the following formula (4):
[0053] Where X(ω) L and X(ω) R They are the amplitude spectra after short-time Fourier transform of the left and right channels, respectively. The energy (amplitude spectrum value) is used to calculate the position of the sound image. Ψ(ω) indicates whether the sound image at frequency ω0 is in the middle, that is, it satisfies When the sound image is biased to the left or right,
[0054] The computing device further determines whether the sound image is shifted to the left or right, and calculates the difference in sound image similarity between the left and right channels using the following formula (5): Δ(ω) = Ψ(ω) L -Ψ(ω) R (5)
[0055] in: When Δ(ω)>0, the sound image is biased to the left, otherwise it is biased to the right.
[0056] The gain values for the left and right channel sound images are calculated using the following formulas (6) and (7): Φ(ω) L =1-α(ω) (6) Φ(ω) R =α(ω) (7)
[0057] in, The value of α(ω) can be adjusted according to the user's needs. The above process can be calculated for a single audio track, and the computing device can obtain S(ω) based on the average value of α(ω) for a group of audio tracks of the same type.
[0058] When the computing device normalizes the sound image of the target human voice or target background music, the normalization is performed through the following process: first, the Ψ(ω) and Δ(ω) of the human voice or background music are calculated, and then the corresponding Φ(ω) is calculated. L and Φ(ω) R Then use S(ω) to calculate the corresponding set of tracks and Then, the amplitude spectrum is normalized using the following formulas (8) and (9):
[0059] Finally, the computing device uses the original short-time Fourier transform phase information of each frame and the above-mentioned amplitude spectrum information to perform inverse transformation and transfer to the time domain.
[0060] In some embodiments, the computing device may obtain a mean loudness value for a set of tracks. In this process, the computing device may determine the loudness of each track in the set and then average the loudness values of all tracks to obtain the mean loudness value. The mean loudness value for the set of tracks is calculated using the following formula (10):
[0061] Where k is the category of the audio track, which can be vocals or background music. LUFS() is the loudness calculation function, x i is the i-th track in the set. For each track in the first and second sets, a specific loudness value is assigned to the entire duration of the track. N is the number of tracks in the first and second sets. The sum divided by N represents the average loudness of the vocals and background music for all songs.
[0062] When the computing device performs loudness normalization on the target vocals or background music, the computing device may adjust the loudness value of the target vocals or background music to the average loudness value of a group of audio tracks of the corresponding category, and then perform normalization using the maximum loudness value of the target vocals or background music.
[0063] At block 308, a target mixed audio is generated based on the processed target vocals and the processed target background music. For example, the computing device normalizes the target vocals and the target background music according to the first set of sound features and the second set of sound features, and then generates the target mixed audio using the normalized target vocals and the target background music.
[0064] In some embodiments, when generating the target mixed audio, the computing device may perform upsampling and downsampling operations on the processed target vocals and the processed target background music, and then generate the target mixed audio using the sampled target vocals and the target background music.
[0065] In some embodiments, when generating the target mixed audio, the computing device generates the target mixed audio by applying the processed target human voice and the processed target background music to a mixing model, where the mixing model is a pre-trained neural network model.
[0066] In some embodiments, a computing device can train a neural network model. During this training process, the computing device first obtains vocal samples from sample tracks in a first set of audio tracks for vocals and background music samples from sample tracks in a second set of audio tracks for background music. The computing device then calculates a first set of sound features for the first set of audio tracks and a second set of sound features for the second set of audio tracks using the methods described above. For example, the computing device can calculate the amplitude spectrum mean, transient amplitude mean, transient variance, average similarity measure for sound images, or mean loudness for the first or second set of audio tracks. The computing device then normalizes the vocal samples and background music samples using the first and second sets of sound features, as described above. Next, the computing device trains the mixing model using the processed vocal samples, the processed background music samples, and a sample mixed audio sample, where the mixed audio sample is a combination of the vocal samples and background music samples. In some embodiments, since the vocal samples and background music samples are separated from a song, the mixed audio sample can be the song itself.
[0067] In some embodiments, the computing device obtains sample magnitude spectra and sample phase spectra of the channels of the sample mixed audio. The computing device also obtains predicted magnitude spectra and predicted phase spectra of the channels of the mixed audio predicted by the mixing model and adjusts parameters of the mixing model based on the sample magnitude spectra and sample phase spectra.
[0068] In some embodiments, the comprehensive spectrum loss function = λ1*||y_L_amp-y_hat_L_amp||+λ2*||y_R_amp-y_hat_R_amp||+λ3*||y_L_phase-y_hat_L_phase||+λ4*||y_R_phase-y_hat_R_phase||. Here, y_L_amp and y_R_amp respectively represent the amplitude spectra of the left and right channels of the original audio signal or the sample mixed audio, y_hat_L_amp and y_hat_R_amp respectively represent the amplitude spectra of the left and right channels of the separated audio signals predicted by the model, y_L_phase and y_R_phase respectively represent the phase spectra of the left and right channels of the original audio signal or the sample mixed audio, y_hat_L_phase and y_hat_R_phase respectively represent the phase spectra of the left and right channels of the separated audio signals predicted by the model, and λ1, λ2, λ3, and λ4 are weight parameters used to balance the losses of different parts.
[0069] In some embodiments, during the training process, the computing device may also use methods such as grid search or random search to try different weight parameter combinations and evaluate their impact on model performance. By observing the performance of the model on the validation set, the optimal weight parameter combination can be found. In the grid search process of the hyperparameter search, a set of possible weight parameter combinations can be defined first. These parameter combinations are then arranged and combined to form a parameter grid. For each parameter combination, the model is trained on the training set and the performance is evaluated on the validation set. The parameter combination with the best performance on the validation set is selected as the optimal weight parameter.
[0070] Through the above method, since the first set of sound features and the second set of sound features are utilized and the target human voice and target background music are normalized based on them, the complex characteristics of the audio signal can be better captured, the degree of automation is improved, and it has strong adaptability. It can achieve more intelligent and adaptive audio mixing processing, thereby improving the user experience.
[0071] A schematic diagram of an example method for mixing audio according to an embodiment of the present disclosure is described above in conjunction with Figure 3. A schematic diagram of an example mixing model architecture for mixing audio according to an embodiment of the present disclosure is described below in conjunction with Figure 4.
[0072] As shown in FIG4 , example 400 is that a computing device performs upsampling and downsampling operations on the processed target vocals and target background music, and then generates a target mixed audio based on the sampled target vocals and target background music.
[0073] The computing device obtains the input 402 of the model, wherein the input 402 includes the left and right channel information of the target vocals and the left and right channel information of the target background music. In addition, the length of each input can be from 1 second to 10 seconds. The computing device further performs downsampling and upsampling operations on the obtained input 402, wherein the downsampling module 404 and the upsampling module 406 illustrate the specific operations of sampling: the computing device performs a one-dimensional convolution on the input and performs group normalization. Finally, the computing device compresses the input through the one-dimensional convolution in the convolution module 408 to obtain the output 410, wherein the output 410 includes the left and right channel information of the mixed audio of the target vocals and the target background music. The mixing model includes four downsampling modules and four upsampling modules, which are only examples and not specific limitations of the present disclosure. It can adopt any suitable number of upsampling modules and downsampling modules, and the size of the convolution module can also be set to any suitable value.
[0074] FIG5 shows a schematic block diagram of an apparatus for mixing audio according to an embodiment of the present disclosure. As shown in FIG5 , the apparatus 500 includes a track acquisition module 510 configured to acquire a target vocal for a first track among a plurality of tracks and a target background music for a second track among a plurality of tracks; a feature determination module 520 configured to determine a first set of sound features for a first set of tracks related to the target vocal and a second set of sound features for a second set of tracks related to the target background music; a normalization module 530 configured to normalize the target vocal and the target background music based on the first set of sound features and the second set of sound features; and a mixed audio generation module 540 configured to generate a target mixed audio based on the processed target vocal and the processed target background music.
[0075] In some embodiments, the track acquisition module 510 includes a dry sound acquisition module configured to determine the dry sound for human voice as the target human voice of the first track; and determine the dry sound for background music as the target background music of the second track.
[0076] In some embodiments, the feature determination module 520 includes an audio signal determination module configured to obtain a first set of sound features based on audio signals in a first set of audio tracks; and obtain a second set of sound features based on audio signals in a second set of audio tracks.
[0077] In some embodiments, the feature determination module 520 also includes a feature calculation module, which is configured to calculate the first set of sound features and the second set of sound features and the average features in the multi-source separation data set, the features including at least one of the following: amplitude spectrum mean, amplitude mean of transient values, variance of transient values, average similarity measure for sound image, or loudness mean.
[0078] In some embodiments, the normalization module 530 includes an equalization normalization module, which is configured to perform equalization normalization on the target human voice based on the amplitude spectrum mean; a dynamic range control normalization module, which is configured to normalize the dynamic range of the target human voice based on the amplitude mean of the transient value and the variance of the transient value; an audio-visual normalization module, which is configured to normalize the audio-visual image of the target human voice based on the average similarity measure for the audio-visual image; and a loudness normalization module, which is configured to normalize the loudness of the target human voice based on the loudness mean.
[0079] In some embodiments, the target mixed audio generation module 540 also includes a sampling module configured to perform upsampling and downsampling operations on the processed target vocals and the processed target background music; and to generate the target mixed audio based on the sampled target vocals and the target background music.
[0080] In some embodiments, the target mixed audio generation module 540 also includes a channel separation module, which is configured to perform channel separation on the processed target vocals and the processed target background music, determine the left channel information and the right channel information of the processed target vocals and the processed target background music; and input the left channel information and the right channel information into the audio mixing model respectively to obtain the target mixed audio.
[0081] In some embodiments, the apparatus 500 further includes: a data preprocessing module configured to perform digital signal processing on the sample audio based on the average features in the multi-track audio source separation dataset.
[0082] In some embodiments, the computing device first calculates the average features of all tracks in the multi-track sound source separation dataset, and then adds the average features to the sample audio features to calculate and obtain the processed sample vocal features.
[0083] This method uses the average feature to pre-process the sample audio and mix the sample audio based on it, which can better capture the complex characteristics of the audio signal, improve the degree of automation, have strong adaptability, and achieve more intelligent and adaptive audio mixing processing, thereby improving the user experience.
[0084] Fig. 6 shows a schematic block diagram of an example device 600 that can be used to implement an embodiment of the present disclosure. The computing device 108 in Fig. 1 can be implemented using device 600. As shown in the figure, device 600 includes a central processing unit (CPU) 601, which can perform various appropriate actions and processes according to the computer program instructions stored in a read-only memory (ROM) 602 or the computer program instructions loaded into a random access memory (RAM) 603 from a storage unit 608. In RAM 603, various programs and data required for the operation of device 600 can also be stored. CPU 601, ROM 602 and RAM 603 are connected to each other via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.
[0085] Multiple components in device 600 are connected to I / O interface 605, including: input unit 606, such as a keyboard, mouse, etc.; output unit 607, such as various types of displays, speakers, etc.; storage page 608, such as a magnetic disk, optical disk, etc.; and communication unit 609, such as a network card, modem, wireless communication transceiver, etc. Communication unit 609 allows device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0086] The various processes and processing described above, such as method 300, may be performed by processing unit 601. For example, in some embodiments, method 300 may be implemented as a computer software program that is tangibly embodied in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed onto device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by CPU 601, one or more actions of method 300 described above may be performed.
[0087] The present disclosure may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present disclosure.
[0088] Computer-readable storage media can be a tangible device that can hold and store instructions used by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination thereof. The computer-readable storage media used herein is not to be interpreted as a transient signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated by a waveguide or other transmission medium (for example, a light pulse by a fiber optic cable), or an electrical signal transmitted by a wire.
[0089] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.
[0090] The computer program instructions for performing the operation of the present disclosure can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, and conventional procedural programming languages such as "C" language or similar programming languages. Computer-readable program instructions can be executed entirely on a user's computer, partially on a user's computer, executed as an independent software package, partially on a user's computer and partially on a remote computer, or executed entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer via any type of network including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., utilizing an Internet service provider to connect via the Internet). In some embodiments, by utilizing the state information of computer-readable program instructions to personalize an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute computer-readable program instructions, thereby realizing various aspects of the present disclosure.
[0091] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0092] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0093] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0094] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the part of the module, program segment or instruction contains one or more executable instructions for realizing the prescribed logical function. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the prescribed function or action, or can be implemented by a combination of dedicated hardware and computer instructions.
[0095] While various embodiments of the present disclosure have been described above, the foregoing description is intended to be illustrative, non-exhaustive, and not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or technical improvements to existing technologies, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A method for mixing audio, comprising: Obtaining a target vocal for a first audio track among a plurality of audio tracks and a target background music for a second audio track among the plurality of audio tracks; determining a first set of sound features for a first set of audio tracks associated with vocals and a second set of sound features for a second set of audio tracks associated with background music; Normalizing the target human voice and the target background music based on the first set of sound features and the second set of sound features; as well as A target mixed audio is generated based on the processed target vocals and the processed target background music.
2. The method according to claim 1, wherein obtaining a target vocal for a first audio track of the plurality of audio tracks and a target background music for a second audio track of the plurality of audio tracks comprises: determining a dry sound for a human voice as the target human voice of the first audio track; as well as The dry sound for the background music is determined as the target background music of the second audio track.
3. The method of claim 1 , wherein determining a first set of sound features for a first set of audio tracks associated with human voices and a second set of sound features for a second set of audio tracks associated with background music comprises: obtaining the first set of sound features based on the audio signals in the first set of audio tracks; as well as The second set of sound features is obtained based on audio signals in the second set of audio tracks.
4. The method of claim 3, wherein the first set of sound features and the second set of sound features include at least one of the following: Amplitude spectrum mean, amplitude mean of transient values, variance of transient values, average similarity measure for sound images, or mean loudness.
5. The method according to claim 4, wherein normalizing the target vocals and the target background music comprises: performing normalization processing on the target human voice using the first set of sound features; as well as The target background music is normalized using the second set of sound features.
6. The method according to claim 5, wherein normalizing the target human voice using the first set of sound features comprises one or more of the following: Based on the amplitude spectrum mean, performing equalization normalization on the target human voice; Normalizing the target human voice for dynamic range control based on the amplitude mean of the transient value and the variance of the transient value; Performing sound and image normalization on the target human voice based on the average similarity metric for the sound and image; or The target human voice is subjected to loudness normalization based on the loudness mean.
7. The method according to claim 1, wherein generating a target mixed audio based on the processed target vocals and the processed target background music comprises: performing upsampling and downsampling operations on the processed target vocals and the processed target background music; as well as The target mixed audio is generated based on the sampled target human voice and the target background music.
8. The method according to claim 1, wherein generating a target mixed audio based on the processed target vocals and the processed target background music comprises: The target mixed audio is generated by applying the processed target vocals and the processed target background music to a mixing model.
9. The method according to claim 8, further comprising: Obtaining a sample vocal from a sample track in the first group of audio tracks and a sample background music from a sample track in the second group of audio tracks; performing normalization processing on the sample human voice and the sample background music based on the first set of sound features and the second set of sound features; as well as The mixing model is trained based on the processed sample human voice, the processed sample background music and the sample mixed audio.
10. The method of claim 9, wherein training the mixing model comprises: determining a sample magnitude spectrum and a sample phase spectrum for a channel of the sample mixed audio; determining a predicted magnitude spectrum and a predicted phase spectrum of the channels of the mixed audio predicted by the mixing model; and Parameters of the mixing model are adjusted based on the sample magnitude spectrum, the sample phase spectrum, the predicted magnitude spectrum, and the predicted phase spectrum.
11. The method of claim 10, wherein adjusting the parameters of the mixing model comprises: Grid search or random search is used to adjust the parameters of the mixing model.
12. The method according to claim 9, further comprising: The sample human voice and the sample background music are obtained by separating the sample mixed audio.
13. An apparatus for mixing audio, comprising: a track acquisition module configured to acquire a target vocal for a first track among the plurality of tracks and a target background music for a second track among the plurality of tracks; a feature determination module configured to determine a first set of sound features for a first set of audio tracks associated with human voices and a second set of sound features for a second set of audio tracks associated with background music; a normalization module, configured to normalize the target human voice and the target background music based on the first set of sound features and the second set of sound features; as well as The mixed audio generation module is configured to generate a target mixed audio based on the processed target human voice and the processed target background music.
14. An electronic device comprising: at least one processor; as well as A storage device for storing at least one program, wherein when the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1 to 12.
15. A computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the method according to any one of claims 1 to 12 when executed by a processor.
16. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 12.