An audio main-accompaniment separation method, device, storage medium and program product

By acquiring the energy envelope difference signal of the song audio, determining the extreme left and extreme right signal intervals, generating and adjusting the dual-channel lead vocal and backing vocal signals, the problems of low accuracy and high cost of lead vocal and backing vocal separation in existing technologies are solved, and highly refined lead vocal and backing vocal separation and sound image information restoration are achieved.

CN119252274BActive Publication Date: 2026-04-07TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-18
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing technologies for separating lead and backing vocals are not very accurate and are costly, especially deep learning-based methods which require a large amount of data for training.

Method used

By acquiring the energy envelope difference signal of the song audio, the extreme left and extreme right signal intervals are determined. The dual-channel lead vocal and backing vocal signals are generated using mono correlated and uncorrelated signals, respectively. Sound image information is added, and the channel signals within the signal intervals are adjusted to improve the degree of separation refinement.

Benefits of technology

It achieves low-cost and highly precise separation of lead and backing vocals, which can restore audio-visual information and enhance user experience and immersion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119252274B_ABST
    Figure CN119252274B_ABST
Patent Text Reader

Abstract

The application discloses a kind of audio main chorus separation method, equipment, storage medium and program product, it is related to signal processing technical field.The method comprises: obtaining song audio, extract monophonic relevant signal and monophonic irrelevant signal from it;Obtain the energy envelope difference signal corresponding to song audio, determine the extreme left signal interval and the extreme right signal interval;Based on monophonic relevant signal, get the left channel signal of main singer and the right channel signal of main singer, based on monophonic irrelevant signal, get the left channel signal of chorus and the right channel signal of chorus;Using monophonic irrelevant signal, respectively, add sound image information to the left channel signal of main singer in extreme left signal interval and the right channel signal of main singer in extreme right signal interval, obtain double-channel main singer signal;Adjust the right channel signal of chorus in extreme left signal interval, the left channel signal of chorus in extreme right signal interval, obtain double-channel chorus signal.Can improve the degree of refinement of main singer and chorus separation at low cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of signal processing technology, and in particular to an audio master-singer separation method, device, storage medium, and program product. Background Technology

[0002] In existing technologies, vocal and backing vocal separation involves extracting the mono vocal and backing vocal signals from a two-channel vocal signal, then simply copying the mono track into two tracks to obtain the two-channel vocal and backing vocal signals; however, this method is not very accurate. Existing technologies also employ deep learning-based vocal and backing vocal separation, specifically by training a large amount of vocal and backing vocal audio data to enable a neural network to separate the vocal and backing vocal signals; however, this requires a large amount of data for model training, resulting in high technical costs. Summary of the Invention

[0003] In view of this, the purpose of this invention is to provide an audio main vocal and backing vocal separation method, device, storage medium, and program product, which can improve the precision of main vocal and backing vocal separation and achieve low-cost, high-fidelity separation of main vocal and backing vocal information. The specific solution is as follows:

[0004] In the first aspect, this application discloses an audio main and backing vocal separation method, including:

[0005] Acquire the song audio and extract the mono-correlated signal and mono-uncorrelated signal from the song audio;

[0006] Obtain the energy envelope difference signal of the song audio, and determine the extreme left signal interval and the extreme right signal interval based on the energy envelope difference signal;

[0007] Based on the mono correlated signal, the lead vocal left channel signal and the lead vocal right channel signal are obtained; based on the mono uncorrelated signal, the backing vocal left channel signal and the backing vocal right channel signal are obtained.

[0008] Using the mono uncorrelated signal, sound image information is added to the lead vocal left channel signal in the far left signal interval and the lead vocal right channel signal in the far right signal interval to obtain a stereo lead vocal signal.

[0009] Adjusting the right channel signal of the backing vocals within the far left signal range and adjusting the left channel signal of the backing vocals within the far right signal range yields a dual-channel backing vocal signal.

[0010] Optionally, obtaining the energy envelope difference signal of the song audio includes:

[0011] The left and right channel signals of the song audio are divided into frames, and the energy envelope of each frame is calculated to obtain the energy envelope of the left channel and the energy envelope of the right channel.

[0012] Based on the difference between the energy envelope of the left channel and the energy envelope of the right channel, the energy envelope difference signal corresponding to the song audio is obtained.

[0013] Optionally, determining the extreme left signal interval and the extreme right signal interval based on the energy envelope difference signal includes:

[0014] Obtain a first energy threshold and a second energy threshold; the first energy threshold is greater than the second energy threshold; the first energy threshold is an absolute value, and the first energy threshold corresponds to a first positive energy threshold and a first negative energy threshold; the second energy threshold is an absolute value, and the second energy threshold corresponds to a second positive energy threshold and a second negative energy threshold;

[0015] The first region with energy greater than the first positive energy threshold is located from the energy envelope difference signal. The first region is used as the center to search along the energy envelope difference signal to both sides to locate two time points with energy equal to the second positive energy threshold. The time period between the two time points is taken as the extreme left signal interval.

[0016] The second region with energy less than the first negative energy threshold is located from the energy envelope difference signal. The second region is used as the center to search along the energy envelope difference signal to both sides to locate two time points with energy equal to the second negative energy threshold. The time period between the two time points is taken as the extreme right signal interval.

[0017] Optionally, calculating the energy envelope of each frame of signal to obtain the energy envelope of the left channel and the energy envelope of the right channel includes:

[0018] Calculate the energy envelope of each frame signal, and perform smoothing and normalization on the energy envelope to obtain the energy envelope of the left channel and the energy envelope of the right channel;

[0019] Smoothing the energy envelope includes:

[0020] The energy envelope is upsampled to obtain an upsampled signal sequence. The upsampled signal sequence is then convolved in the time domain to obtain a convolved sequence. Finally, the convolved sequence is downsampled to obtain a downsampled signal sequence.

[0021] Optionally, the step of using the mono uncorrelated signal to add image information to the lead vocal left channel signal in the far left signal interval and the lead vocal right channel signal in the far right signal interval to obtain a two-channel lead vocal signal includes:

[0022] The updated vocal left channel signal is obtained by comparing the lead vocal left channel signal within the extreme left signal interval with the unrelated signal within the extreme left signal interval.

[0023] The updated vocal right channel signal is obtained by comparing the difference between the vocal right channel signal within the far right signal interval and the unrelated signal within the far right signal interval.

[0024] Based on the updated left channel signal and the updated right channel signal of the lead vocalist, a dual-channel lead vocal signal is obtained.

[0025] Optionally, adjusting the right channel signal of the backing vocals within the far left signal range and adjusting the left channel signal of the backing vocals within the far right signal range to obtain a dual-channel backing vocal signal includes:

[0026] The right channel signal of the backing vocals in the far left signal interval is adjusted to 0 to obtain the updated right channel signal of the backing vocals.

[0027] The left channel signal of the backing vocals in the far right signal range is adjusted to 0 to obtain the updated left channel signal of the backing vocals.

[0028] Based on the updated right channel backing vocal signal and the updated left channel backing vocal signal, a dual-channel backing vocal signal is obtained.

[0029] Optionally, after adjusting the right channel signal of the backing vocals within the far left signal range and adjusting the left channel signal of the backing vocals within the far right signal range to obtain a dual-channel backing vocal signal, the method further includes:

[0030] The two-channel accompaniment signal is subjected to decorrelation processing to obtain the decorrelation-processed two-channel accompaniment signal.

[0031] Optionally, the decorrelation processing of the two-channel accompaniment signal includes:

[0032] The dual-channel accompaniment signal is delayed and sampled for decorrelation processing.

[0033] Alternatively, decorrelation processing can be performed by changing the phase of the two-channel accompaniment signal.

[0034] Secondly, this application discloses an electronic device, comprising:

[0035] Memory, used to store computer programs;

[0036] A processor is used to execute the computer program to implement the aforementioned audio main and backing vocal separation method.

[0037] Thirdly, this application discloses a computer-readable storage medium for storing a computer program; wherein the computer program, when executed by a processor, implements the aforementioned audio main and backing vocal separation method.

[0038] In this application, the following steps are taken: First, obtain the audio of a song; extract mono-correlated and mono-uncorrelated signals from the audio. Second, obtain the energy envelope difference signal of the audio; determine the extreme left and extreme right signal intervals based on the energy envelope difference signal. Third, obtain the lead vocal left and right channel signals based on the mono-correlated signals, and obtain the backing vocal left and right channel signals based on the mono-uncorrelated signals. Fourth, use the mono-uncorrelated signals to add image information to the lead vocal left channel signal in the extreme left signal interval and the lead vocal right channel signal in the extreme right signal interval to obtain a stereo lead vocal signal. Fifth, adjust the backing vocal right channel signal in the extreme left signal interval and the backing vocal left channel signal in the extreme right signal interval to obtain a stereo backing vocal signal. It can be seen that by detecting the energy envelope difference between the left and right channels of the song audio, the possible ranges of extreme left and extreme right signals can be determined. The extreme left and extreme right signals are used to give the lead singer and backing vocals sound image information, and finally the two-channel lead singer signal and two-channel backing vocal signal are obtained, which improves the precision of the separation of lead singer and backing vocals and achieves low-cost separation of lead singer and backing vocals that restores sound image information. Attached Figure Description

[0039] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0040] Figure 1 A flowchart of an audio main and backing vocal separation method provided in this application;

[0041] Figure 2 A specific energy envelope diagram is provided for this application;

[0042] Figure 3 A flowchart of a specific audio main and backing vocal separation method provided in this application;

[0043] Figure 4 A schematic diagram of a specific audio master-accompaniment separation system framework provided in this application;

[0044] Figure 5 A flowchart of a specific audio main and backing vocal separation method provided in this application;

[0045] Figure 6 This application provides a structural diagram of an electronic device. Detailed Implementation

[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0047] In existing technologies, vocal and backing vocal separation involves extracting the mono vocal and backing vocal signals from a two-channel vocal signal and then simply copying the mono signal into two tracks to obtain the two-channel vocal and backing vocal signals; however, this method has low accuracy. Existing technologies also employ deep learning-based vocal and backing vocal separation, specifically training a large amount of vocal and backing vocal audio data to enable a neural network to separate the vocal and backing vocal signals; however, this requires a large amount of data for model training, resulting in high technical costs. To overcome the above technical problems, this application proposes an audio vocal and backing vocal separation method that can improve the precision of vocal and backing vocal separation, achieving low-cost separation that restores the original sound image information.

[0048] This application discloses an audio main and backing vocal separation method. See also Figure 1 As shown, the method may include the following steps:

[0049] Step S11: Obtain the song audio and extract the mono-correlated signal and mono-uncorrelated signal from the song audio.

[0050] First, acquire the song audio. It's important to note that the acquired audio at this stage includes vocal signals, specifically the signals from the lead singer and backing vocals. In other words, this is the song audio that needs to be separated into lead and backing vocals. The lead singer is the vocalist in a band or musical group who plays the primary vocal role; the backing vocalist sings alongside the lead singer, complementing their performance.

[0051] Specifically, this involves extracting the mono correlation signal from the song audio, i.e., extracting the lead vocal signal; and simultaneously extracting the stereo correlation signal from the song audio, i.e., extracting the backing vocal signal. The extraction of these correlation and uncorrelated signals can be achieved using digital signal processing algorithms, such as M / S (Mid / Side) decomposition, adaptive filtering, and PCA-based signal decomposition, to separate the correlation and uncorrelated signals from the two-channel vocal signal.

[0052] Taking M / S decomposition as an example, the formula is as follows:

[0053] ; ;

[0054] in, This is the left channel signal. This is the right channel signal. In Mid / Side stereo, the Mid channel is located in the center of the stereo image, and the Side channels are located on either side of the stereo image. In typical song mixing, the lead vocal image is usually mixed in the center of the stereo image, while the backing vocals are mixed to the sides of the stereo image. This results in the lead vocal mono signal Mid(n) and the backing vocal mono signal Side(n).

[0055] Step S12: Obtain the energy envelope difference signal of the song audio, and determine the extreme left signal interval and the extreme right signal interval based on the energy envelope difference signal.

[0056] Based on the energy envelope difference signal corresponding to the song's audio, the extreme left and extreme right signal intervals are determined. An extreme left signal refers to a stereo audio file where only the left channel has a signal; similarly, an extreme right signal refers to a stereo audio file where only the right channel has a signal. It's understandable that, since an extreme left signal means only the left channel has a signal in a stereo audio file, the signal energy of the left channel within the extreme left signal interval must be much greater than the signal energy of the right channel; the same applies to the right channel. Therefore, the occurrence and end times of the extreme left and extreme right signals, i.e., the start and end points of the extreme left and extreme right signal intervals, can be located based on the energy envelope difference between the left and right channels.

[0057] The acquisition of the energy envelope difference signal of the song audio can include: dividing the left and right channel signals of the song audio into frames and calculating the energy envelope of each frame to obtain the energy envelope of the left and right channels; and obtaining the energy envelope difference signal corresponding to the song audio based on the difference between the energy envelopes of the left and right channels. Specifically, calculating the energy envelope of each frame to obtain the energy envelopes of the left and right channels includes: calculating the energy envelope of each frame, smoothing and normalizing the energy envelopes to obtain the energy envelopes of the left and right channels; and smoothing the energy envelopes specifically includes: upsampling the energy envelopes to obtain an upsampled signal sequence, performing temporal convolution on the upsampled signal sequence to obtain a convolved sequence, and downsampling the convolved sequence to obtain a downsampled signal sequence.

[0058] To calculate the signal energy envelope, firstly, the input signal is divided into frames, and the energy envelope of each frame is calculated. The energy envelope calculation formula can adopt the short-time average energy calculation formula, as follows:

[0059] ;

[0060] Where x(n) represents the nth frame of the speech signal, and N represents the window length, this formula is used to calculate the short-time average energy of the speech signal at a certain moment. After obtaining the energy envelope of each frame, the energy envelope is then smoothed and normalized. The smoothing process includes: upsampling the energy envelope, then performing a temporal convolution window function, followed by downsampling; finally, normalization is performed to obtain the smoothed energy envelope, so that the energy envelope difference signal can be obtained later based on the energy envelope difference.

[0061] In some embodiments, determining the extreme left signal interval and the extreme right signal interval based on the energy envelope difference signal specifically includes the following steps: obtaining a first energy threshold and a second energy threshold; the first energy threshold is greater than the second energy threshold; the first energy threshold is an absolute value, corresponding to a first positive energy threshold and a first negative energy threshold; the second energy threshold is an absolute value, corresponding to a second positive energy threshold and a second negative energy threshold; locating a first region with energy greater than the first positive energy threshold from the energy envelope difference signal, searching outwards along the energy envelope difference signal with the first region as the center, locating two time points where the energy is equal to the second positive energy threshold, and taking the time period between the two time points as the extreme left signal interval; locating a second region with energy less than the first negative energy threshold from the energy envelope difference signal, searching outwards along the energy envelope difference signal with the second region as the center, locating two time points where the energy is equal to the second negative energy threshold, and taking the time period between the two time points as the extreme right signal interval.

[0062] In other words, this embodiment can use a dual-threshold detection method to extract the extreme left and extreme right intervals, that is, to determine the start and end points of the extreme left and extreme right signals. Extracting the start and end points of the extreme left and extreme right signals can be considered as an endpoint detection task based on the smoothed energy envelope difference signal of the left and right channels. The dual-threshold detection method is used for extraction, that is, a higher first energy threshold T1 is selected; signals with an absolute energy envelope value higher than the T1 threshold are considered extreme left and extreme right signals. Then, the search is expanded outwards from this intersection point to find two points where the absolute energy value intersects with the lower second energy threshold T2, which are considered the start and end points of the extreme left and extreme right signals. For example... Figure 2 As shown, the shaded area represents the left channel waveform of the song audio, and the curve (energy diff) represents the smoothed energy envelope difference signal. In this embodiment, the smoothed energy envelope of the left channel is subtracted from the smoothed energy envelope of the right channel. Of course, the smoothed energy envelope of the right channel can also be subtracted from the smoothed energy envelope of the left channel. Figure 2 The signal in the signal contains a leftmost signal interval and a rightmost signal interval.

[0063] Step S13: Obtain the lead vocal left channel signal and lead vocal right channel signal based on the mono correlated signal, and obtain the backing vocal left channel signal and backing vocal right channel signal based on the mono uncorrelated signal.

[0064] Since the extracted lead vocals and backing vocals are mono signals, they need to be copied into dual tracks to become stereo. Specifically, the lead vocal left and right channel signals are obtained by copying the correlated mono signals, and the backing vocal left and right channel signals are obtained by copying the uncorrelated mono signals. That is, stereo consists of two independent audio signals, left and right, played through the left and right speakers respectively. However, the copied lead vocal left and right channel signals at this point lack image information and are overly correlated, losing the phase information of the original vocals; the same applies to the backing vocal left and right channel signals. Therefore, it is necessary to use extreme left and extreme right signals to give the lead vocals and backing vocals image information. The spatial position of each vocal part as perceived by the listener, and the resulting sound image, is called the sound image.

[0065] In addition, the order of steps S12 and S13 is not set. Step S13 can be executed first and then step S12. Of course, they can also be executed simultaneously.

[0066] Step S14: Using the mono uncorrelated signal, add sound image information to the lead vocal left channel signal in the far left signal interval and the lead vocal right channel signal in the far right signal interval to obtain a two-channel lead vocal signal.

[0067] That is, for the sound image of the lead vocal signal, the extreme left and extreme right signals are weighted onto the stereo signal that was copied from the mono channel to obtain a stereo lead vocal signal containing sound image information.

[0068] In some embodiments, the step of adding imaging information to the lead vocal left channel signal in the far left signal interval and the lead vocal right channel signal in the far right signal interval using the mono uncorrelated signal to obtain a stereo lead vocal signal may specifically include: adding imaging information to the lead vocal left channel signal in the far left signal interval using the uncorrelated signal in the far left signal interval to obtain an updated lead vocal left channel signal; adding imaging information to the lead vocal right channel signal in the far right signal interval using the uncorrelated signal in the far right signal interval to obtain an updated lead vocal right channel signal; that is, obtaining the updated lead vocal left channel signal based on the difference between the lead vocal left channel signal in the far left signal interval and the uncorrelated signal in the far left signal interval; obtaining the updated lead vocal right channel signal based on the difference between the lead vocal right channel signal in the far right signal interval and the uncorrelated signal in the far right signal interval; and finally, obtaining a stereo lead vocal signal based on the updated lead vocal left channel signal and the updated lead vocal right channel signal.

[0069] Understandably, when the vocal image of the lead vocalist and backing vocalist is assigned to extreme left or extreme right signals, the separation of the lead vocalist and backing vocalist based on relevant and irrelevant signals is not entirely accurate. Therefore, after determining the start and end points of several extreme left and extreme right signals, further vocal image assignment is needed to improve the precision of the separation. The irrelevant signals obtained in the previous separation are considered to be waveform data of extreme left and extreme right signals. For the lead vocalist, the stereo sound of the lead vocalist without vocal image information obtained by copying the relevant signals into two tracks includes: lead vocalist left channel LV_L and lead vocalist right channel LV_R. If the start and end points of the detected extreme left signal are [t1, t2], then LV_L[t1 : t2] = LV_L[t1 : t2] - irrelevant signal [t1 : t2]. Similarly, for the extreme right signal, the start and end points of the extreme right signal are [t3, t4], and LV_R[t3 : t4] = LV_R[t3 : t4] - irrelevant signal [t3 : t4].

[0070] Step S15: Adjust the right channel signal of the backing vocals in the far left signal range and adjust the left channel signal of the backing vocals in the far right signal range to obtain a dual-channel backing vocal signal.

[0071] As can be seen, the relevant and unrelated signals of the mono channel are first extracted from the human voice stereo. Then, by detecting the energy envelope difference between the left and right channels in the human voice stereo, the possible extreme left and right sound image information is extracted and assigned to the separated main and backing vocal stereo, resulting in a two-channel lead vocal and two-channel backing vocal containing sound image information.

[0072] The execution order of steps S14 and S15 is not limited; step S15 can be executed first, or they can be executed simultaneously.

[0073] The above-mentioned adjustment of the right channel signal of the backing vocals within the far left signal interval and the adjustment of the left channel signal of the backing vocals within the far right signal interval to obtain a two-channel backing vocal signal may include: adjusting the right channel signal of the backing vocals within the far left signal interval to 0 to obtain an updated right channel signal of the backing vocals; adjusting the left channel signal of the backing vocals within the far right signal interval to 0 to obtain an updated left channel signal of the backing vocals; and obtaining a two-channel backing vocal signal based on the updated right channel signal of the backing vocals and the updated left channel signal of the backing vocals.

[0074] It is understandable that the right channel can be considered to have no signal in the far left signal range, and the same applies to the far right signal range. Therefore, for backing vocals, a stereo backing vocal without image information is generated by copying unrelated signals into two tracks: backing vocal left channel BV_L and backing vocal right channel BV_R. If the start and end points of the far left signal are detected as [t1, t2], then BV_R[t1 : t2] = 0; the same applies to the far right signal, BV_L[t3 : t4] = 0.

[0075] As can be seen from the above, in this embodiment, the song audio is acquired, and mono-correlated signals and mono-uncorrelated signals are extracted from the song audio; the energy envelope difference signal of the song audio is acquired, and the extreme left signal interval and extreme right signal interval are determined based on the energy envelope difference signal; the lead vocal left channel signal and lead vocal right channel signal are obtained based on the mono-correlated signal, and the backing vocal left channel signal and backing vocal right channel signal are obtained based on the mono-uncorrelated signal; using the mono-uncorrelated signal, sound image information is added to the lead vocal left channel signal in the extreme left signal interval and the lead vocal right channel signal in the extreme right signal interval, respectively, to obtain a stereo lead vocal signal; the backing vocal right channel signal in the extreme left signal interval and the backing vocal left channel signal in the extreme right signal interval are adjusted to obtain a stereo backing vocal signal. It can be seen that by detecting the energy envelope difference between the left and right channels of the song audio, the possible ranges of extreme left and extreme right signals can be determined. The extreme left and extreme right signals are used to give the lead singer and backing vocals sound image information, and finally the two-channel lead singer signal and two-channel backing vocal signal are obtained, which improves the precision of the separation of lead singer and backing vocals and achieves low-cost separation of lead singer and backing vocals that restores sound image information.

[0076] This application discloses a specific method for separating the main and backing vocals in audio. (See also...) Figure 3 As shown, the method may include the following steps:

[0077] Step S21: Obtain the song audio and extract the mono-correlated signal and mono-uncorrelated signal from the song audio.

[0078] Step S22: Obtain the energy envelope difference signal of the song audio, and determine the extreme left signal interval and the extreme right signal interval based on the energy envelope difference signal.

[0079] Step S23: Obtain the lead vocal left channel signal and lead vocal right channel signal based on the mono correlated signal, and obtain the backing vocal left channel signal and backing vocal right channel signal based on the mono uncorrelated signal.

[0080] Step S24: Using the mono uncorrelated signal, add sound image information to the lead vocal left channel signal in the far left signal interval and the lead vocal right channel signal in the far right signal interval respectively to obtain a two-channel lead vocal signal.

[0081] Step S25: Adjust the right channel signal of the backing vocals in the far left signal interval and adjust the left channel signal of the backing vocals in the far right signal interval to obtain a dual-channel backing vocal signal.

[0082] Step S26: Perform decorrelation processing on the dual-channel accompaniment signal to obtain the decorrelation-processed dual-channel accompaniment signal.

[0083] Because the lead vocal and backing vocal signals are in phase, it will feel like both are coming from directly in front of the listener, resulting in a poor user experience. To enhance the immersive experience, the lead and backing vocals need to be positioned differently in space. Specifically, this is achieved by decorrelation processing the dual-channel backing vocal signals, making the backing vocals feel like they are coming from both sides. The resulting lead and backing vocal signals can then be used in immersive sound production.

[0084] The aforementioned decorrelation processing of the two-channel backing vocal signal can include: decorrelation processing by delaying sampling of the two-channel backing vocal signal; or decorrelation processing by changing the phase of the two-channel backing vocal signal. That is, decorrelation can be performed using delay sampling or phase adjustment, and of course, any method capable of decorrelation is acceptable.

[0085] The aforementioned decorrelation processing includes, but is not limited to: time-domain sampling point delay, passing through an all-pass filter, and phase shifting by 90 degrees. The principle behind time-domain sampling point delay is to delay the signal data by a certain number of sampling points. The formula is: y(n) = x(n - t); where y(n) is the delayed signal, x(n) is the input signal, and t is the number of delayed sampling points.

[0086] An all-pass filter is a filter that allows all frequencies to pass through; unlike other common filter types (such as low-pass or high-pass filters), its main function is to change the phase response of a signal while maintaining its amplitude. Taking a first-order all-pass filter as an example, the formula is:

[0087] ;

[0088] H(z) is the z-transform transfer function of the all-pass filter in the frequency domain, and c is the filter coefficient;

[0089] The frequency domain representation of the output signal Y(z) can be obtained by multiplying the frequency domain representation X(z) of the input signal by the transfer function, Y(n) = X(n)H(z).

[0090] Among them, a 90-degree phase shift is also a way to achieve decorrelation by changing the phase of the signal; its advantage is that it gives each frequency component in the signal a phase shift of plus or minus 90 degrees.

[0091] The specific processes of steps S21-S25 can be found in the relevant content disclosed in the foregoing embodiments, and will not be repeated here.

[0092] As can be seen from the above, this embodiment takes the song audio and extracts mono-correlated and mono-uncorrelated signals from it; obtains the energy envelope difference signal of the song audio and determines the extreme left and extreme right signal intervals based on the energy envelope difference signal; obtains the lead vocal left and right channel signals based on the mono-correlated signals, and obtains the backing vocal left and right channel signals based on the mono-uncorrelated signals; uses the mono-uncorrelated signals to add image information to the lead vocal left channel signal in the extreme left signal interval and the lead vocal right channel signal in the extreme right signal interval, respectively, to obtain a stereo lead vocal signal; adjusts the backing vocal right channel signal in the extreme left signal interval and adjusts the backing vocal left channel signal in the extreme right signal interval to obtain a stereo backing vocal signal; and performs decorrelation processing on the stereo backing vocal signal to obtain a decorrelation-processed stereo backing vocal signal. It can be seen that by decorrelation processing the backing vocals, different spatial positions of the lead and backing vocals are achieved, enhancing the user's experience and immersion.

[0093] The system framework used in the audio master-accompaniment separation scheme of this application can be found in [reference needed]. Figure 4 As shown, this may specifically include: a backend server and a number of user terminals that establish communication connections with the backend server. The user terminals include, but are not limited to, tablets, laptops, smartphones, and personal computers (PCs).

[0094] For example Figure 5 As shown in this application, the user terminal sends the song audio (vocals) to be separated to the backend server. This audio is a stereo track. The backend server performs the following steps in the audio main / backing vocal separation method: extracting mono correlated signals and mono uncorrelated signals from the song audio; obtaining the energy envelope difference signal of the song audio; determining the extreme left and extreme right signal intervals based on the energy envelope difference signal; obtaining the lead vocal left and right channel signals based on the mono correlated signals, and obtaining the backing vocal left and right channel signals based on the mono uncorrelated signals, thereby obtaining stereo from mono; and using the mono uncorrelated signals, adding image information to the lead vocal left channel signal in the extreme left signal interval and the lead vocal right channel signal in the extreme right signal interval to obtain a two-channel lead vocal signal. (Vocal); Adjust the right channel signal of the backing vocals within the far left signal range and adjust the left channel signal of the backing vocals within the far right signal range to obtain a two-channel backing vocal signal (background vocal). The background server pushes the two-channel vocal signal and the two-channel backing vocal signal corresponding to the song audio to the user terminal.

[0095] The following section uses a music app as an example to illustrate the technical solution of this application.

[0096] Assuming a user has installed this music app on their device, they send the audio of a song they want to separate for vocals and backing vocals to the app's backend server, along with a separation request. The backend server receives the song audio, extracts the mono-correlated signal to obtain the mono vocal signal, and extracts the mono-uncorrelated signal to obtain the mono backing vocal signal. It then calculates the energy envelope difference signal corresponding to the song audio and determines all extreme left and extreme right signal intervals present in the song audio based on this energy envelope difference signal. The lead vocal left and right channel signals are obtained by copying the mono correlated signal, and the backing vocal left and right channel signals are obtained by copying the mono uncorrelated signal. Using the uncorrelated signal within the extreme left signal interval, image information is added to the lead vocal left channel signal within that interval to obtain an updated lead vocal left channel signal. Similarly, using the uncorrelated signal within the extreme right signal interval, image information is added to the lead vocal right channel signal within that interval to obtain an updated lead vocal right channel signal. Based on the updated lead vocal left and right channel signals, a stereo lead vocal signal is obtained. The backing vocal right channel signal within the extreme left signal interval and the backing vocal left channel signal within the extreme right signal interval are adjusted to obtain a stereo backing vocal signal.

[0097] Accordingly, this application also discloses an audio master / backing vocal separation device, which includes:

[0098] Signal extraction module 11 is used to acquire song audio and extract mono-related signals and mono-unrelated signals from the song audio.

[0099] The signal interval determination module 12 is used to acquire the energy envelope difference signal of the song audio and determine the extreme left signal interval and the extreme right signal interval based on the energy envelope difference signal.

[0100] The dual-channel generation module 13 is used to obtain the lead vocal left channel signal and the lead vocal right channel signal based on the mono-correlated signal, and to obtain the backing vocal left channel signal and the backing vocal right channel signal based on the mono-uncorrelated signal.

[0101] The vocal imaging adjustment module 14 is used to add imaging information to the vocal left channel signal in the far left signal range and the vocal right channel signal in the far right signal range using the mono unrelated signal to obtain a dual-channel vocal signal.

[0102] The backing vocal image adjustment module 15 is used to adjust the backing vocal right channel signal in the extreme left signal range and the backing vocal left channel signal in the extreme right signal range to obtain a dual-channel backing vocal signal.

[0103] As can be seen from the above, in this embodiment, the song audio is acquired, and mono-correlated signals and mono-uncorrelated signals are extracted from the song audio; the energy envelope difference signal of the song audio is acquired, and the extreme left signal interval and extreme right signal interval are determined based on the energy envelope difference signal; the lead vocal left channel signal and lead vocal right channel signal are obtained based on the mono-correlated signal, and the backing vocal left channel signal and backing vocal right channel signal are obtained based on the mono-uncorrelated signal; using the mono-uncorrelated signal, sound image information is added to the lead vocal left channel signal in the extreme left signal interval and the lead vocal right channel signal in the extreme right signal interval, respectively, to obtain a stereo lead vocal signal; the backing vocal right channel signal in the extreme left signal interval and the backing vocal left channel signal in the extreme right signal interval are adjusted to obtain a stereo backing vocal signal. It can be seen that by detecting the energy envelope difference between the left and right channels of the song audio, the possible ranges of extreme left and extreme right signals can be determined. The extreme left and extreme right signals are used to give the lead singer and backing vocals sound image information, and finally the two-channel lead singer signal and two-channel backing vocal signal are obtained, which improves the precision of the separation of lead singer and backing vocals and achieves low-cost separation of lead singer and backing vocals that restores sound image information.

[0104] In some specific embodiments, the signal interval determination module 12 may specifically include:

[0105] The energy envelope acquisition unit is used to divide the left channel signal and the right channel signal of the song audio into frames respectively, and calculate the energy envelope of each frame signal to obtain the energy envelope of the left channel and the energy envelope of the right channel.

[0106] The energy envelope difference signal determination unit is used to obtain the energy envelope difference signal corresponding to the song audio based on the difference between the energy envelope of the left channel and the energy envelope of the right channel.

[0107] In some specific embodiments, the signal interval determination module 12 may specifically include:

[0108] A threshold acquisition unit is used to acquire a first energy threshold and a second energy threshold; the first energy threshold is greater than the second energy threshold; the first energy threshold is an absolute value, and the first energy threshold corresponds to a first positive energy threshold and a first negative energy threshold; the second energy threshold is an absolute value, and the second energy threshold corresponds to a second positive energy threshold and a second negative energy threshold.

[0109] The extreme left signal interval determination unit is used to locate a first region with energy greater than a first positive energy threshold from the energy envelope difference signal, search along the energy envelope difference signal to both sides with the first region as the center, locate two time points with energy equal to a second positive energy threshold, and take the time period between the two time points as the extreme left signal interval.

[0110] The far-right signal interval determination unit is used to locate a second region with energy less than a first negative energy threshold from the energy envelope difference signal, search along the energy envelope difference signal to both sides with the second region as the center, locate two time points with energy equal to the second negative energy threshold, and take the time period between the two time points as the far-right signal interval.

[0111] In some specific embodiments, the energy envelope acquisition unit may specifically include:

[0112] A smoothing and normalization processing unit is used to smooth and normalize the energy envelope;

[0113] The smoothing normalization process may specifically include:

[0114] The smoothing unit is used to upsample the energy envelope to obtain an upsampled signal sequence, perform temporal convolution on the upsampled signal sequence to obtain a convolutional sequence, and downsample the convolutional sequence to obtain a downsampled signal sequence.

[0115] In some specific embodiments, the lead vocal imaging adjustment module 14 may specifically include:

[0116] The lead vocal left channel signal adjustment unit is used to obtain the updated lead vocal left channel signal based on the difference between the lead vocal left channel signal in the extreme left signal interval and the unrelated signal in the extreme left signal interval.

[0117] The lead vocal right channel signal adjustment unit is used to obtain the updated lead vocal right channel signal based on the difference between the lead vocal right channel signal in the far right signal interval and the unrelated signal in the far right signal interval.

[0118] The dual-channel vocal signal determination unit is used to obtain an updated vocal right channel signal based on the difference between the vocal right channel signal in the far right signal interval and the unrelated signal in the far right signal interval; and to obtain a dual-channel vocal signal based on the updated vocal left channel signal and the updated vocal right channel signal.

[0119] In some specific embodiments, the backing vocal image adjustment module 15 may specifically include:

[0120] The backing vocal right channel signal adjustment unit is used to adjust the backing vocal right channel signal in the far left signal interval to 0, so as to obtain the updated backing vocal right channel signal.

[0121] The backing vocal left channel signal adjustment unit is used to adjust the backing vocal left channel signal in the far right signal interval to 0, so as to obtain the updated backing vocal left channel signal.

[0122] A dual-channel accompaniment signal determination unit is used to obtain a dual-channel accompaniment signal based on the updated right accompaniment signal and the updated left accompaniment signal.

[0123] In some specific embodiments, the audio master and backing vocal separation device may specifically include:

[0124] The decorrelation unit is used to adjust the right channel signal of the backing vocals in the far left signal range and the left channel signal of the backing vocals in the far right signal range to obtain a stereo backing vocal signal, and then perform decorrelation processing on the stereo backing vocal signal to obtain a decorrelation-processed stereo backing vocal signal.

[0125] In some specific embodiments, the decorrelation unit may specifically include:

[0126] The first decorrelation unit is used to perform decorrelation processing by delaying sampling of the dual-channel accompaniment signal;

[0127] Alternatively, a second decorrelation unit is used to perform decorrelation processing by changing the phase of the two-channel accompaniment signal.

[0128] Furthermore, this application also discloses an electronic device, see [link to relevant documentation]. Figure 6 As shown, the content in the figure should not be considered as any limitation on the scope of use of this application.

[0129] Figure 6 This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of this application. Specifically, the electronic device 20 may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the audio main / backing vocal separation method disclosed in any of the foregoing embodiments.

[0130] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.

[0131] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored on it include operating system 221, computer program 222 and data 223 including song audio, etc. The storage method can be temporary storage or permanent storage.

[0132] The operating system 221 manages and controls the various hardware devices on the electronic device 20 and the computer program 222 to enable the processor 21 to perform calculations and processing on the massive amounts of data 223 in the memory 22. The operating system 221 can be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of performing the audio main / backing vocal separation method executed by the electronic device 20 as disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs capable of performing other specific tasks.

[0133] Furthermore, this application also discloses a computer program product, including a computer program, which, when executed by a processor, implements the audio main and backing vocal separation method steps disclosed in any of the foregoing embodiments.

[0134] Furthermore, this application also discloses a computer storage medium storing computer-executable instructions. When the computer-executable instructions are loaded and executed by a processor, they implement the audio main and backing vocal separation method steps disclosed in any of the foregoing embodiments.

[0135] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.

[0136] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0137] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0138] The above provides a detailed description of the audio master-accompaniment separation method, device, storage medium, and program product provided by the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A method for separating main and backing vocals in audio, characterized in that, include: Acquire the song audio and extract the mono-correlated signal and mono-uncorrelated signal from the song audio; Obtain the energy envelope difference signal of the song audio, and determine the extreme left signal interval and the extreme right signal interval based on the energy envelope difference signal; Based on the mono correlated signal, the lead vocal left channel signal and the lead vocal right channel signal are obtained; based on the mono uncorrelated signal, the backing vocal left channel signal and the backing vocal right channel signal are obtained. Using the mono uncorrelated signal, sound image information is added to the lead vocal left channel signal in the far left signal interval and the lead vocal right channel signal in the far right signal interval to obtain a stereo lead vocal signal. Adjust the right channel signal of the backing vocals within the far left signal range, and adjust the left channel signal of the backing vocals within the far right signal range to obtain a dual-channel backing vocal signal; The step of using the mono uncorrelated signal to add sound image information to the lead vocal left channel signal in the far left signal interval and the lead vocal right channel signal in the far right signal interval to obtain a two-channel lead vocal signal includes: Using the unrelated signals within the extreme left signal interval, add sound image information to the vocal left channel signal within the extreme left signal interval to obtain the updated vocal left channel signal. Using the unrelated signals within the extreme right signal interval, add sound image information to the lead vocal right channel signal within the extreme right signal interval to obtain the updated lead vocal right channel signal; Based on the updated lead vocal left channel signal and the updated lead vocal right channel signal, the stereo lead vocal signal is obtained; The process of adjusting the right channel signal of the backing vocals within the far left signal range and adjusting the left channel signal of the backing vocals within the far right signal range to obtain a dual-channel backing vocal signal includes: The right channel signal of the backing vocals in the far left signal interval is adjusted to 0 to obtain the updated right channel signal of the backing vocals. The left channel signal of the backing vocals in the far right signal range is adjusted to 0 to obtain the updated left channel signal of the backing vocals. Based on the updated right channel backing vocal signal and the updated left channel backing vocal signal, a dual-channel backing vocal signal is obtained.

2. The audio main and backing vocal separation method according to claim 1, characterized in that, The step of obtaining the energy envelope difference signal of the song audio includes: The left and right channel signals of the song audio are divided into frames, and the energy envelope of each frame is calculated to obtain the energy envelope of the left channel and the energy envelope of the right channel. Based on the difference between the energy envelope of the left channel and the energy envelope of the right channel, the energy envelope difference signal corresponding to the song audio is obtained.

3. The audio main and backing vocal separation method according to claim 2, characterized in that, The step of determining the extreme left signal interval and the extreme right signal interval based on the energy envelope difference signal includes: Obtain a first energy threshold and a second energy threshold; the first energy threshold is greater than the second energy threshold; the first energy threshold is an absolute value, and the first energy threshold corresponds to a first positive energy threshold and a first negative energy threshold; the second energy threshold is an absolute value, and the second energy threshold corresponds to a second positive energy threshold and a second negative energy threshold; The first region with energy greater than the first positive energy threshold is located from the energy envelope difference signal. The first region is used as the center to search along the energy envelope difference signal to both sides to locate two time points with energy equal to the second positive energy threshold. The time period between the two time points is taken as the extreme left signal interval. The second region with energy less than the first negative energy threshold is located from the energy envelope difference signal. The second region is used as the center to search along the energy envelope difference signal to both sides to locate two time points with energy equal to the second negative energy threshold. The time period between the two time points is taken as the extreme right signal interval.

4. The audio main and backing vocal separation method according to claim 2, characterized in that, The calculation of the energy envelope of each frame of signal to obtain the energy envelope of the left channel and the energy envelope of the right channel includes: Calculate the energy envelope of each frame signal, and perform smoothing and normalization on the energy envelope to obtain the energy envelope of the left channel and the energy envelope of the right channel; Smoothing the energy envelope includes: The energy envelope is upsampled to obtain an upsampled signal sequence. The upsampled signal sequence is then convolved in the time domain to obtain a convolved sequence. Finally, the convolved sequence is downsampled to obtain a downsampled signal sequence.

5. The audio main and backing vocal separation method according to claim 1, characterized in that, The process of adding sound image information to the lead vocal left channel signal in the far left signal range and the lead vocal right channel signal in the far right signal range using the mono uncorrelated signal to obtain a two-channel lead vocal signal includes: The updated vocal left channel signal is obtained by comparing the lead vocal left channel signal within the extreme left signal interval with the unrelated signal within the extreme left signal interval. The updated vocal right channel signal is obtained by comparing the difference between the vocal right channel signal within the far right signal interval and the unrelated signal within the far right signal interval. Based on the updated left channel signal and the updated right channel signal of the lead vocalist, a dual-channel lead vocal signal is obtained.

6. The audio main and backing vocal separation method according to any one of claims 1 to 5, characterized in that, After adjusting the right channel signal of the backing vocals within the far left signal range and adjusting the left channel signal of the backing vocals within the far right signal range to obtain a dual-channel backing vocal signal, the method further includes: The two-channel accompaniment signal is subjected to decorrelation processing to obtain the decorrelation-processed two-channel accompaniment signal.

7. The audio main and backing vocal separation method according to claim 6, characterized in that, The decorrelation processing of the dual-channel accompaniment signal includes: The dual-channel accompaniment signal is delayed and sampled for decorrelation processing. Alternatively, decorrelation processing can be performed by changing the phase of the two-channel accompaniment signal.

8. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the audio lead / backing vocal separation method as described in any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that, Used to store computer programs; wherein the computer programs, when executed by a processor, implement the audio main and backing vocal separation method as described in any one of claims 1 to 7.

10. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the audio main and backing vocal separation method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method and device for processing chorus audio, and storage medium

    CN113192486A

  • Method and its device for extracting accompanying music from songs

    CN1945689A