Sound signal separation method and apparatus, electronic device, and storage medium
By using a method based on short-time energy and cross-correlation coefficients, coherent sound and ambient sound in stereo are separated, solving the problem of separation difficulties in existing technologies and improving the immersion and quality of the sound.
Patent Information
- Application Number
- CN202411973957.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2044-12-30
AI Technical Summary
In existing stereo audio formats, coherent sound and ambient sound are difficult to separate effectively, affecting the spatial positioning accuracy of sound and the spatial sound reproduction effect.
The short-time energy of ambient sound is determined based on the short-time energy and cross-correlation coefficient of each channel signal. The ambient sound mask is then calculated using the short-time energy of the ambient sound and the short-time energy of the channel signals, thereby separating coherent sound from ambient sound.
It achieves effective separation of coherent sound and ambient sound, enhances the immersiveness and quality of the sound, improves the clarity and fidelity of the sound, and adapts to personalized needs.
Smart Images

Figure CN119811415B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of sound processing technology, and in particular to a sound signal separation method, apparatus, electronic device, and storage medium. Background Technology
[0002] Spatial sound reproduction technology is a technique that uses technical means to reproduce the distribution and propagation of sound in three-dimensional space. Its main purpose is to create a realistic sound field around the listener, enabling the listener to perceive the location, distance, and depth of the sound, thereby obtaining a more immersive auditory experience.
[0003] To achieve better spatial sound reproduction, it is necessary to separate coherent sound from ambient sound in stereo audio so that they can be processed separately during post-production sound mixing. However, due to limitations of existing stereo audio formats, coherent sound and ambient sound are often recorded and transmitted together, making effective separation difficult. Therefore, how to effectively separate coherent sound and ambient sound has become a key focus of current research. Summary of the Invention
[0004] This invention provides a sound signal separation method, apparatus, electronic device, and storage medium to address the shortcomings of related technologies in separating coherent sound from ambient sound.
[0005] This invention provides a method for separating sound signals, comprising:
[0006] The short-time energy of the ambient sound is determined based on the short-time energy of each channel signal and the cross-correlation coefficient between the channel signals, and the ambient sound of each channel has the same short-time energy.
[0007] Based on the short-time energy of the ambient sound and the short-time energy of the signals of each channel, the ambient sound mask of each channel is determined;
[0008] Based on the ambient sound mask of each channel and the signal of each channel, the ambient sound and coherent sound of each channel are determined.
[0009] According to a sound signal separation method provided by the present invention, determining the short-time energy of ambient sound based on the short-time energy of each channel signal and the cross-correlation coefficient between the channel signals includes:
[0010] A short-time Fourier transform is performed on the time-domain representation of each channel signal to obtain the time-frequency domain representation of each channel signal;
[0011] Based on the time-frequency domain representation of each channel signal, the linear combination energy representation of each channel signal is determined;
[0012] Based on the time-frequency domain representation of each channel signal, the energy representation of the cross-correlation coefficient between each channel signal is determined;
[0013] The short-time energy of the ambient sound is determined based on the linear combination energy representation, the energy representation of the cross-correlation coefficient, the short-time energy of each channel signal, and the cross-correlation coefficient between each channel signal.
[0014] According to a sound signal separation method provided by the present invention, determining the linear combination energy representation of the channel signals based on the time-frequency domain representation of each channel signal includes:
[0015] Based on the time-frequency domain representation of each channel signal and the equivalence relationship between the short-time energy of the ambient sound of each channel, the linear combination energy representation of each channel signal is determined.
[0016] According to a sound signal separation method provided by the present invention, each channel signal includes a first channel signal and a second channel signal. The time-domain representation of the first channel signal is obtained by superimposing the time-domain representation of the coherent sound and the time-domain representation of the first channel ambient sound. The time-domain representation of the second channel signal is obtained by superimposing the time-domain representation of the coherent sound and the time-domain representation of the second channel ambient sound after multiplying the time-domain representation of the coherent sound with a difference factor. The difference factor is the amplitude difference factor between the first channel coherent sound and the second channel coherent sound.
[0017] According to a sound signal separation method provided by the present invention, determining the ambient sound mask for each channel based on the short-time energy of the ambient sound and the short-time energy of the signals of each channel includes:
[0018] Construct the time-frequency domain representation relationship between the ambient sound of each channel and the signal of each channel;
[0019] Based on the time-frequency domain representation relationship, the linear combination energy representation of each channel signal, and the short-time energy of the ambient sound and the short-time energy of each channel signal, the ambient sound mask for each channel is determined.
[0020] According to a sound signal separation method provided by the present invention, determining the ambient sound and coherent sound of each channel based on the ambient sound mask of each channel and the signal of each channel includes:
[0021] Multiply the ambient sound mask of any channel with the signal of any channel to obtain the ambient sound of any channel;
[0022] Subtract the ambient sound of any channel from the signal of any channel to obtain the coherent sound of any channel.
[0023] The present invention also provides a sound signal separation device, comprising:
[0024] An energy determination unit is used to determine the short-time energy of ambient sound based on the short-time energy of each channel signal and the cross-correlation coefficient between the channel signals, wherein the ambient sound of each channel has the same short-time energy.
[0025] A mask determination unit is used to determine the ambient sound mask for each channel based on the short-time energy of the ambient sound and the short-time energy of the signals of each channel.
[0026] The sound separation unit is used to determine the ambient sound and coherent sound of each channel based on the ambient sound mask of each channel and the signal of each channel.
[0027] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the sound signal separation method as described above.
[0028] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the sound signal separation method as described above.
[0029] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the sound signal separation method as described above.
[0030] The sound signal separation method, apparatus, electronic device, and storage medium provided by this invention can determine the short-time energy of ambient sound based on the short-time energy of each channel signal and the cross-correlation coefficient between the channel signals. Based on the short-time energy of the ambient sound and the short-time energy of each channel signal, the ambient sound mask of each channel can be further determined, that is, the mask of ambient sound in each channel. Thus, based on the ambient sound mask of each channel, ambient sound and coherent sound can be separated from each channel signal, ensuring the effectiveness and reliability of the separation of coherent sound and ambient sound. The entire separation process is simple to calculate, easy to program and implement, and applicable to various stereo sound scenarios. Attached Figure Description
[0031] To more clearly illustrate the technical solutions in this invention or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0032] Figure 1 This is a flowchart illustrating the sound signal separation method provided by the present invention;
[0033] Figure 2 This is a schematic flowchart of the coherent sound and ambient sound separation method based on time-frequency domain mask provided by the present invention;
[0034] Figure 3 This is a schematic diagram of the sound signal separation device provided by the present invention;
[0035] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0036] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0037] To provide users with a more realistic sound experience in immersive systems such as virtual reality, spatial sound reproduction technology is particularly important. To achieve better spatial sound reproduction, coherent sound and ambient sound in stereo sound need to be separated so that they can be processed separately during post-production sound mixing. Coherent sound refers to localizable sound signals, such as the sound signals produced by dialogue, while ambient sound refers to non-localizable sound signals, such as background sound effects.
[0038] However, current spatial sound reproduction technology faces a major challenge: the mixing of coherent sound and ambient sound. In existing stereo audio formats, coherent sound and ambient sound are usually recorded or stored together. Due to their interweaving and interference in the audio signal, it is difficult to completely separate them using simple signal processing methods. This mixing not only affects the spatial positioning accuracy of sound but also limits the effectiveness of spatial sound reproduction technology.
[0039] To address this, the present invention provides a sound signal separation method based on time-frequency domain masking, which can separate mixed sound in stereo into localizable coherent sound (such as dialogue between people) and non-localizable ambient sound (such as background sound effects). The two separated signals can be further processed in post-production to create a more immersive and realistic auditory atmosphere, which helps to achieve flexible spatial sound reproduction, thereby overcoming the above-mentioned defects.
[0040] It should be noted that the significance of separating coherent sound and ambient sound in this invention is: (1) enhancing immersion by accurately locating coherent sound sources and finely controlling ambient sound, making users feel as if they are there; (2) improving sound quality by enhancing the clarity and fidelity of sound; and (3) adapting to personalized needs by customizing for the preferences of different users.
[0041] Figure 1 This is a flowchart illustrating the sound signal separation method provided by the present invention, as shown below. Figure 1 As shown, the method includes:
[0042] Step 110: Determine the short-time energy of the ambient sound based on the short-time energy of each channel signal and the cross-correlation coefficient between the channel signals, wherein the ambient sound of each channel has the same short-time energy.
[0043] It should be noted that each channel signal refers to the sound signal acquired from different channels. For example, each channel signal can be a two-channel signal or a multi-channel signal, and the embodiments of the present invention do not specifically limit this. For ease of understanding, the embodiments of the present invention will mainly use a two-channel signal as an example for introduction.
[0044] Specifically, when each channel signal is a two-channel signal, each channel signal can include a first channel signal and a second channel signal. Here, the first channel signal and the second channel signal refer to the sound signals acquired by the first channel and the second channel respectively. The first channel and the second channel can be the dual channels required to construct a stereo effect, that is, two independent audio channels, left and right. In other words, the first channel signal can be acquired through the left channel and the second channel signal can be acquired through the right channel.
[0045] It is understandable that the sound signals acquired in each channel can be interpreted as the sound signals obtained by superimposing coherent sound and ambient sound. That is, the first channel signal and the second channel signal are both mixed sound signals of coherent sound and ambient sound.
[0046] Furthermore, based on the characteristics of coherent sound and ambient sound, coherent sound is perfectly correlated across different channels, while ambient sound is uncorrelated across different channels. Moreover, coherent sound is uncorrelated with the ambient sound in each channel. Here, perfect correlation means that two or more sound signals have a high degree of similarity and consistency in waveform, frequency, phase, etc., usually implying that these sound signals originate from the same sound source; uncorrelatedness means that two or more sound signals do not have significant similarity and consistency in waveform, frequency, phase, etc., usually implying that these sound signals originate from different sound sources.
[0047] For the first and second channel signals acquired from different channels, time-domain representations of the first and second channel signals can be established. Here, since the sound signal in each channel can be represented as a superposition of coherent sound and ambient sound, the time-domain representation of the first channel signal can be understood as a linear superposition of the time-domain representations of the coherent sound and the ambient sound in the first channel; the time-domain representation of the second channel signal is obtained by a linear superposition of the time-domain representations of the coherent sound and the ambient sound in the second channel.
[0048] Since the coherent sounds between different channels are perfectly correlated, a difference factor can be defined between the coherent sounds in the first channel and the coherent sounds in the second channel. That is, the coherent sounds in the second channel can be represented as a product of the difference factor and the coherent sounds in the first channel. Based on this, the time-domain representation of the second channel signal can be determined by linearly superimposing the time-domain representation of the coherent sounds in the first channel, multiplied by the difference factor, with the time-domain representation of the ambient sound in the second channel.
[0049] After establishing the time-domain representation of each channel signal, a short-time Fourier transform can be performed on each channel to obtain its representation in the Fourier transform domain, i.e., the time-frequency domain representation of each channel signal. Here, the short-time Fourier transform is a mathematical transform related to the Fourier transform, used to determine the frequency and phase of a local sinusoidal wave in a time-varying signal. Unlike the Fourier transform, the short-time Fourier transform takes into account the time-varying nature of the signal. By applying a time window to the signal, it is divided into multiple short time intervals (or frames), and then a Fourier transform is performed on each short time interval. In this way, the spectral information of the signal in different time intervals can be obtained, thus providing a better understanding of the signal's time-frequency characteristics.
[0050] Based on the time-frequency domain representation of each channel signal, a relationship can be established between the short-time energy of each channel signal and the short-time energy of coherent sound and ambient sound in each channel. Here, short-time energy represents the total energy of the sound signal within a time period and can be used to reflect the strength of the sound signal within that time period. For example, when each channel signal is a two-channel signal, based on the time-frequency domain representation of the first and second channel signals, a relationship can be established between the short-time energy of the first channel signal and the short-time energy of ambient sound and coherent sound in the first channel, as well as a relationship between the short-time energy of the second channel signal and the short-time energy of ambient sound and coherent sound in the second channel.
[0051] Since both the first and second channel signals are known sound signals acquired through sampling, the short-time energy values of the first and second channel signals can be directly calculated. By substituting the short-time energy values of the first and second channel signals into the established relationship between short-time energy values, the short-time energy values of ambient sound and coherent sound, as well as the difference factor of coherent sound in different channels, can be obtained.
[0052] It should be understood that, based on the characteristics of ambient sound, the short-time energy of ambient sound collected in different channels is the same. Therefore, the short-time energy value of ambient sound obtained is unique, meaning that the short-time energy value of ambient sound in each channel is consistent. Furthermore, based on the characteristics of coherent sound, the short-time energy of coherent sound collected in different channels is different, and the coherent sound between different channels is completely correlated. Therefore, the short-time energy value of coherent sound obtained can represent the short-time energy value of coherent sound in one channel. Multiplying this short-time energy value of coherent sound by the difference factor between channels yields the short-time energy value of coherent sound in another channel.
[0053] Step 120: Determine the ambient sound mask for each channel based on the short-time energy of the ambient sound and the short-time energy of the signals of each channel.
[0054] Specifically, based on the calculated short-time energy of the ambient sound (i.e., the solved short-time energy value) and the short-time energy of each channel signal (i.e., the known short-time energy value), the ambient sound mask for each channel can be calculated, that is, the mask for the ambient sound under each channel. For example, since the short-time energy of the ambient sound is the same under each channel, the ambient sound mask under the first channel can be calculated based on the short-time energy of the ambient sound and the short-time energy of the first channel signal; the ambient sound mask under the second channel can be calculated based on the short-time energy of the ambient sound and the short-time energy of the second channel signal.
[0055] Here, the ambient sound mask for each channel refers to the mask used to extract or separate ambient sound from the original signal in the sound signal separation task. It can be regarded as a filter or weight matrix that acts on the original signal to extract the ambient sound components of interest. It should be understood that the ambient sound mask for each channel is a quantity related to the autocorrelation and cross-correlation of the signal in each channel. In order to achieve extraction, the ambient sound mask for each channel can be restricted to a positive real number.
[0056] Step 130: Based on the ambient sound mask of each channel and the signal of each channel, determine the ambient sound and coherent sound of each channel.
[0057] Specifically, after calculating the ambient sound mask for each channel, the separated ambient sound can be obtained by multiplying this mask by the channel signal. Then, the separated ambient sound is subtracted from the original channel signal to obtain the coherent sound.
[0058] For example, after calculating the ambient sound mask for the first channel and the ambient sound mask for the second channel, multiplying the ambient sound mask for the first channel by the signal for the first channel yields the ambient sound of the first channel. Subtracting the estimated ambient sound of the first channel from the signal for the first channel yields the coherent sound of the first channel. Similarly, multiplying the ambient sound mask for the second channel by the signal for the second channel yields the ambient sound of the second channel. Subtracting the estimated ambient sound of the second channel from the signal for the second channel yields the coherent sound of the second channel.
[0059] The method provided in this invention determines the short-time energy of ambient sound based on the short-time energy of each channel signal and the cross-correlation coefficient between the channel signals. Based on the short-time energy of the ambient sound and the short-time energy of each channel signal, the ambient sound mask for each channel can be further determined, i.e., the mask for ambient sound in each channel. Therefore, based on the ambient sound mask of each channel, ambient sound and coherent sound can be separated from the channel signals, ensuring the effectiveness and reliability of the separation of coherent sound and ambient sound. The entire separation process is computationally simple, easy to program, and applicable to various stereo sound scenarios.
[0060] Based on any of the above embodiments, each channel signal includes a first channel signal and a second channel signal. The time-domain representation of the first channel signal is obtained by superimposing the time-domain representation of the coherent sound and the time-domain representation of the first channel ambient sound. The time-domain representation of the second channel signal is obtained by superimposing the time-domain representation of the coherent sound and the time-domain representation of the second channel ambient sound after multiplying the time-domain representation of the coherent sound with a difference factor. The difference factor is the amplitude difference factor between the first channel coherent sound and the second channel coherent sound.
[0061] Specifically, assuming the first channel is the left channel and the second channel is the right channel, the time-domain representation of the first channel signal is... Time-domain representation of the second channel signal They are defined as follows:
[0062]
[0063] In the formula, The coherent acoustic component in the left channel. The ambient sound component in the left channel. This is the amplitude difference factor between the coherent acoustic components in the left and right channels. This is the coherent sound from the right channel. The ambient sound is located in the right channel. This represents the time-domain sampling point.
[0064] Understandable, and These are the coherent sound and ambient sound under the first channel, respectively. That is, the signal of the first channel can be obtained by adding the coherent sound and ambient sound under the first channel. and The second channel signal is obtained by adding the coherent sound and ambient sound from the second channel.
[0065] Based on the above embodiments, step 110 specifically includes:
[0066] Step 111: Perform a short-time Fourier transform on the time-domain representation of each channel signal to obtain the time-frequency domain representation of each channel signal.
[0067] It should be noted that in practical applications, the vast majority of sound signals are non-stationary, meaning their frequency components change over time. Traditional time-domain sound signals cannot simultaneously provide time and frequency information, thus failing to effectively process non-stationary signals. The Short-Time Fourier Transform (SFT) captures the time-varying frequency characteristics of the sound signal by dividing the sound signal into multiple short time intervals (windows) and performing a Fourier transform on each interval.
[0068] Specifically, in order to better apply it to non-stationary states and situations where multiple sound sources exist simultaneously, the signal model can be transformed into the short-time Fourier transform domain for processing. That is, by performing a short-time Fourier transform on the time-domain representation of the left and right channel signals in the above formula (1), the representation of the left and right channel signals in the Fourier transform domain (i.e., the frequency domain representation) can be obtained.
[0069]
[0070] in, The time-frequency domain representation of the left channel signal (i.e., the first channel signal). The time-frequency domain representation of the right channel signal (i.e., the second channel signal). This represents a coherent acoustic signal in the time-frequency domain. This represents the ambient sound signal in the time-frequency domain of the left channel. This represents the ambient sound signal in the time-frequency domain of the right channel. Indicates the time frame index. Indicates frequency point index, This represents the difference factor between the coherent sounds of the left and right channels in the time-frequency domain. To maintain the simplicity of the formula, the part in the above formula can be omitted. According to the above formula (2), the parameters can be estimated. , , and This allows for the extraction of coherent sound and ambient sound.
[0071] Step 112: Based on the time-frequency domain representation of each channel signal, determine the linear combination energy representation of each channel signal.
[0072] Specifically, based on the time-frequency domain representation of each channel signal, the short-time energy of the first channel signal can be determined as the sum of the short-time energy of the coherent sound of the first channel and the short-time energy of the ambient sound of the first channel, and the short-time energy of the second channel signal can be determined as the sum of the short-time energy of the coherent sound of the second channel and the short-time energy of the ambient sound of the second channel. Thus, the linear combination energy representations of the first and second channel signals can be obtained.
[0073] Furthermore, step 112 specifically includes:
[0074] Based on the time-frequency domain representation of each channel signal and the equivalence relationship between the short-time energy of the ambient sound of each channel, the linear combination energy representation of each channel signal is determined.
[0075] Specifically, the linear combination energy representation includes the short-time energy representation of the first channel signal and the short-time energy representation of the second channel signal. Taking the first channel signal as the left channel signal as an example, the short-time energy of the first channel signal can be represented as follows:
[0076]
[0077] In the formula, This represents the short-time energy of the first channel signal. This is the time-frequency domain representation of the first channel signal. This represents the short-time average. Similarly, the short-time energy of the second channel signal (i.e., the right channel signal) can be expressed as:
[0078]
[0079] Assuming that the ambient sound of the first channel signal and the second channel signal in formula (2) have the same short-time energy, denoted as The short-time energy of coherent sound is denoted as Based on the short-time energy representations of the first and second channel signals described above, the formula is omitted. Formula (2) above can be expressed as:
[0080]
[0081] In the formula, and These are the short-time energies of the first channel signal and the second channel signal, respectively. This represents the short-time energy of the coherent sound in the first channel. This represents the short-time energy of the coherent sound in the second channel. Let be the short-time energy of the ambient sound. Formula (3) here represents the linear combination energy of the signals from each channel. Considering that the short-time energy of the ambient sound is the same across different channels, the short-time energy of the ambient sound can be expressed as follows in both the linear combination energy representation of the first channel signal and the linear combination energy representation of the second channel signal: .
[0082] Step 113: Based on the time-frequency domain representation of each channel signal, determine the energy representation of the cross-correlation coefficient between each channel signal.
[0083] Specifically, after obtaining the time-frequency domain representations of the first channel signal and the second channel signal, the cross-correlation coefficient between the first channel signal and the second channel signal can be calculated. During the calculation of the cross-correlation coefficient, the time-frequency domain representations of the first channel signal and the second channel signal can be substituted into the calculation formula for the cross-correlation coefficient. The resulting cross-correlation coefficient is expressed in the form of short-time energy of the sound signal. In this embodiment of the invention, the cross-correlation coefficient expressed in the form of short-time energy is denoted as the energy representation of the cross-correlation coefficient.
[0084] Here, the formula for calculating the cross-correlation coefficient between the first channel signal and the second channel signal can be:
[0085]
[0086] In the formula, For cross-correlation coefficients, and These are the time-frequency domain representations of the first channel signal and the second channel signal, respectively.
[0087] Substituting formula (2) into formula (4), we get the following formula:
[0088]
[0089] Since the coherent sound is uncorrelated with the ambient sound in each channel, , Furthermore, since the ambient sound in different channels is uncorrelated, therefore .
[0090] Therefore, the formula for calculating the cross-correlation coefficient can be expressed in the following form, which is the energy representation of the cross-correlation coefficient:
[0091]
[0092] In the formula, the cross-relation number It can be determined by the short-time energy of the first channel signal. Short-time energy of the second channel signal Short-time energy of coherent sound and the time-frequency domain representation of the difference factor express.
[0093] Step 114: Determine the short-time energy of the ambient sound based on the linear combination energy representation, the energy representation of the cross-correlation coefficient, the short-time energy of each channel signal, and the cross-correlation coefficient between each channel signal.
[0094] Specifically, since both the first and second channel signals are known sound signals, the short-time energy values of the first and second channel signals can be directly calculated, and the cross-correlation coefficient between the first and second channel signals can be calculated. Therefore, the short-time energy values of the first and second channel signals, as well as the cross-correlation coefficient between them, can be substituted into the above-mentioned linear combination energy representation and cross-correlation coefficient energy representation, i.e., substituted into formulas (3) and (5), to solve for the short-time energy values and difference factors of the ambient sound and coherent sound in the above-mentioned linear combination energy representation and cross-correlation coefficient energy representation. This yields the short-time energy values of the ambient sound and coherent sound, as well as the difference factor values of the coherent sound in different channels.
[0095] Specifically, in formulas (3) and (5) above, and It is about and The function of ambient sound can be used to solve for the short-time energy of ambient sound in the following form. and the short-time energy of coherent sound and differential factors :
[0096]
[0097] in,
[0098]
[0099] Understandably, since both the first and second channel signals are known sound signals, the short-time energy of the first channel signal... Short-time energy of the second channel signal and cross-correlation coefficients This is known. Based on this, the parameters can be calculated. and The value of this value allows for further calculation of the short-time energy of the ambient sound. Short-time energy of coherent sound and differential factors .
[0100] The method provided in this invention, based on the linear combination relationship between the ambient sound and coherent sound contained in the first channel signal and the second channel signal, realizes the short-time energy solution of the ambient sound and coherent sound, thus providing conditions for achieving sound separation.
[0101] Based on any of the above embodiments, step 120 specifically includes:
[0102] Step 121: Construct the time-frequency domain representation relationship between the ambient sound of each channel and the signal of each channel.
[0103] Specifically, since the short-time energy of the ambient sound in the first and second channels is the same, the ambient sound of each channel can be obtained by multiplying the ambient sound mask of that channel with the sound signal of that channel. Based on this, the time-frequency domain representation relationship between the ambient sound of the first channel and the signal of the first channel, as well as the time-frequency domain representation relationship between the ambient sound of the second channel and the signal of the second channel, can be constructed:
[0104]
[0105] In the formula, This is the time-frequency domain representation of the ambient sound in the first channel. It is the ambient sound mask for the first channel. It is the time-frequency domain representation of the first channel signal. This is the time-frequency domain representation of the ambient sound in the second channel. It is the ambient sound mask for the second channel. This is the time-frequency domain representation of the second channel signal. Among them, the ambient sound mask... , It is a quantity related to the autocorrelation and cross-correlation of the channel signal. To achieve extraction, it can be... , The numbers are restricted to positive real numbers.
[0106] Step 122: Based on the time-frequency domain representation relationship, the linear combination energy representation of each channel signal, and the short-time energy of the ambient sound and the short-time energy of each channel signal, determine the ambient sound mask for each channel.
[0107] Specifically, in the above formula (3), since the ambient sound of each channel has the same short-time energy, and is denoted as Therefore, formula (7) can be transformed into:
[0108]
[0109] in, and These are the short-time energies of the first and second channel signals, respectively, both of which are known. This is the short-time energy of the ambient sound, which has been solved in formula (6). Therefore, the ambient sound masks for the first and second channels can be calculated according to formula (8). and .
[0110] Based on any of the above embodiments, step 130 specifically includes:
[0111] Step 131: Multiply the ambient sound mask of any channel with the signal of any channel to obtain the ambient sound of any channel;
[0112] Step 132: Subtract the ambient sound of any channel from the signal of any channel to obtain the coherent sound of any channel.
[0113] Specifically, for the first channel, the ambient sound of the first channel can be obtained by multiplying the ambient sound mask of the first channel with the signal of the first channel, thereby completing the extraction of the ambient sound component from the signal of the first channel. Subtracting the ambient sound of the first channel from the signal of the first channel yields the remaining component of the signal of the first channel. This remaining component is used as the coherent sound component of the first channel, thus completing the separation and extraction of the ambient sound and coherent sound of the first channel.
[0114] Similarly, for the second channel, by multiplying the ambient sound mask of the second channel with the signal of the second channel, the ambient sound of the second channel can be obtained, thus realizing the extraction of the ambient sound component from the signal of the second channel. Subtracting the ambient sound of the second channel from the signal of the second channel yields the remaining components of the signal of the second channel. These remaining components are the coherent sound components of the second channel, thereby completing the separation and extraction of the ambient sound and coherent sound of the second channel.
[0115] Based on any of the above embodiments Figure 2 This is a schematic flowchart of the coherent sound and ambient sound separation method based on time-frequency domain mask provided by the present invention, as shown below. Figure 2 As shown, the core idea of this invention is to assume that the short-time energy of the ambient sound in the left and right channels is the same. By solving three simultaneous equations, the short-time energy of the ambient sound can be calculated. Then, the masking effect of the ambient sound in each channel is further calculated, and multiplied by the original channel signal to obtain the ambient sound. Subtracting the estimated ambient sound from the original channel signal yields the coherent sound, which has strong theoretical interpretability. The method includes the following steps:
[0116] Step S1: Obtain the time-domain input signal of the stereo left channel. and the time-domain input signal of the right channel .
[0117] Step S2: Perform short-time Fourier transforms on the time-domain input signals of the left and right channels respectively to obtain the representations of the left and right channel signals in the Fourier transform domain. and .
[0118] Step S3: Calculate the short-time energy of the left and right channel signals based on their representations in the Fourier transform domain. , and cross-correlation coefficients .
[0119] Step S4, based on the short-time energy of the left and right channel signals , and cross-correlation coefficients Calculate the short-time energy of ambient sound .
[0120] Step S5, based on the short-time energy of ambient sound and the short-time energy of the left and right channel signals , Calculate the ambient sound mask for the left and right channels in the time-frequency domain. , .
[0121] Step S6: Multiply the ambient sound mask of each channel with the signal of that channel to obtain the ambient sound signal of each channel. The remaining signal is the coherent sound signal.
[0122] The method provided in this invention features a relatively simple sound separation calculation process that is easy to program. This method exhibits good adaptability to different types of sound signals and can be adapted to various stereo scenarios. Furthermore, this method achieves superior sound separation results, effectively separating coherent sound from ambient sound.
[0123] The sound separation device provided by the present invention will be described below. The sound separation device described below can be referred to in correspondence with the sound separation method described above.
[0124] Based on any of the above embodiments Figure 3 This is a schematic diagram of the sound signal separation device provided by the present invention, as shown below. Figure 3 As shown, the device includes:
[0125] The energy determination unit 310 is used to determine the short-time energy of the ambient sound based on the short-time energy of each channel signal and the cross-correlation coefficient between the channel signals, wherein the ambient sound of each channel has the same short-time energy.
[0126] The mask determination unit 320 is used to determine the ambient sound mask for each channel based on the short-time energy of the ambient sound and the short-time energy of the signals of each channel.
[0127] The sound separation unit 330 is used to determine the ambient sound and coherent sound of each channel based on the ambient sound mask of each channel and the signal of each channel.
[0128] The apparatus provided in this invention can determine the short-time energy of ambient sound based on the short-time energy of each channel signal and the cross-correlation coefficient between the channel signals. Based on the short-time energy of the ambient sound and the short-time energy of each channel signal, the ambient sound mask of each channel can be further determined, that is, the mask of ambient sound in each channel. Thus, based on the ambient sound mask of each channel, ambient sound and coherent sound can be separated from each channel signal, ensuring the effectiveness and reliability of the separation of coherent sound and ambient sound. The entire separation process is simple to calculate, easy to program and implement, and applicable to various stereo sound scenarios.
[0129] Based on any of the above embodiments, the energy determination unit 310 includes:
[0130] The transform subunit is used to perform a short-time Fourier transform on the time-domain representation of each channel signal to obtain the time-frequency domain representation of each channel signal;
[0131] The first energy representation subunit is used to determine the linear combination energy representation of the channel signals based on the time-frequency domain representation of the channel signals;
[0132] The second energy representation subunit is used to determine the energy representation of the cross-correlation coefficient between the channel signals based on the time-frequency domain representation of each channel signal;
[0133] The short-time energy determination subunit is used to determine the short-time energy of the ambient sound based on the linear combination energy representation, the energy representation of the cross-correlation coefficient, the short-time energy of each channel signal, and the cross-correlation coefficient between the channel signals.
[0134] Based on any of the above embodiments, the first energy representation subunit is specifically used for:
[0135] Based on the time-frequency domain representation of each channel signal and the equivalence relationship between the short-time energy of the ambient sound of each channel, the linear combination energy representation of each channel signal is determined.
[0136] Based on any of the above embodiments, each channel signal includes a first channel signal and a second channel signal. The time-domain representation of the first channel signal is obtained by superimposing the time-domain representation of the coherent sound and the time-domain representation of the first channel ambient sound. The time-domain representation of the second channel signal is obtained by superimposing the time-domain representation of the coherent sound and the time-domain representation of the second channel ambient sound after multiplying the time-domain representation of the coherent sound with a difference factor. The difference factor is the amplitude difference factor between the first channel coherent sound and the second channel coherent sound.
[0137] Based on any of the above embodiments, the mask determination unit 320 is specifically used for:
[0138] Construct the time-frequency domain representation relationship between the ambient sound of each channel and the signal of each channel;
[0139] Based on the time-frequency domain representation relationship, the linear combination energy representation of each channel signal, and the short-time energy of the ambient sound and the short-time energy of each channel signal, the ambient sound mask for each channel is determined.
[0140] Based on any of the above embodiments, the sound separation unit 330 is specifically used for:
[0141] Multiply the ambient sound mask of any channel with the signal of any channel to obtain the ambient sound of any channel;
[0142] Subtract the ambient sound of any channel from the signal of any channel to obtain the coherent sound of any channel.
[0143] Figure 4 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 4 As shown, the electronic device may include a processor 410, a communication interface 420, a memory 430, and a communication bus 440, wherein the processor 410, the communication interface 420, and the memory 430 communicate with each other through the communication bus 440. The processor 410 can call logical instructions in the memory 430 to execute a sound signal separation method, which includes: determining the short-time energy of ambient sound based on the short-time energy of each channel signal and the cross-correlation coefficient between the channel signals, wherein the ambient sound of each channel has the same short-time energy; determining the ambient sound mask of each channel based on the short-time energy of the ambient sound and the short-time energy of each channel signal; and determining the ambient sound and coherent sound of each channel based on the ambient sound mask of each channel and the channel signal.
[0144] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to related technologies, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0145] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the sound signal separation method provided by the above methods. The method includes: determining the short-time energy of ambient sound based on the short-time energy of each channel signal and the cross-correlation coefficient between the channel signals, wherein the ambient sound of each channel has the same short-time energy; determining the ambient sound mask of each channel based on the short-time energy of the ambient sound and the short-time energy of each channel signal; and determining the ambient sound and coherent sound of each channel based on the ambient sound mask of each channel and the channel signal.
[0146] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the sound signal separation method provided by the above methods. The method includes: determining the short-time energy of ambient sound based on the short-time energy of each channel signal and the cross-correlation coefficient between the channel signals, wherein the ambient sound of each channel has the same short-time energy; determining the ambient sound mask of each channel based on the short-time energy of the ambient sound and the short-time energy of each channel signal; and determining the ambient sound and coherent sound of each channel based on the ambient sound mask of each channel and the channel signal.
[0147] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0148] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of software products. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0149] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for separating sound signals, characterized in that, include: The short-time energy of the ambient sound is determined based on the short-time energy of each channel signal and the cross-correlation coefficient between the channel signals, and the ambient sound of each channel has the same short-time energy. Based on the short-time energy of the ambient sound and the short-time energy of the signals of each channel, the ambient sound mask of each channel is determined; Based on the ambient sound mask of each channel and the signal of each channel, the ambient sound and coherent sound of each channel are determined; The determination of the short-time energy of ambient sound based on the short-time energy of each channel signal and the cross-correlation coefficient between the channel signals includes: A short-time Fourier transform is performed on the time-domain representation of each channel signal to obtain the time-frequency domain representation of each channel signal; Based on the time-frequency domain representation of each channel signal, the linear combination energy representation of each channel signal is determined; Based on the time-frequency domain representation of each channel signal, the energy representation of the cross-correlation coefficient between each channel signal is determined; The short-time energy of the ambient sound is determined based on the linear combination energy representation, the energy representation of the cross-correlation coefficient, the short-time energy of each channel signal, and the cross-correlation coefficient between each channel signal.
2. The sound signal separation method according to claim 1, characterized in that, The determination of the linear combination energy representation of the signals in each channel based on the time-frequency domain representation of each channel signal includes: Based on the time-frequency domain representation of each channel signal and the equivalence relationship between the short-time energy of the ambient sound of each channel, the linear combination energy representation of each channel signal is determined.
3. The sound signal separation method according to claim 1, characterized in that, Each channel signal includes a first channel signal and a second channel signal. The time-domain representation of the first channel signal is obtained by superimposing the time-domain representation of the coherent sound and the time-domain representation of the first channel ambient sound. The time-domain representation of the second channel signal is obtained by superimposing the time-domain representation of the coherent sound and the time-domain representation of the second channel ambient sound after multiplying the time-domain representation of the coherent sound with a difference factor. The difference factor is the amplitude difference factor between the first channel coherent sound and the second channel coherent sound.
4. The sound signal separation method according to claim 1, characterized in that, The determination of the ambient sound mask for each channel based on the short-time energy of the ambient sound and the short-time energy of the signals from each channel includes: Construct the time-frequency domain representation relationship between the ambient sound of each channel and the signal of each channel; Based on the time-frequency domain representation relationship, the linear combination energy representation of each channel signal, and the short-time energy of the ambient sound and the short-time energy of each channel signal, the ambient sound mask for each channel is determined.
5. The sound signal separation method according to any one of claims 1 to 4, characterized in that, The determination of ambient sound and coherent sound for each channel based on the ambient sound mask and the signals of each channel includes: Multiply the ambient sound mask of any channel with the signal of any channel to obtain the ambient sound of any channel; Subtract the ambient sound of any channel from the signal of any channel to obtain the coherent sound of any channel.
6. A sound signal separation device, characterized in that, include: An energy determination unit is used to determine the short-time energy of ambient sound based on the short-time energy of each channel signal and the cross-correlation coefficient between the channel signals, wherein the ambient sound of each channel has the same short-time energy. A mask determination unit is used to determine the ambient sound mask for each channel based on the short-time energy of the ambient sound and the short-time energy of the signals of each channel. A sound separation unit is used to determine the ambient sound and coherent sound of each channel based on the ambient sound mask of each channel and the signal of each channel; The energy determination unit is specifically used for: A short-time Fourier transform is performed on the time-domain representation of each channel signal to obtain the time-frequency domain representation of each channel signal; Based on the time-frequency domain representation of each channel signal, the linear combination energy representation of each channel signal is determined; Based on the time-frequency domain representation of each channel signal, the energy representation of the cross-correlation coefficient between each channel signal is determined; The short-time energy of the ambient sound is determined based on the linear combination energy representation, the energy representation of the cross-correlation coefficient, the short-time energy of each channel signal, and the cross-correlation coefficient between each channel signal.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the sound signal separation method as described in any one of claims 1 to 5.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the sound signal separation method as described in any one of claims 1 to 5.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the sound signal separation method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Filtering summation multi-channel voice separation method based on simplified attention coding and decoding network
CN115910092A
Array microphone noise reduction method and device based on blind source separation
CN118335097A