Sound separation method, apparatus, electronic device, and storage medium
By acquiring sound signals from different channels, and based on linear combination relationships and weight mapping, the ambient sound and coherent sound in stereo audio are separated, solving the problem of difficult separation in stereo audio formats and achieving effective spatial sound reproduction.
Patent Information
- Application Number
- CN202411779092.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-05
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2044-12-05
AI Technical Summary
Due to the limitations of stereo audio formats, coherent sound and ambient sound are recorded and transmitted together, making it difficult to separate them in the later stages and achieve effective spatial sound reproduction.
By acquiring first and second sound signals from different channels, the short-time energy values and difference factors of ambient sound and coherent sound are determined based on a linear combination relationship. A weighted mapping relationship is established, and linear fitting is performed to separate ambient sound and coherent sound.
It achieves effective separation of ambient sound and coherent sound, ensuring the effectiveness and reliability of the separation. It is computationally simple, adaptable to various stereo scenarios, and easy to program.
Smart Images

Figure CN119741935B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of sound processing, and in particular to a sound separation method and device, electronic equipment and a storage medium. BACKGROUND
[0002] In order to optimize the hearing effect of users, spatial sound playback technology has emerged.
[0003] A core link of the spatial sound playback technology is to separate the coherent sound and the ambient sound in the mixed sound, so that the coherent sound and the ambient sound can be processed respectively in the later tuning, thereby creating a more immersive and immersive auditory atmosphere.
[0004] However, under the limitation of the current stereo audio format, the coherent sound and the ambient sound are often mixed together for recording and transmission, which makes it difficult to be completely distinguished when separated later, which limits the implementation effect of the spatial sound playback technology to some extent. SUMMARY
[0005] The present application provides a sound separation method, device, electronic equipment and storage medium to solve the defect of difficulty in separating coherent sound and ambient sound in related technologies.
[0006] The present application provides a sound separation method, comprising:
[0007] Obtaining a first sound signal and a second sound signal, the first sound signal and the second sound signal being obtained by different channels;
[0008] Based on the short-time energy value of the first sound signal and the second sound signal respectively, and the linear combination relationship between the ambient sound and the coherent sound contained in the first sound signal and the second sound signal respectively, the short-time energy value of the ambient sound and the coherent sound respectively, and the difference factor of the coherent sound in the different channels are determined;
[0009] Based on the weight mapping relationship, the short-time energy value of the ambient sound and the coherent sound respectively, and the difference factor are mapped to the sound separation weight, and the weight mapping relationship is obtained by linear fitting based on the linear combination relationship;
[0010] Based on the sound separation weight, the ambient sound and the coherent sound are separated from the first sound signal and the second sound signal.
[0011] According to the sound separation method provided by the application, the short-time energy values of the ambient sound and the coherent sound are determined based on the short-time energy values of the first sound signal and the second sound signal respectively, and the linear combination relationship between the ambient sound and the coherent sound contained in the first sound signal and the second sound signal respectively, and the difference factor of the coherent sound on the different channels, comprising:
[0012] The linear combination energy representation of the first sound signal and the second sound signal is determined based on the linear combination relationship;
[0013] The energy representation of the cross-correlation coefficient between the first sound signal and the second sound signal is determined based on the linear combination relationship;
[0014] The short-time energy values of the ambient sound and the coherent sound are determined, and the difference factor of the coherent sound on the different channels is determined based on the linear combination energy representation and the energy representation of the cross-correlation coefficient, the short-time energy values of the first sound signal and the second sound signal respectively, and the cross-correlation coefficient between the first sound signal and the second sound signal.
[0015] According to the sound separation method provided by the application, the linear combination energy representation of the first sound signal and the second sound signal is determined based on the linear combination relationship, comprising:
[0016] The linear combination energy representation of the first sound signal and the second sound signal is determined based on the linear combination relationship and the equal relationship between the short-time energy values of the ambient sound under the different channels.
[0017] According to the sound separation method provided by the application, the determination of the weight mapping relationship comprises:
[0018] The linear composition relationship of the estimated sound component is constructed, the estimated sound component is the estimated ambient sound or the estimated coherent sound, and the linear composition relationship comprises the initial weight of the first sound signal and the second sound signal respectively;
[0019] The estimated error of the estimated sound component is determined based on the linear composition relationship and the linear combination relationship;
[0020] The initial weight in the linear composition relationship is linearly fitted to obtain the weight mapping relationship, with the estimated error being irrelevant to the first sound signal and the second sound signal respectively.
[0021] According to the sound separation method provided by the application, the initial weight in the linear combination relationship is linearly fitted with the target of making the estimation error irrelevant to the first sound signal and the second sound signal respectively, so as to obtain the weight mapping relationship, which comprises:
[0022] A first cross-correlation function between the estimation error and the first sound signal and a second cross-correlation function between the estimation error and the second sound signal are constructed;
[0023] The estimation error and the linear combination relationship are substituted into the first cross-correlation function and the second cross-correlation function with the target of making the values of the first cross-correlation function and the second cross-correlation function both zero, so as to obtain the weight mapping relationship.
[0024] According to the sound separation method provided by the application, the linear combination relationship comprises that the coherent sound and the ambient sound of the first channel are added to obtain the first sound signal, and the coherent sound is multiplied by the difference factor and then added with the ambient sound of the second channel to obtain the second sound signal.
[0025] The application further provides a sound separation device, which comprises:
[0026] A signal acquisition unit is configured to acquire a first sound signal and a second sound signal, wherein the first sound signal and the second sound signal are collected by different channels;
[0027] A parameter acquisition unit is configured to determine the short-time energy value of the ambient sound and the coherent sound respectively and the difference factor of the coherent sound on the different channels based on the short-time energy value of the first sound signal and the second sound signal respectively and a linear combination relationship between the ambient sound and the coherent sound contained in the first sound signal and the second sound signal;
[0028] A weight mapping unit is configured to map the short-time energy value of the ambient sound and the coherent sound respectively and the difference factor into sound separation weights based on a weight mapping relationship, wherein the weight mapping relationship is obtained by linear fitting based on the linear combination relationship;
[0029] A sound separation unit is configured to separate the ambient sound and the coherent sound from the first sound signal and the second sound signal based on the sound separation weights.
[0030] The application further provides an electronic device, which comprises a memory, a processor, and a computer program stored in the memory and capable of running on the processor, wherein the processor implements the sound separation method according to any one of the above-mentioned sound separation methods when executing the program.
[0031] The application further provides a non-transitory computer-readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the sound separation method.
[0032] The application further provides a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the sound separation method.
[0033] The sound separation method, device, electronic equipment and storage medium provided by the application establish a linear combination relationship between the ambient sound and coherent sound contained in the first sound signal and the second sound signal respectively, perform linear fitting on the weight mapping relationship required for sound separation based on the linear combination relationship, so that the sound separation weight required for sound separation of the first sound signal and the second sound signal can be determined based on the weight mapping relationship in actual application, and then the separation of the ambient sound and the coherent sound is realized, thereby guaranteeing the effectiveness and reliability of the separation of the ambient sound and the coherent sound, and the sound separation process is simple in calculation and easy to program and implement, and can be adapted to various stereo sound scenes. BRIEF DESCRIPTION OF DRAWINGS
[0034] In order to more clearly illustrate the technical solutions in the application or the related art, the following will briefly introduce the drawings needed to be used in the embodiments or the related art description. Obviously, the drawings in the following description are some embodiments of the application, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.
[0035] Figure 1 is a flowchart of the sound separation method provided by the application.
[0036] Figure 2 is a flowchart of the short-time energy value determination method provided by the application.
[0037] Figure 3 is a flowchart of the determination method of the weight mapping relationship provided by the application.
[0038] Figure 4 is a structural schematic diagram of the sound separation device provided by the application.
[0039] Figure 5 is a structural schematic diagram of the electronic equipment provided by the application. DETAILED DESCRIPTION
[0040] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below in conjunction with the drawings in the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0041] The spatial sound playback technology, also known as three-dimensional audio technology, is a technology capable of constructing a virtual auditory space, which enables the listener to have a specific spatial perception. With the development of artificial intelligence audio-visual technology, constructing an immersive virtual auditory space has gradually become a key component to realize realistic virtual reality perception, and the spatial sound playback technology has been applied in the fields of audio-entertainment and voice communication.
[0042] In the spatial sound playback technology, the coherent sound and the environmental sound in the mixed sound are separated and extracted, and then the two are rendered respectively, which can effectively optimize the effect of spatial sound playback. However, due to the limitation of the current stereo audio format, the coherent sound and the environmental sound are often mixed together, which makes it difficult to be completely distinguished during the later separation. Therefore, how to realize the effective separation of the coherent sound and the environmental sound has become a problem to be solved at present.
[0043] Figure 1 is a flowchart of the sound separation method provided by the present application, as shown in Figure 1 The sound separation method provided by the present application can be used for the separation of the coherent sound and the environmental sound, and the method comprises the following steps:
[0044] In step 110, a first sound signal and a second sound signal are obtained, wherein the first sound signal and the second sound signal are obtained by different channels.
[0045] Specifically, the first sound signal and the second sound signal are two sound signals obtained by different channels. Here, the different channels can also be understood as two independent audio channels required to construct a stereo sound effect, i.e. left and right channels. For example, the first sound signal can be obtained by a left channel, and the second sound signal can be obtained by a right channel.
[0046] In the embodiments of the present application, the sound signal obtained by the channel can be understood as a sound signal obtained by superimposing the coherent sound and the environmental sound, i.e. the first sound signal and the second sound signal are both mixed sound signals of the coherent sound and the environmental sound. Here, the coherent sound refers to a locatable sound signal, for example, a sound signal generated by a dialogue between characters, and the environmental sound refers to a non-locatable sound signal, for example, a background sound effect.
[0047] Further, according to the characteristics of the coherent sound and the ambient sound, the coherent sound is fully correlated among different channels, the ambient sound is not correlated among different channels, and the coherent sound is not correlated with the ambient sound in each channel. Here, fully correlated means that two or more sound signals have high similarity and consistency in waveform, frequency, phase, etc., which usually means that the sound signals come from the same sound source; not correlated means that two or more sound signals have no significant similarity and consistency in waveform, frequency, phase, etc., which usually means that the sound signals come from different sound sources.
[0048] In step 120, based on the short-time energy values of the first sound signal and the second sound signal respectively, and the linear combination relationship between the ambient sound and the coherent sound contained in the first sound signal and the second sound signal respectively, the short-time energy values of the ambient sound and the coherent sound respectively, and the difference factor of the coherent sound on the different channels are determined.
[0049] Specifically, for the first sound signal and the second sound signal collected in different channels, a linear combination relationship for the first sound signal and the second sound signal can be established.
[0050] Here, the linear combination relationship can be understood as that the first sound signal is a linear superposition of the coherent sound and the ambient sound in the channel corresponding to the first sound signal, and the second sound signal is a linear superposition of the coherent sound and the ambient sound in the channel corresponding to the second sound signal. For ease of description, in the embodiments of the present application, the channel corresponding to the first sound signal is referred to as the first channel, and the channel corresponding to the second sound signal is referred to as the second channel.
[0051] And since the coherent sound is fully correlated among different channels, a difference factor between the coherent sound in the first channel and the coherent sound in the second channel can be defined, that is, the coherent sound in the second channel can be expressed in the form of multiplying the coherent sound in the first channel by the difference factor.
[0052] Therefore, in the linear combination relationship, the first sound signal can be expressed in the form of linearly adding the coherent sound and the ambient sound in the first channel, and the second sound signal can be expressed in the form of linearly adding the coherent sound and the ambient sound in the second channel.
[0053] After the linear combination relationship is determined, the linear combination relationship can be converted to the frequency domain to calculate the relationship between the short-time energy of the first sound signal and the short-time energy of the ambient sound and the short-time energy of the coherent sound in the first channel, and the relationship between the short-time energy of the second sound signal and the short-time energy of the ambient sound and the short-time energy of the coherent sound in the second channel under the linear combination relationship. Here, conversion to the frequency domain can be achieved by Fourier transformation; the short-time energy represents the total energy of the sound signal in a time period, and can be used to reflect the intensity of the sound signal in the time period.
[0054] Since the first sound signal and the second sound signal are both known sound signals, the short-time energy value of the first sound signal and the short-time energy value of the second sound signal can be directly calculated, and the short-time energy values of the first sound signal and the second sound signal are substituted into the above relationship between the short-time energies to solve, thereby obtaining the short-time energy value of the ambient sound and the short-time energy value of the coherent sound, and the difference factor of the coherent sound in different channels.
[0055] It can be understood that, according to the characteristics of the ambient sound, the short-time energy of the ambient sound collected in different channels is consistent, so the short-time energy of the ambient sound in the first channel and the second channel is consistent in step 120, and the short-time energy value of the ambient sound obtained by solving is unique, and can represent the short-time energy value of the ambient sound in the first sound signal and the short-time energy value of the ambient sound in the second sound signal.
[0056] In addition, according to the characteristics of the coherent sound, the short-time energy of the coherent sound collected in different channels is different, and the coherent sound between different channels is completely related, so the short-time energy value of the coherent sound obtained by solving in step 120 can represent the short-time energy value of the coherent sound in one of the channels, and the short-time energy value of the coherent sound combined with the difference factor between channels can represent the short-time energy value of the coherent sound in the other channel.
[0057] Step 130, based on the weight mapping relationship, mapping the short-time energy value of the ambient sound and the short-time energy value of the coherent sound, and the difference factor, to the sound separation weight, the weight mapping relationship is obtained by linear fitting based on the linear combination relationship.
[0058] Specifically, to achieve sound separation for the first sound signal and the second sound signal, a linear composition relationship for the coherent sound or for the ambient sound can be established. Taking the coherent sound as an example, the linear composition relationship for the coherent sound can represent that the coherent sound is obtained by weighted summation of the first sound signal and the second sound signal, and the weight used for the weighted summation is the sound separation weight required for separating the coherent sound and the ambient sound, which is denoted as sound separation weight in the embodiment of the present application.
[0059] After the linear combination relationship is determined, the sound separation weight required for separating the coherent sound and the ambient sound from the first sound signal and the second sound signal can also be linearly fitted based on the linear combination relationship, so as to obtain a weight mapping relationship capable of reflecting the relationship between the respective short-time energy of the coherent sound and the ambient sound and the weight.
[0060] It can be understood that the weight mapping relationship is a function of the respective short-time energy of the ambient sound and the coherent sound and the difference factor as independent variables, and the sound separation weight as dependent variable. Thus, after the respective short-time energy of the ambient sound and the coherent sound and the difference factor are obtained through step 120, the respective short-time energy of the ambient sound and the coherent sound and the difference factor can be substituted into the weight mapping relationship, so as to obtain the value of the sound separation weight.
[0061] Step 140, separating the ambient sound and the coherent sound from the first sound signal and the second sound signal based on the sound separation weight.
[0062] Specifically, after the sound separation weight is obtained, the first sound signal and the second sound signal can be subjected to sound separation with respect to the sound separation weight. For example, in the case that the sound separation weight is the sound separation weight with respect to the coherent sound, the first sound signal and the second sound signal can be weighted and summed based on the sound separation weight, and the result of the weighted sum can be taken as the separated coherent sound, and the part of the first sound signal and the second sound signal subtracted from the coherent sound can be taken as the ambient sound.
[0063] In the method provided in the embodiments of the present application, by establishing the linear combination relationship between the ambient sound and the coherent sound contained in the first sound signal and the second sound signal respectively, and linearly fitting the weight mapping relationship required for sound separation based on the linear combination relationship, the sound separation weight required for sound separation of the first sound signal and the second sound signal can be determined based on the weight mapping relationship in actual application, so as to realize the separation of the ambient sound and the coherent sound, thereby ensuring the effectiveness and reliability of the separation of the ambient sound and the coherent sound, and the calculation in the sound separation process is simple and easy to program, and can be adapted to various stereo sound scenes.
[0064] Based on the above embodiments, the linear combination relationship includes that the first sound signal is obtained by adding the coherent sound and the ambient sound of the first channel, and the second sound signal is obtained by adding the coherent sound multiplied by the difference factor and the ambient sound of the second channel.
[0065] Specifically, assuming that the first channel is the left channel and the second channel is the right channel, the linear combination relationship can be expressed in the following form:
[0066]
[0067]
[0068] wherein, and are time-domain input signals of left and right channels respectively. is coherent sound under the left channel, is ambient sound under the left channel, is an amplitude difference factor between coherent sounds under the left and right channels, is coherent sound under the right channel, is ambient sound under the right channel.
[0069] wherein, may represent a first sound signal, may represent a second sound signal. and are coherent and ambient sounds under the first channel, i.e., the coherent and ambient sounds under the first channel are added to obtain the first sound signal. and are coherent and ambient sounds under the second channel, i.e., the coherent and ambient sounds under the second channel are added to obtain the second sound signal.
[0070] According to the above formulae, the first sound signal and the second sound signal can each be represented as a linear combination of coherent and ambient sounds, i.e., the first sound signal and the second sound signal can each be represented as a linear combination of ambient and coherent sounds.
[0071] Based on any of the above embodiments, Figure 2 is a flowchart of a short-time energy value determination method provided by the present application, as shown in Figure 2 step 120, determining, based on the linear combination relationship between the ambient sound and the coherent sound included in each of the first sound signal and the second sound signal, the short-time energy value of each of the ambient sound and the coherent sound, and the difference factor of the coherent sound on the different channels, comprising:
[0072] step 121, determining, based on the linear combination relationship, a linear combination energy representation of each of the first sound signal and the second sound signal.
[0073] Specifically, in the case that the linear combination relationship is obtained, the linear combination relationship can be converted to the frequency domain, and the relationship between the short-time energy of the first sound signal and the short-time energy of the ambient sound and the short-time energy of the coherent sound under the first channel, and the relationship between the short-time energy of the second sound signal and the short-time energy of the ambient sound and the short-time energy of the coherent sound under the second channel are calculated in the frequency domain under the linear combination relationship. In the embodiment of the application, the relationship between the short-time energy of the first sound signal and the short-time energy of the ambient sound and the short-time energy of the coherent sound under the first channel is denoted as the linear combination energy representation of the first sound signal. In addition, the relationship between the short-time energy of the second sound signal and the short-time energy of the ambient sound and the short-time energy of the coherent sound under the second channel is denoted as the linear combination energy representation of the second sound signal.
[0074] is denoted as and respectively. Taking the linear combination relationship as an example, the linear combination relationship is converted to the frequency domain, and the linear combination relationship in the frequency domain is obtained as follows:
[0075]
[0076]
[0077] wherein, and are the frequency domain representations of the first sound signal and the second sound signal respectively, is the frequency domain representation of the coherent sound , and are the frequency domain representations of the ambient sound under the first channel and the ambient sound under the second channel respectively. is the frequency domain representation of the difference factor .
[0078] wherein, denotes the time frame index, denotes the frequency point index. In order to keep the formula writing simple, the above formula can omit . Thus, the linear combination relationship in the frequency domain in the simplified form can be denoted as:
[0079]
[0080]
[0081] that is, in the frequency domain, the frequency domain representation of the first sound signal is the frequency domain representation of the coherent sound of the first channel a frequency domain representation of the ambient sound of the first channel a result of the linear addition; a frequency domain representation of the second sound signal a frequency domain representation of the coherent sound of the second channel a frequency domain representation of the ambient sound of the second channel a result of the linear addition; a frequency domain representation of the second sound signal a frequency domain representation of the difference factor a frequency domain representation of the coherent sound of the first channel a product of the frequency domain representation of the coherent sound of the first channel
[0082] Based on the linear combination relationship in the frequency domain, the short-time energy of the first sound signal can be determined as the sum of the short-time energy of the coherent sound of the first channel and the short-time energy of the ambient sound of the first channel, and the short-time energy of the second sound signal can be determined as the sum of the short-time energy of the coherent sound of the second channel and the short-time energy of the ambient sound of the second channel. Thus, the linear combination energy representation of the first sound signal and the linear combination energy representation of the second sound signal can be obtained.
[0083] Further, in some embodiments, step 121 comprises:
[0084] determining the linear combination energy representation of the first sound signal and the linear combination energy representation of the second sound signal based on the linear combination relationship and an identical relationship between the short-time energy values of the ambient sound in different channels.
[0085] Specifically, in the process of determining the linear combination energy representation of the first sound signal and the linear combination energy representation of the second sound signal, in addition to the linear combination relationship in the frequency domain, the characteristics of the ambient sound in different channels can also be considered. That is, the ambient sound in different channels has the same short-time energy value, that is, the short-time energy values of the ambient sound in different channels are the same.
[0086] Taking the first sound signal as an example, the short-time energy of the first sound signal can be calculated by the following formula:
[0087]
[0088] In the formula, is the short-time energy of the first sound signal, is the frequency domain representation of the first sound signal, represents the short-time average. Based on the above formula for calculating the short-time energy, the short-time energy of the second sound signal can also be calculated as follows:
[0089] Based on the above two, the linear combination energy representation of the first sound signal and the linear combination energy representation of the second sound signal can be obtained, which can be embodied as the following formula:
[0090]
[0091]
[0092] wherein, and are short-time energies of the first sound signal and the second sound signal, respectively, is a short-time energy of the coherent sound under the first channel, is a short-time energy of the coherent sound under the second channel, is a short-time energy of the ambient sound. Here, considering that the short-time energies of the ambient sound under different channels are equivalent, the short-time energy of the ambient sound can be expressed as in both the linear combination energy representation of the first sound signal and the linear combination energy representation of the second sound signal.
[0093] Step 122, based on the linear combination relationship, determining an energy representation of a cross-correlation coefficient between the first sound signal and the second sound signal.
[0094] Specifically, after obtaining the linear combination relationship of the first sound signal and the second sound signal respectively, the cross-correlation coefficient between the first sound signal and the second sound signal can be calculated, and in the calculation process of the cross-correlation coefficient, the linear combination relationship of the first sound signal and the linear combination relationship of the second sound signal can be substituted into the calculation formula of the cross-correlation coefficient. The cross-correlation coefficient obtained in this way, i.e. the cross-correlation coefficient in the form of short-time energy of coherent sound, ambient sound, etc. In an embodiment of the present application, the cross-correlation coefficient in the form of short-time energy is denoted as the energy representation of the cross-correlation coefficient.
[0095] Here, the calculation formula of the cross-correlation coefficient between the first sound signal and the second sound signal can be:
[0096]
[0097] wherein, is the cross-correlation coefficient, and are frequency domain representations of the first sound signal and the second sound signal, respectively.
[0098] Substituting the above linear combination relationship in the frequency domain into the calculation formula of the cross-correlation coefficient, the following formula can be obtained:
[0099]
[0100] Since the coherent sound and the ambient sound under each channel are uncorrelated, , . In addition, since the ambient sounds under different channels are uncorrelated, .
[0101] Thus, the formula of the cross-correlation coefficient can be expressed as follows, which is the energy representation of the cross-correlation coefficient:
[0102]
[0103] In the formula, the cross-correlation coefficient can be represented by the short-time energy of the first sound signal , the short-time energy of the second sound signal , the short-time energy of the coherent sound , and the frequency domain representation of the difference factor .
[0104] Step 123, based on the linear combination energy representation and the energy representation of the cross-correlation coefficient, the short-time energy values of the first sound signal and the second sound signal, and the cross-correlation coefficient between the first sound signal and the second sound signal, determine the short-time energy values of the ambient sound and the coherent sound, and the difference factor of the coherent sound on the different channels.
[0105] Specifically, since the first sound signal and the second sound signal are both known sound signals, the short-time energy values of the first sound signal and the second sound signal can be directly calculated, and the cross-correlation coefficient between the first sound signal and the second sound signal can be calculated.
[0106] Thus, the short-time energy values of the first sound signal and the second sound signal, and the cross-correlation coefficient between the two, can be substituted into the above linear combination energy representation and the energy representation of the cross-correlation coefficient, to solve the short-time energy values of the ambient sound and the coherent sound, and the difference factor in the above linear combination energy representation and the energy representation of the cross-correlation coefficient, thereby obtaining the short-time energy values of the ambient sound and the coherent sound, and the difference factor of the coherent sound on the different channels.
[0107] Further, based on the above linear combination energy representation and the energy representation of the cross-correlation coefficient, the short-time energy of the ambient sound and the short-time energy of the coherent sound , and the difference factor can be solved as follows:
[0108]
[0109]
[0110]
[0111] wherein,
[0112]
[0113]
[0114] In the method provided in the embodiments of the present application, the short-time energy of the ambient sound and the coherent sound is solved based on the linear combination relationship between the ambient sound and the coherent sound contained in the first sound signal and the second sound signal, thereby providing a condition for realizing sound separation.
[0115] Based on any of the above embodiments, Figure 3 is a flowchart of the method for determining the weight mapping relationship provided in the present application. As shown in Figure 3 , the determination of the weight mapping relationship comprises:
[0116] Step 310, constructing a linear composition relationship of an estimated sound component, the estimated sound component being an estimated ambient sound or an estimated coherent sound, the linear composition relationship including initial weights of the first sound signal and the second sound signal respectively.
[0117] Specifically, in order to realize sound separation for the first sound signal and the second sound signal, a linear composition relationship for the coherent sound or for the ambient sound can be established. It can be understood that the linear composition relationship here is used to estimate the ambient sound or the coherent sound in the form of weighted summation of the first sound signal and the second sound signal.
[0118] In the embodiments of the present application, the estimated value for the ambient sound can be denoted as an estimated ambient sound, and the estimated value for the coherent sound can be denoted as an estimated coherent sound. On this basis, a linear composition relationship can be constructed for the estimated ambient sound, or a linear composition relationship can be constructed for the estimated coherent sound.
[0119] For example, the linear composition relationship constructed for the estimated coherent sound can be expressed in the following form:
[0120]
[0121] In the formula, is the estimated coherent sound, and are initial weights of the first sound signal and the second sound signal respectively.
[0122] Step 320, determining an estimation error of the estimated sound component based on the linear composition relationship and the linear combination relationship.
[0123] Specifically, after obtaining the linear combination relationship for the estimated sound component, the estimated sound component can be subtracted from the real sound component to obtain the estimation error. For example, the estimated coherent sound can be subtracted from the real coherent sound, or the estimated ambient sound can be subtracted from the real ambient sound.
[0124] In the process of calculating the estimation error, the linear combination relationship can be substituted into the calculation formula of the estimation error, so that the formula representation of the estimation error can be in the following form:
[0125]
[0126] In the formula, is the estimation error for the coherent sound, and are the real coherent sound and the estimated coherent sound, respectively. By substituting the linear combination relationship into the linear combination relationship of the estimated coherent sound, the estimation error represented by the initial weight and , the difference factor , and the real coherent sound and the real ambient sound can be obtained.
[0127] Step 330, linearly fitting the initial weight in the linear combination relationship to obtain the weight mapping relationship, aiming at the estimation error being irrelevant to the first sound signal and the second sound signal, respectively.
[0128] Specifically, after obtaining the estimation error represented by the initial weight and the parameters in the linear combination relationship, the linear fitting for the initial weight can be performed, and the representation form of the optimal estimation obtained by the linear fitting for the initial weight can be taken as the weight mapping relationship. Here, the optimal estimation of the initial weight is the sound separation weight, and the representation form of the optimal estimation is the weight mapping relationship, which can reflect the relationship between the short-time energy of the coherent sound and the ambient sound and the sound separation weight.
[0129] The linear fitting for the initial weight can be implemented based on the least squares (LS). In the process of linear fitting based on the least squares, the optimal estimation of the initial weight is taken when the estimation error is irrelevant to the first sound signal and the estimation error is irrelevant to the second sound signal.
[0130] In specific implementation, the objective function can be constructed by taking the estimation error being irrelevant to the first sound signal and the estimation error being irrelevant to the second sound signal, and the formula representation of the estimation error is substituted into the above objective function for solving, so that the formula representation of the initial weight, i.e. the weight mapping relationship, is obtained.
[0131] Based on any of the above embodiments, in step 330, the initial weight in the linear combination relationship is linearly fitted with the target of making the estimation error irrelevant to the first sound signal and the second sound signal respectively, to obtain the weight mapping relationship, including:
[0132] constructing a first cross-correlation function between the estimation error and the first sound signal, and a second cross-correlation function between the estimation error and the second sound signal;
[0133] with the target of making the values of the first cross-correlation function and the second cross-correlation function both zero, substituting the estimation error and the linear combination relationship into the first cross-correlation function and the second cross-correlation function to obtain the weight mapping relationship.
[0134] Specifically, for the linear fitting target of making the estimation error irrelevant to the first sound signal and the second sound signal, the first cross-correlation function and the second cross-correlation function can be established respectively.
[0135] wherein the first cross-correlation function is the cross-correlation function between the estimation error and the first sound signal, and the value of the first cross-correlation function can be set to 0 for the target of making the estimation error irrelevant to the first sound signal, that is, specifically can be expressed in the following form:
[0136]
[0137] wherein, and are the estimation error and the first sound signal respectively, denotes short-time average.
[0138] Similarly, the second cross-correlation function is the cross-correlation function between the estimation error and the second sound signal, and the value of the second cross-correlation function can be set to 0 for the target of making the estimation error irrelevant to the second sound signal, that is, specifically can be expressed in the following form:
[0139]
[0140] wherein, and are the estimation error and the second sound signal respectively, denotes short-time average.
[0141] After obtaining the first cross-correlation function and the second cross-correlation function, substituting the formulaic expression of the estimation error and the linear combination relationship for the first sound signal and the second sound signal into the first cross-correlation function and the second cross-correlation function, the formulaic expression of the initial weight can be solved, that is, the weight mapping relationship in the following form:
[0142]
[0143]
[0144] wherein the sound separation weight and are expressed in terms of the short-time energies of the ambient sound and the coherent sound respectively , and a difference factor .
[0145] Based on any of the above embodiments, a sound separation method can comprise the following steps:
[0146] First, a first sound signal and a second sound signal are obtained.
[0147] Second, the short-time energy value of the first sound signal and the short-time energy value of the second sound signal are calculated, and the cross-correlation coefficient between the first sound signal and the second sound signal is calculated.
[0148] Subsequently, based on the values of , and , the short-time energy of the ambient sound and the short-time energy of the coherent sound , and the difference factor are calculated.
[0149] Then, the values of , and are substituted into the weight mapping relationship to obtain the sound separation weight and .
[0150] Finally, based on the sound separation weight and , the first sound signal and the second sound signal are weighted and summed (i.e., the value of is calculated), the estimated coherent sound is obtained, and is taken as the estimated ambient sound under the first channel, and is taken as the estimated ambient sound under the second channel, and thus the sound separation is achieved.
[0151] In the sound separation method provided in the embodiments of the present application, only matrix operation and Fourier transform are needed to realize sound separation, and the sound separation process is simple in calculation and easy to program and realize. In addition, the method has good adaptability to various types of sound signals and can adapt to various stereo sound scenes. In addition, the method can achieve a better sound separation effect to realize efficient and stable sound separation.
[0152] The sound separation device provided in the present application is described below, and the sound separation device described below can be correspondingly referred to the sound separation method described above.
[0153] Figure 4 is a structural schematic diagram of the sound separation device provided in the present application. As shown in the figure, Figure 4 The device comprises:
[0154] The signal acquisition unit 410 is configured to acquire a first sound signal and a second sound signal, and the first sound signal and the second sound signal are collected by different channels.
[0155] The parameter acquisition unit 420 is configured to determine a short-time energy value of an ambient sound and a coherent sound respectively, and a difference factor of the coherent sound on different channels based on the short-time energy values of the first sound signal and the second sound signal respectively, and a linear combination relationship between the ambient sound and the coherent sound contained in the first sound signal and the second sound signal.
[0156] The weight mapping unit 430 is configured to map the short-time energy values of the ambient sound and the coherent sound respectively, and the difference factor to a sound separation weight based on a weight mapping relationship, and the weight mapping relationship is obtained by linear fitting based on the linear combination relationship.
[0157] The sound separation unit 440 is configured to separate the ambient sound and the coherent sound from the first sound signal and the second sound signal based on the sound separation weight.
[0158] Based on any of the above embodiments, the parameter acquisition unit is specifically configured to:
[0159] Determine a linear combination energy representation of the first sound signal and the second sound signal respectively based on the linear combination relationship.
[0160] Determine an energy representation of a cross-correlation coefficient between the first sound signal and the second sound signal based on the linear combination relationship.
[0161] determining the short-time energy values of the ambient sound and the coherent sound respectively, and the difference factor of the coherent sound on the different channels based on the linear combination energy representation and the energy representation of the cross-correlation coefficient, the short-time energy values of the first sound signal and the second sound signal respectively, and the cross-correlation coefficient between the first sound signal and the second sound signal.
[0162] According to any one of the above embodiments, the parameter obtaining unit is specifically configured to:
[0163] determining the linear combination energy representation of the first sound signal and the second sound signal based on the linear combination relationship and the equal relationship between the short-time energy values of the ambient sound under the different channels.
[0164] According to any one of the above embodiments, the weight mapping unit is further configured to:
[0165] constructing a linear composition relationship of an estimated sound component, the estimated sound component being an estimated ambient sound or an estimated coherent sound, the linear composition relationship including initial weights of the first sound signal and the second sound signal respectively;
[0166] determining an estimation error of the estimated sound component based on the linear composition relationship and the linear combination relationship;
[0167] performing linear fitting on the initial weights in the linear composition relationship to obtain the weight mapping relationship, with the estimation error being irrelevant to the first sound signal and the second sound signal respectively as the target.
[0168] According to any one of the above embodiments, the weight mapping unit is specifically configured to:
[0169] constructing a first cross-correlation function between the estimation error and the first sound signal, and a second cross-correlation function between the estimation error and the second sound signal;
[0170] substituting the estimation error and the linear combination relationship into the first cross-correlation function and the second cross-correlation function to obtain the weight mapping relationship, with the values of the first cross-correlation function and the second cross-correlation function being zero as the target.
[0171] According to any one of the above embodiments, the linear combination relationship includes that the coherent sound is added to the ambient sound of the first channel to obtain the first sound signal, and the coherent sound is multiplied by the difference factor and then added to the ambient sound of the second channel to obtain the second sound signal.
[0172] Figure 5 An example of a schematic diagram of the physical structure of an electronic device is shown as follows: Figure 5As shown, the electronic device can include a processor 510, a communications interface 520, a memory 530, and a communications bus 540, wherein the processor 510, the communications interface 520, and the memory 530 complete mutual communication through the communications bus 540. The processor 510 can invoke a logical instruction in the memory 530 to execute a sound separation method, which includes:
[0173] Obtaining a first sound signal and a second sound signal, the first sound signal and the second sound signal being obtained by different channels;
[0174] Based on the short-time energy values of the first sound signal and the second sound signal respectively, and the linear combination relationship between the ambient sound and the coherent sound contained in the first sound signal and the second sound signal respectively, determining the short-time energy values of the ambient sound and the coherent sound respectively, and the difference factor of the coherent sound on the different channels;
[0175] Based on a weight mapping relationship, mapping the short-time energy values of the ambient sound and the coherent sound respectively, and the difference factor, to a sound separation weight, the weight mapping relationship being obtained by linear fitting based on the linear combination relationship;
[0176] Based on the sound separation weight, separating the ambient sound and the coherent sound from the first sound signal and the second sound signal.
[0177] In addition, the logical instruction in the memory 530 described above can be implemented in the form of a software function unit and sold or used as an independent product, which can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application or the part that contributes to the related art or part of the technical solutions can be embodied in the form of a software product, which is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various program code storage media.
[0178] In another aspect, the present application also provides a computer program product, which comprises a computer program, the computer program being stored in a non-transitory computer-readable storage medium and executable by a processor, so that the computer can execute the sound separation method provided by the above method, which comprises:
[0179] Based on the linear combination relationship between the ambient sound and the coherent sound contained in the first sound signal and the second sound signal respectively, and the short-time energy values of the ambient sound and the coherent sound respectively, and the difference factor of the coherent sound on the different channels is determined;
[0180] Based on the weight mapping relationship, the short-time energy values of the ambient sound and the coherent sound respectively, and the difference factor are mapped to the sound separation weight, and the weight mapping relationship is obtained by linear fitting based on the linear combination relationship;
[0181] Based on the sound separation weight, the ambient sound and the coherent sound are separated from the first sound signal and the second sound signal.
[0182] In another aspect, the present application also provides a non-transitory computer-readable storage medium, which stores a computer program, and the computer program is executable by a processor to implement the sound separation method provided by the above method, which comprises:
[0183] Obtaining the first sound signal and the second sound signal, the first sound signal and the second sound signal are collected by different channels;
[0184] Based on the linear combination relationship between the ambient sound and the coherent sound contained in the first sound signal and the second sound signal respectively, and the short-time energy values of the ambient sound and the coherent sound respectively, and the difference factor of the coherent sound on the different channels is determined;
[0185] Based on the weight mapping relationship, the short-time energy values of the ambient sound and the coherent sound respectively, and the difference factor are mapped to the sound separation weight, and the weight mapping relationship is obtained by linear fitting based on the linear combination relationship;
[0186] Based on the sound separation weight, the ambient sound and the coherent sound are separated from the first sound signal and the second sound signal.
[0187] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected to achieve the purposes of the embodiments according to actual needs. Those skilled in the art can understand and implement without creative labor.
[0188] Through the description of the above embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software and the necessary general hardware platform, and of course can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.
[0189] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A sound separation method characterized by, The method comprises: acquiring a first sound signal and a second sound signal, the first sound signal and the second sound signal being acquired by different channels; determining, based on respective short-time energy values of the first sound signal and the second sound signal and a linear combination relationship between ambient sound and coherent sound contained in the first sound signal and the second sound signal respectively, respective short-time energy values of the ambient sound and the coherent sound and a difference factor of the coherent sound on the different channels; mapping, based on a weight mapping relationship, the respective short-time energy values of the ambient sound and the coherent sound and the difference factor into sound separation weights, the weight mapping relationship being obtained by linear fitting based on the linear combination relationship; separating, based on the sound separation weights, the ambient sound and the coherent sound from the first sound signal and the second sound signal; the determining, based on respective short-time energy values of the first sound signal and the second sound signal and a linear combination relationship between ambient sound and coherent sound contained in the first sound signal and the second sound signal respectively, respective short-time energy values of the ambient sound and the coherent sound and a difference factor of the coherent sound on the different channels, comprises: determining, based on the linear combination relationship, linear combination energy representations of the first sound signal and the second sound signal, the short-time energy values of the ambient sound under the different channels being the same; determining, based on the linear combination relationship, energy representations of cross-correlation coefficients between the first sound signal and the second sound signal; determining, based on the linear combination energy representations and the energy representations of the cross-correlation coefficients, the respective short-time energy values of the first sound signal and the second sound signal and the cross-correlation coefficients between the first sound signal and the second sound signal, the respective short-time energy values of the ambient sound and the coherent sound and the difference factor of the coherent sound on the different channels; the determining of the weight mapping relationship comprises: constructing a linear composition relationship of an estimated sound component, the estimated sound component being an estimated ambient sound or an estimated coherent sound, the linear composition relationship comprising initial weights of the first sound signal and the second sound signal respectively; determining an estimation error of the estimated sound component based on the linear composition relationship and the linear combination relationship; performing linear fitting on the initial weights in the linear composition relationship to obtain the weight mapping relationship, with the estimation error being irrelevant to the first sound signal and the second sound signal respectively as a target.
2. The sound separation method of claim 1, wherein, the determining, based on the linear combination relationship, linear combination energy representations of the first sound signal and the second sound signal, comprises: determining, based on the linear combination relationship and an identical relationship between the short-time energy values of the ambient sound under the different channels, linear combination energy representations of the first sound signal and the second sound signal.
3. The voice separation method according to claim 1, wherein The initial weight in the linear combination relationship is linearly fitted with the estimation error being irrelevant to the first sound signal and the second sound signal respectively, to obtain the weight mapping relationship, including: A first cross-correlation function between the estimation error and the first sound signal is constructed, and a second cross-correlation function between the estimation error and the second sound signal is constructed; The estimation error and the linear combination relationship are substituted into the first cross-correlation function and the second cross-correlation function, with the values of the first cross-correlation function and the second cross-correlation function being zero, to obtain the weight mapping relationship.
4. The sound separation method according to any one of claims 1 to 3, characterized by, The linear combination relationship includes that the coherent sound is added to the ambient sound of the first channel to obtain the first sound signal, and the coherent sound is multiplied by the difference factor and then added to the ambient sound of the second channel to obtain the second sound signal.
5. A sound separating device, characterized in that Including: The signal acquisition unit is configured to acquire a first sound signal and a second sound signal, the first sound signal and the second sound signal being collected by different channels; The parameter acquisition unit is configured to determine a short-time energy value of the ambient sound and a short-time energy value of the coherent sound, and a difference factor of the coherent sound on the different channels based on a short-time energy value of the first sound signal and a short-time energy value of the second sound signal, and a linear combination relationship between ambient sound and coherent sound included in the first sound signal and the second sound signal; The weight mapping unit is configured to map the short-time energy value of the ambient sound and the short-time energy value of the coherent sound, and the difference factor to a sound separation weight based on a weight mapping relationship, the weight mapping relationship being obtained by linear fitting based on the linear combination relationship; The sound separation unit is configured to separate the ambient sound and the coherent sound from the first sound signal and the second sound signal based on the sound separation weight; The parameter acquisition unit is specifically configured to: Determine a linear combination energy representation of the first sound signal and the second sound signal based on the linear combination relationship, the short-time energy value of the ambient sound under the different channels being the same; Determine an energy representation of a cross-correlation coefficient between the first sound signal and the second sound signal based on the linear combination relationship; Determine the short-time energy value of the ambient sound and the short-time energy value of the coherent sound, and the difference factor of the coherent sound on the different channels based on the linear combination energy representation, the energy representation of the cross-correlation coefficient, the short-time energy value of the first sound signal and the second sound signal, and the cross-correlation coefficient between the first sound signal and the second sound signal; The weight mapping unit is specifically configured to: Construct a linear combination relationship of an estimated sound component, the estimated sound component being an estimated ambient sound or an estimated coherent sound, the linear combination relationship including an initial weight of the first sound signal and the second sound signal; Determine an estimation error of the estimated sound component based on the linear combination relationship and the linear combination relationship; The initial weight in the linear composition relationship is linearly fitted with the estimated error being irrelevant to the first sound signal and the second sound signal respectively, to obtain the weight mapping relationship.
6. An electronic device comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, The processor implements the sound separation method according to any one of claims 1 to 4 when executing the computer program.
7. A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the sound separation method according to any one of claims 1 to 4.
8. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the sound separation method according to any one of claims 1 to 4. The computer program is executed by the processor to implement the sound separation method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Non-orthogonal joint diagonalization instantaneous blind source separation method based on double iteration
CN103780522A
Method and device for positioning sound area, storage medium and electronic equipment
CN113380267A